A High-Dimensional Small-Sample Data Augmentation Method Based on Acceptable Regions

By demarcating acceptable areas under the conditions of small sample data sets and using multivariate joint probability distribution sampling, the problems of inaccurate virtual data generation and error in feature combination in the prior art are solved, and virtual data that conforms to the characteristics of small sample data is generated in high-dimensional space.

CN115169470BActive Publication Date: 2025-05-30XIAN TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210840445.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2025-05-30
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to generate reasonable virtual data under the conditions of small sample data sets, there are errors in the combination of virtual data features, and it is difficult to define an effective range to limit the generation of virtual data.

Method used

By analyzing the data, identifying the overall trend and distribution trend of input features, demarcate acceptable areas, including the general allowable range and mutual influence existence range, and sampling in the acceptable areas using multiple joint probability distributions to generate virtual data that conforms to the characteristics of small sample data.

Benefits of technology

It effectively avoids the problem of combining errors of virtual input features, and the generated virtual data is closer to real samples, avoids the uncertainty brought by intermediate models, and can generate virtual data that meets the characteristics of small sample data in high-dimensional space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115169470B_ABST
    Figure CN115169470B_ABST
Patent Text Reader

Abstract

The present invention is a method for augmenting high-dimensional small-sample data based on an acceptable region, which overcomes the problems in the prior art that it is difficult to generate reasonable virtual data, there are feature combination errors in the virtual data, and it is difficult to delimit an effective range to restrict the generation of virtual data. The present invention includes the following steps: Step 1: Analyze the data and determine the overall trend of the data; Step 2: Determine the distribution trend between the input features; Step 3: Delimit the acceptable region: For each feature, delimit its acceptable region Q, and the acceptable region consists of two parts, one part is the general allowable range Q a , and the other part is the range where mutual influence exists Q β ; Step 4: Generate virtual data: First, sample based on the multivariate joint probability distribution of the small sample and the acceptable region within the acceptable region Q X in the input feature space, and then map it to the output feature space through the relationship between y q and X, and finally form virtual data within the acceptable region Q Y in the output feature space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention belongs to the technical field of virtual data augmentation, and relates to a method for augmenting high-dimensional small-sample data based on an acceptable region. It can be used to generate virtual samples for time series in practical problems under the condition of a small-sample data set, and use the generated samples for modeling. Background Art:

[0002] Small samples have the problem that it is impossible to construct an effective machine learning model when the data volume is scarce. There are mainly two technical approaches. One is data augmentation, and the other is model optimization. The present invention belongs to the category of data augmentation methods.

[0003] Currently, the mainstream data augmentation methods include virtual sample augmentation techniques based on distribution and virtual sample augmentation techniques based on prior knowledge. Virtual sample augmentation techniques based on distribution include Bootstrap, Moving Trend Diffusion (MTD), etc. Bootstrap is a resampling technique. Its advantage is that it can simulate the real distribution through the sampling distribution, but its disadvantage is that it does not generate new samples. This method is just a redistribution of the original sample set. MTD is a recognized effective virtual sample augmentation technique in some scenarios. However, since this method generates each input feature separately, there is a problem that the effectiveness of virtual samples is poor due to incorrect combination of virtual data. Virtual sample augmentation techniques based on prior knowledge generate virtual data by using prior knowledge or extracting knowledge from limited samples. In this type of method, the accuracy of the prior knowledge directly determines the quality of the virtual samples. Therefore, whether accurate prior knowledge can be obtained becomes a determining factor for using this type of method. The method of extracting knowledge from limited samples to generate virtual data has problems such as how to extract effective knowledge and judge whether the extracted knowledge is applicable to this type of object. Summary of the Invention:

[0004] The purpose of the present invention is to provide a method for augmenting high-dimensional small-sample data based on an acceptable region, which overcomes the problems in the prior art that it is difficult to generate reasonable virtual data, there are feature combination errors in the virtual data, and it is difficult to delimit an effective range to limit the generation of virtual data. The present invention avoids the combination error problem when generating virtual input features and avoids the uncertainty brought by various intermediate models, and can effectively generate virtual data that conforms to the characteristics of small-sample data in a high-dimensional space.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A method for augmenting high-dimensional small-sample data based on an acceptable region, characterized by comprising the following steps:

[0007] Step 1: Analyze the data and determine the overall trend of the data;

[0008] Step 2: Determine the distribution trend among the input features;

[0009] Step 3: Define the acceptable region:

[0010] For each feature, define its acceptable region Q, which consists of two parts. One part is the general allowable range Q a , and the other part is the range where mutual influence exists Q β ;

[0011] The relationship between the acceptable region Q and the general allowable range Q a , and the range where mutual influence exists Q β is as follows:

[0012]

[0013] Step 4: Generate virtual data:

[0014] First, sample based on the multivariate joint probability distribution of the small sample and the acceptable region within the acceptable region Q X in the input feature space, and then map it to the output feature space through the relationship between y q and X, finally forming virtual data within the acceptable region Q Y in the output feature space.

[0015] Step 3 includes the following steps

[0016] 3.1 Set the general allowable range;

[0017] 3.2 Set the range where mutual influence exists;

[0018] 3.3 Determine the acceptable region of the input features;

[0019] 3.4 Determine the acceptable region of the output features.

[0020] Step 4 includes the following steps:

[0021] 4.1 Generate virtual input features within the acceptable region Q X :

[0022] Construct a skewed distribution with each input feature X in the acceptable region Q obtained in Step 3 as the mode and the distribution trend as the median with the range restricted within the acceptable region of the input features to obtain the multivariate joint probability distribution U of all input features combined, and finally perform a smooth transition among the discrete to obtain the virtual data within the acceptable region Q of the input featuresX For any y within the output feature acceptable region Q Y there is a corresponding q joint probability distribution, and sampling is performed within this joint probability distribution;

[0023]

[0024] 4.2 Calculate the output feature y' corresponding to each group of virtual input features according to the relationship between y and x 1 , x 2 obtained in step 1, where i Y' i ={y' i |1≤i≤n'}; i i

[0025]

[0026] 4.3 Finally, obtain the virtual sample set D',

[0027] D'={(X' i , Y' i )|1≤i≤n', i∈N}.

[0028] In step 1, the least squares method is used for linear fitting to describe the overall change trend of each output feature and all input feature data.

[0029] In step 2, the least squares method is used for linear fitting to describe the distribution trend of the input features in the data space.

[0030] Compared with the prior art, the advantages and effects of the present invention are as follows:

[0031] (1) The present invention provides a virtual data augmentation method for high-dimensional small samples. Compared with existing methods, by defining an acceptable region to determine the effective range of virtual data generation, the virtual data generated in this region is closer to real samples;

[0032] (2) The present invention avoids the combination error problem in generating virtual input features through high-dimensional sampling, thereby greatly increasing the effectiveness of the generated virtual data and preventing the situation where the virtual samples are too far from the real samples;

[0033] (3) The present invention uses the least squares method to describe the trend of data and generates virtual samples based on the joint probability distribution on this basis, avoiding the uncertainty brought by various intermediate models and being able to effectively generate virtual data conforming to the characteristics of small sample data in a high-dimensional space. BRIEF DESCRIPTION OF THE DRAWINGS:

[0034] Figure 1 is the overall flowchart of the present invention;

[0035] Figure 2 The original full sample set and the extracted small sample set of the experimental objects;

[0036] Figure 3 The overall trend of the small sample data set;

[0037] Figure 4 The distribution trends of the input features F1 and F2 (in different coordinate systems);

[0038] Figure 5 The generalized allowable ranges of the input features F1 and F2 (in different coordinate systems);

[0039] Figure 6 The range where the mutual influence of the input features F1 and F2 exists (in different coordinate systems);

[0040] Figure 7 The acceptable regions of the input features F1 and F2 (in different coordinate systems);

[0041] Figure 8 The acceptable regions of the input features F1 and F2 (in the same coordinate system);

[0042] Figure 9 The virtual input features generated within the acceptable regions of the input features;

[0043] Figure 10 The comparison between the virtual input features and the original full sample input features;

[0044] Figure 11 The virtual samples generated from the small sample set;

[0045] Figure 12 The comparison between the virtual samples and the original full samples;

[0046] Figure 13 The comparison of the overall trends of different data sets. Specific implementation manner:

[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] The present invention relates to a method for augmenting high-dimensional small-sample data based on an acceptable region, which determines the effective range for generating virtual data by delimiting the acceptable region and solves the problem of combinatorial errors between various input features of the virtual data through high-dimensional sampling. First, the overall trend of the data and the distribution trend of the input features are described by the least squares method; then, prior knowledge is used to determine the general allowable range of the data, and the deviation limitation method or the deviation expansion method is selected according to the distribution of the data to determine the range where the mutual influence of the data exists, thereby delimiting the acceptable region and restricting the range for generating virtual data; finally, virtual data is generated within the acceptable region according to the overall trend of the data through a reasonable sampling method. The present invention mainly solves the following problems: (1) it is difficult to generate reasonable virtual data with less prior knowledge; (2) there are combinatorial errors of features in generating virtual data through one-dimensional samples; (3) it is difficult to delimit an effective range to restrict the generation of virtual data.

[0049] The present invention specifically includes the following steps:

[0050] Step 1: Analyze the data and determine the overall trend of the data.

[0051] Use the least squares method for linear fitting to describe the overall change trend of each output feature and all input feature data;

[0052] Step 2: Determine the distribution trend among the input features.

[0053] Use the least squares method for linear fitting to describe the distribution trend of the input features in the data space.

[0054] Step 3: Delimit the acceptable region.

[0055] For each feature, delimit its acceptable region Q. The acceptable region consists of two parts, one part is the general allowable range Q a , and the other part is the range where the mutual influence exists Q β .

[0056] The general allowable range only stipulates the theoretical upper and lower limits of each feature. This range is usually determined by physical factors such as materials and forms, and is an objective limitation. Especially for some pipeline products, each parameter has a rated range and a maximum fluctuation range, and this maximum fluctuation range is the general allowable range here.

[0057] The range where the mutual influence exists needs to be determined according to the distribution among the features. The reason is that there may be some correlation factors among multiple input features, such that one input feature will change with the change of other input features, and the output feature will also change with the change of the input features.

[0058] Therefore, the acceptable region Q and the general allowable range Q a、Interaction scope Q β The relationship is:

[0059]

[0060] Step 4: Generate dummy data

[0061] There are two ways to generate virtual data. One is to directly sample in the acceptable region Q, and the other is to first sample in the acceptable region Q of the input feature space. X Based on the multivariate joint probability distribution sampling of small samples and acceptable regions, and then through y q The relationship between X and X is mapped to the output feature space, and finally forms the acceptable region Q in the output feature space. Y Virtual data inside.

[0062] Example:

[0063] See also Figure 1 The basic idea of ​​the method of the present invention is to first determine the overall trend of the data and the distribution trend of the input features through data analysis, then define the acceptable area in the data space, limit the scope of virtual data generation, and finally generate virtual data based on the overall trend of the data and the acceptable area.

[0064] Based on the above basic idea, the present invention provides a high-dimensional small sample data expansion method based on acceptable region, comprising the following steps:

[0065] Step 1: Analyze the data and identify the overall trend of the data

[0066] For the existing original full sample set n # =326 and the small sample data set D extracted from it = {(X i ,Y i )|1≤i≤n,i∈N}, n=34, data visualization is as follows Figure 2 As shown, we can determine the overall trend of the data: as F1 increases and F2 decreases, Capacity has an obvious overall upward trend.

[0067] Use the least squares method to obtain the overall change trend of the output feature y and all input feature data X, that is, to describe the relationship between y and x 1 ,x 2 The relationship between Figure 3 shown.

[0068] y=f(X)=f(x 1 ,x 2 )=0.9452+0.0003*(x 1 )-0.0001*(x 2)

[0069] Step 2: Determine the distribution trend among the input features

[0070] Describe x with a unary linear relationship 1 and x 2 for the overall distribution trend therebetween. As Figure 4 shown

[0071] F(x 1 , x 2 )

[0072] then for x 1 there is

[0073]

[0074] then for x 2 there is

[0075]

[0076] Step 3: Define the acceptable region

[0077] 3.1 Set the general allowable range

[0078] The object in this example is a lithium-ion battery, whose features have clear meanings. Through prior knowledge, it can be known that the general allowable range of x 1 is: [1400, 3600], the general allowable range of x 2 is: [1100, 1600], and the general allowable range of y is: [0, 2]. As Figure 5 shown

[0079] 3.2 Set the range where mutual influence exists

[0080] In this example, there are two input features. It is necessary to calculate the range where mutual influence exists for each input feature, and calculate the range where mutual influence exists for the output feature y

[0081] 3.2.1 Calculate the range where mutual influence exists for x 1 using the deviation limit method

[0082] Calculate the value for each distance and denote it as Take its maximum and minimum values to get

[0083]

[0084] In this example, the additional margins are all 30% of the maximum and minimum deviations. Then

[0085]

[0086] Then x 1 The range of the mutual influence of

[0087]

[0088] That is

[0089]

[0090] 3.2.2 Calculate x 2 using the deviation limit method for the range of the mutual influence.

[0091] Calculate the value for each distance and denote it as Take its maximum and minimum values to get

[0092]

[0093] In this example, the additional margins are all 30% of the maximum and minimum deviations, then

[0094]

[0095] Then the range of the mutual influence of x 2 is:

[0096]

[0097] That is

[0098]

[0099] x 1 and the range of the mutual influence of x 2 are as shown in Figure 6 and the acceptable region is as shown in Figure 7 as shown.

[0100] 3.2.3 Calculate the intersection of the ranges of the mutual influence of two input features. To facilitate the calculation of the intersection, normalize the two ranges of the mutual influence to the same coordinate system.

[0101] Since the relationship between x 1 and x 2 is known, normalize it to a coordinate system with x 1 as the abscissa and x 2 as the ordinate. Then the range of the mutual influence of x 1 can be transformed into:

[0102]

[0103] That is

[0104]

[0105] Then the range of existence Q of the mutual influence of the input features Xβ is

[0106]

[0107] 3.2.4 Use the deviation limit method to calculate the range of existence of the mutual influence of the output feature y

[0108] Calculate for each y i the value of the distance f(X i ), denoted as e yi , and take its maximum and minimum values to obtain

[0109]

[0110] In this example, the increased margins are all 30% of the maximum and minimum deviations, then

[0111]

[0112] Then the range of existence Q of the mutual influence of y Yβ is

[0113] [0.9452 + 0.0003*(x 1 ) - 0.0001*(x 2 ), 0.9452 + 0.0003*(x 1 ) - 0.0001*(x 2 ) + 0.03]

[0114] That is

[0115] [0.9252 + 0.0003*(x 1 ) - 0.0001*(x 2 ), 0.9752 + 0.0003*(x 1 ) - 0.0001*(x 2 )]

[0116] 3.3 Determine the acceptable region of the input features

[0117] Intersect the range of existence Q of the mutual influence of the input features Xβ with the general allowable range [1400, 3600] of x 1 and the general allowable range [1100, 1600] of x 2 to obtain the final acceptable region Q of the input features X . As Figure 8 shown

[0118] 3.4 Determine the acceptable region of the output features

[0119] Take the intersection of the range Q of the mutual influence of the output features Yβ with the general allowable range [0, 2] to obtain the acceptable region Q of the final output features Y .

[0120] Step 4: Generate virtual data

[0121] In this example, first sample within the acceptable region Q of the input feature space X , and then map it to the output feature space through the relationship between y q and X, and generate virtual data within the acceptable region Q of the output feature space Y .

[0122] 4.1 Generate virtual input features within the acceptable region Q X

[0123] Construct a skewed distribution with each input feature X as the mode in the acceptable region Q obtained in Step 3 with the distribution trend as the median and the range restricted within the acceptable region of the input features . Combine the multivariate joint probability distribution U obtained from all input features, and finally perform a smooth transition between the discrete to obtain, within the acceptable region Q of the input features, for any y within the range of the acceptable region Q of the output features X , there is a corresponding Y y q joint probability distribution, and sample within this joint probability distribution. The obtained virtual input features X' are as shown, and the comparison between the virtual input features and the original full-sample input features is as Figure 9 shown Figure 10 .

[0124]

[0125] 4.2 Calculate the output feature y' 1 corresponding to each group of virtual input features according to the relationship between y and x 2 obtained in Step 1 . i .

[0126]

[0127] where μ ~ U(-0.02, 0.03), to obtain Y' i .

[0128] Y' i ​={y' i |1 ≤ i ≤ n'}

[0129] 4.3 Finally, obtain the virtual sample set D'.

[0130] D' = {(X' i , Y' i )|1 ≤ i ≤ n', i ∈ N}

[0131] The virtual samples are as Figure 11 shown, and the comparison of the effects with the original full sample set is as Figure 12 shown.

[0132] In order to verify whether the generated virtual sample set is effective, compare the small sample set, the virtual sample set, and the original full sample set in terms of various indicators, and use the least squares method to estimate the overall trend of each sample for comparison. Table 1 shows the comparison of the mean and standard deviation of each feature.

[0133] Table 1 Mean and Standard Deviation of Each Feature

[0134]

[0135] Table 2 shows the comparison of the overall trends of different sample sets, corresponding to Figure 13 .

[0136] Table 2 Overall Trends of Different Sample Sets

[0137]

[0138] As mentioned above, it is only the preferred embodiment of the present invention, and is not used to limit the protection scope of the present invention. Any equivalent structural changes made using the description and drawings of the present invention shall be included in the patent protection scope of the invention.

Claims

1. A method for high-dimensional small-sample data augmentation based on an acceptable region, characterized in that: It includes the following steps: Step 1: Analyze the data and determine the overall trend of the data; Step 2: Determine the distribution trend among the input features; Step 3: Define the acceptable region: For each feature, an acceptable region Q is delimited. The acceptable region consists of two parts. One part is the general allowable range Q a , and the other part is the range Q where mutual influence exists β ; Acceptable region Q and generalized allowable range Q a , and the range Q of mutual influence β have the following relationship: Step 4: Generate virtual data: First, sample based on the multivariate joint probability distribution of the small sample and the acceptable region within the acceptable region Q of the input feature space, and then map through the relationship between y X and X to the output feature space, and finally form virtual data within the acceptable region Q q of the output feature space. Y ​ 2. The method for high-dimensional small-sample data augmentation based on an acceptable region according to claim 1, characterized in that: Step 3 includes the following steps 3.1 Set the general allowable range; 3.2 Set the range where mutual influence exists; 3.3 Determine the acceptable region of the input features; 3.4 Determine the acceptable region of the output features.

3. The method for high-dimensional small-sample data augmentation based on an acceptable region according to claim 2, characterized in that: Step 4 includes the following steps: 4.1 Generate virtual input features within the acceptable region Q X as follows: The acceptable region Q obtained in step 3 X Each input feature is constructed is majority Based on distribution trend is the median The range is limited to the acceptable region of input features The skewed distribution within the y-axis and the multivariate joint probability distribution U obtained by combining all input features are finally obtained in the discrete A smooth transition is made between the input feature acceptable region Q X For any range in the output feature acceptable region Q Y of y q There is a corresponding a joint probability distribution within which to sample; 4.2 Calculate the output feature y' corresponding to each set of virtual input features based on the relationship between y and x obtained in step 1 1 , x 2 between the output feature y' i , Y′ i ={y′ i | 1 ≤ i ≤ n'}; 4.3 Finally obtain the virtual sample set D'; D' = {(X′ i , Y′ i ) | 1 ≤ i ≤ n', i ∈ N}.

4. The method for high-dimensional small-sample data augmentation based on an acceptable region according to claim 3, characterized in that: In Step 1, the overall change trend of each output feature and all input feature data is described by linear fitting with the least squares method.

5. The method for high-dimensional small-sample data augmentation based on an acceptable region according to claim 4, characterized in that: In Step 2, the distribution trend of the input features in the data space is described by linear fitting with the least squares method.

Citation Information

Patent Citations

  • Medical data expansion method based on generative adversarial network

    CN112215339A

  • Comparative knowledge driven iris and periocular confrontation adaptive fusion recognition method

    CN114596622A