Unbalanced sample data enhancement method based on adaptive loss function

The unbalanced sample data is enhanced through the adaptive loss function, which solves the problems of poor model characterization ability and overfitting, improves the prediction accuracy and adaptability of the model, and is suitable for a variety of unbalanced data scenarios.

CN120408190APending Publication Date: 2025-08-01XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471154.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When processing unbalanced sample data, the model characterization ability is poor or there is a tendency to overfit, resulting in insufficient identification ability in key categories, affecting the generalization ability and prediction accuracy of the model.

Method used

The adaptive loss function is used to enhance the imbalanced sample data. Through feature importance analysis and loss function factor grouping, the adaptive loss function is designed for model training, and combined with Z-score normalization processing, the model parameters are optimized.

Benefits of technology

It enhances the model's adaptability and prediction accuracy to imbalanced data, maintains data integrity, is suitable for a variety of imbalanced data scenarios, and improves the generalization ability and prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408190A_ABST
    Figure CN120408190A_ABST
Patent Text Reader

Abstract

The invention discloses an unbalanced sample data enhancement method based on an adaptive loss function, and the method comprises the steps: building a model, and obtaining a data set; preprocessing the data of the data set; performing importance analysis on the features of the model to obtain important features after feature dimension reduction, and dividing a preprocessed data set corresponding to the important features into dense data and sparse data; grouping the dense data and the sparse data in proportion; setting different loss function factors according to the groups, and designing a loss function by using the loss function factors; performing model training according to the loss function; the data set is put into the model and the trained model for prediction, and a prediction value and a test value are obtained; and performing normalization processing on the predicted value and the test value to obtain a loss function average absolute percentage error and a correlation coefficient, and evaluating a model fitting effect and accuracy. According to the method, sample data features are accurately and rapidly obtained by using unbalanced samples, and the prediction precision of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent manufacturing, and relates to a method for enhancing unbalanced sample data based on an adaptive loss function. Background Art

[0002] Unbalanced sample data is one of the classic problems faced in the fields of machine learning and data mining for a long time. The core contradiction lies in the significant difference in the quantity distribution of different category samples. In real scenarios, such as industrial equipment fault warning, production quality control, and industrial Internet of Things security, key categories (such as micro-cracks, metal fatigue signs, component misalignment) often only account for a very small proportion of the overall data. However, the training mechanism of traditional machine learning algorithms based on empirical risk minimization is prone to causing the model to be overly biased towards the majority class, forming an "accuracy trap" - even if the model maintains a high prediction accuracy for the majority class, its recognition ability for the key minority class is severely insufficient. This bias not only causes the distortion of evaluation indicators but also leads to serious consequences in practical applications. For example, it may cause defective products to flow into the market, and a new energy vehicle company once triggered a large-scale recall event due to this.

[0003] In the field of welding process intelligence, the manifestation form of data distribution imbalance has dynamic complexity. When the sample size of curved welding (such as circumferential welds of pipelines, corner welding of special-shaped components) is significantly less than that of straight welding, this imbalance will pose special modeling challenges. Different from conventional cognition, the data scarcity of curved welding often stems from its process particularity: in high-end manufacturing scenarios such as nuclear power plant pressure vessels and aerospace fuel storage tanks, curved welding needs to meet strict process certification standards, resulting in limited accumulation of effective data in actual production (such as only performing key curved welds several times per month), while straight welding, as a basic process, generates a large amount of conventional data. As a result, the samples of curved paths are severely insufficient, leading to extreme class imbalance in the dataset. This data distribution characteristic is likely to cause the model to misjudge the samples of curved paths as outliers, making it difficult to accurately capture their process characteristics, and ultimately affecting the generalization ability and prediction accuracy of the model.

[0004] Currently, the methods for dealing with data imbalance mainly include: 1) undersampling techniques, which achieve class balance by randomly removing majority-class samples. Although it can highlight the characteristics of minority classes and reduce the computational complexity, it will cause the loss of effective process information. Especially when the initial class ratio is extremely different (such as 1:20), excessive undersampling may cause the model to lose the ability to represent the mainstream process characteristics. 2) Algorithms represented by SMOTE generate synthetic samples through interpolation. Although it can maintain the integrity of the original data, it has two risks: on the one hand, random interpolation in the high-dimensional process parameter space may generate physically infeasible virtual samples; on the other hand, the sharp increase in the sample size will significantly increase the computational resource requirements and exacerbate the overfitting tendency of the model. Summary of the Invention

[0005] The present invention aims to solve the technical problems of poor model representation ability or the tendency of model overfitting in the prior art. The present invention provides an unbalanced sample data augmentation method based on an adaptive loss function, and the technical solution adopted is as follows:

[0006] An unbalanced sample data augmentation method based on an adaptive loss function, comprising the steps of:

[0007] S1. Establish a model, obtain a data set, and divide the data set into a training data set and a test data set;

[0008] S2. Preprocess the data in the training data set to obtain a preprocessed data set;

[0009] S3. Analyze the importance of the features of the model to obtain important features after reducing the feature dimension, and divide the preprocessed data set corresponding to the important features into dense data and sparse data;

[0010] S4. Group the dense data and the sparse data according to a ratio, including a large data group, a small dense group, and a large sparse group;

[0011] S5. Set different loss function factors according to the grouping, and design a loss function by using the loss function factors;

[0012] S6. Train the model according to the loss function to obtain a trained model;

[0013] S7. Put the test data set into the model for prediction to obtain a predicted value, and put the test data set into the trained model for testing to obtain a test value;

[0014] S8. Perform normalization processing on the predicted value and the test value to obtain the mean absolute percentage error and the correlation coefficient of the loss function, and evaluate the model fitting effect and accuracy.

[0015] In one embodiment of the present invention, the step S2 includes:

[0016] S21. Remove blank line data;

[0017] S22. Merge the feature vectors;

[0018] S23. Convert the categorical feature vector into a one-hot encoded vector form.

[0019] In one embodiment of the present invention, the step S3 includes:

[0020] Decompose the preprocessed data set corresponding to the important features along the X-axis and Y-axis. The data set with values on both the X-axis and Y-axis is dense data, and the data set with only X-axis or Y-axis values is sparse data.

[0021] In one embodiment of the present invention, step S4 includes:

[0022] The sample group consisting entirely of the dense data is a large dense group; the sample group with a ratio of the dense data to the sparse data of 5:1 is a small sparse group; the sample group with a ratio of the dense data to the sparse data of 1:10 is a large sparse group.

[0023] In one embodiment of the present invention, in step S5:

[0024] The formula of the loss function factor is expressed as:

[0025]

[0026] In formula (1), η is the loss function factor, m represents the number of dense data, n represents the number of sparse data, η m represents the loss factor corresponding to the dense data, and η n represents the loss factor corresponding to the sparse data.

[0027] In one embodiment of the present invention, the loss factors η m corresponding to the dense data of the large dense group, the small sparse group, and the large sparse group, and the loss factor η n corresponding to the sparse data are different:

[0028] In the large dense group,

[0029] In the small sparse group,

[0030] In the large sparse group, Φ is an adaptive parameter, and its range is 0 to 1.

[0031] In one embodiment of the present invention, in step S5:

[0032] The formula of the loss function is expressed as:

[0033]

[0034] In formula (2), mLoss represents the loss function, η is the loss function factor, M represents the total number of each data set, f′ represents the predicted stress result of the test data after passing through the trained model, and f represents the stress result obtained by the test data through the real experiment.

[0035] In one embodiment of the present invention, the step S8 includes:

[0036] Normalize the predicted value and the test value by using Z - score normalization, and the formula of the Z - score normalization is expressed as:

[0037]

[0038] In formula (3), X is a certain value in the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and z is the data after Z - score normalization.

[0039] In one embodiment of the present invention, in the step S8:

[0040] The formula of the mean absolute percentage error of the loss function is expressed as:

[0041]

[0042] In formula (4), MAPE represents the mean absolute percentage error of the loss function, n represents the total number of samples, i represents the i - th sample, y i represents the true value of the i - th sample, represents the predicted value of the i - th sample.

[0043] In one embodiment of the present invention, in the step S8:

[0044] The formula of the correlation coefficient is expressed as:

[0045]

[0046] In formula (5), R 2 represents the correlation coefficient, n represents the total number of samples, i represents the i - th sample, y i represents the true value of the i - th sample, represents the predicted value of the i - th sample, represents the average value of the true values of the samples.

[0047] Advantages of the present invention:

[0048] The unbalanced sample data enhancement method based on an adaptive loss function of the present invention uses a method of adaptively allocating weights for an unbalanced sample set, without adding or reducing sample data, thereby enhancing the prediction accuracy of the model and preserving the integrity of the data; the unbalanced sample data enhancement method of the present invention processes the unbalanced data set using the loss function factor η formula, and then trains the model through the loss function mLoss to find the optimal parameters with the best evaluation index, optimizing the parameter control accuracy; and the unbalanced sample data enhancement method of the present invention has strong adaptability and can be applied to various unbalanced data scenarios, and is applicable to the processing of unbalanced data models such as financial fraud detection and network intrusion recognition, and can perform model learning for different scenarios and different parameter types. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flowchart of the unbalanced sample data enhancement method based on an adaptive loss function provided by an embodiment of the present invention;

[0050] Figure 2 is an object instance model provided by an embodiment of the present invention;

[0051] Figure 3 is a flow chart illustration of the unbalanced sample data enhancement method based on the adaptive loss function of the composite path metal sealing process provided by an embodiment of the present invention;

[0052] Figure 4 is a flow chart illustration of the data preprocessing provided by an embodiment of the present invention;

[0053] Figure 5 is a rectangular parameter importance analysis diagram obtained by analyzing through the random forest importance algorithm provided by an embodiment of the present invention;

[0054] Figure 6 is a comparison diagram of the predicted values and the true values of the MLP model using the mLoss and MSE loss functions provided by an embodiment of the present invention;

[0055] Figure 7 is a flow chart illustration of the training of the multi-layer perceptron (MLP) model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0057] The present invention provides an unbalanced sample data enhancement method based on an adaptive loss function, which uses a method of an adaptive loss function to allocate weights to unbalanced data, enabling fewer data samples to be fully learned, enhancing the adaptability of the model to the data, so as to more accurately and quickly obtain sample data features using unbalanced samples and enhancing the prediction accuracy of the model.

[0058] Referring to the attached Figure 1 , the unbalanced sample data augmentation method based on an adaptive loss function includes the following steps:

[0059] S1. Establish a model, obtain a data set, and divide the data set into a training data set and a test data set;

[0060] S2. Preprocess the data in the training data set to obtain a preprocessed data set;

[0061] S3. Analyze the importance of the features of the model to obtain the important features after reducing the feature dimension, and divide the preprocessed data set corresponding to the important features into dense data and sparse data;

[0062] S4. Group the dense data and sparse data according to a ratio, including a large data group, a small dense group, and a large sparse group;

[0063] S5. Set different loss function factors according to the grouping, and design a loss function using the loss function factors;

[0064] S6. Train the model according to the loss function to obtain a trained model;

[0065] S7. Put the test data set into the model for prediction to obtain predicted values, and put the test data set into the trained model for testing to obtain test values;

[0066] S8. Normalize the predicted values and test values to obtain the mean absolute percentage error and correlation coefficient of the loss function, and evaluate the model fitting effect and accuracy.

[0067] In one embodiment of the present invention, an object instance model is established. Referring to the attached Figure 2 , taking the electronic packaging housing model as an example, and dividing its regions: three regions, namely, a thermal loading region, a transition region, and a region far from the heat source. The main function of this region segmentation is: the thermal influence region of laser sealing welding is small, which improves the simulation calculation accuracy and efficiency. A hybrid heat source is formed by loading a double ellipsoidal heat source and a three-dimensional conical heat source to simulate the laser sealing welding heat source effect, and the environmental convection parameter is set to 20 °C to obtain a temperature field data set. The data set is obtained through simulation and experiment, and the data set is divided into a training data set and a test data set according to 3:1. Then, unbalanced sample processing and model training are performed. As shown in the attached Figure 3 , the flowchart of the unbalanced sample data augmentation method based on the adaptive loss function of the composite path metal sealing welding process.

[0068] First, design the input process parameter features, obtain experimental data, and preprocess the data. Referring to the attached Figure 4 , the preprocessing of the data in step S2 of the present invention includes:

[0069] S21. Remove the blank line data in the experiment to prevent abnormal blank data from occurring.

[0070] S22. Use the principle of mathematical orthogonal vector combination to combine the eigenvectors. For example, perform vector synthesis on the X-axis and Y-axis velocities to reduce the feature dimension and calculation amount.

[0071] S23. Convert the categorical feature vector into the form of a one-hot encoded vector. In the one-hot encoded vector, index 0 converts the categorical feature vector into an integer vector form.

[0072] In the embodiment of the present invention, the feature dimensions include laser power, welding speed (X-axis), welding speed (Y-axis), inert gas flow rate, laser angle, material thickness, and gap, and simulation results such as stress and strain are obtained. Perform dimensionality reduction on the process parameter feature dimensions. Taking the importance analysis of process parameter features by the random forest importance algorithm as an example, the features with a greater impact on the welding result are obtained. Taking the electronic packaging housing model as an example, as shown in the appendix Figure 5 As shown, the inert gas flow velocity has the greatest impact on the result, and the importance ratio can reach 21%. Perform cardinality superposition on the features, and according to the 95% principle, find the features with the largest range affecting the result, simplify the model structure, and save computing performance. Taking the case shown in the appendix Figure 2 As shown, the most important features are mainly affected by six factors: the flow rate of inert gas, material thickness, focus position, angular position, laser power, and welding speed.

[0073] The method in step S3 of the present invention for dividing the preprocessed data set corresponding to the important features into dense data and sparse data includes: decomposing the preprocessed data set corresponding to the important features along the X-axis and Y-axis. The data set with values on both the X-axis and Y-axis is dense data, and the data set with only X-axis or Y-axis values is sparse data. Taking the velocity in the features of this embodiment as an example, decompose the velocity vector along the X-axis and Y-axis. The data set with values on both the X-axis and Y-axis is dense data, and the data set with only X-axis or Y-axis values is sparse data. The data set with velocities on both the X-axis and Y-axis being dense data means that, as shown in the appendix Figure 2 The established model is a rectangular model. Align the length of the rectangle with the Y-axis and the width with the X-axis. There will be X-axis velocity and Y-axis velocity at the bend.

[0074] The method in step S4 of the present invention for grouping the dense data and sparse data in proportion includes: the sample group with all dense data is the large dense group; the sample group with the ratio of dense data to sparse data being 5:1 is the small sparse group; the sample group with the ratio of dense data to sparse data being 1:10 is the large sparse group. The present invention processes the imbalanced samples by proportional grouping to reduce the impact of imbalanced samples on the model.

[0075] The total number of samples in each sample group is 500, and the formula of the loss function factor η for each group is different. The loss factor η corresponding to the dense data in the large dense group, the small sparse group, and the large sparse group m , and the loss factor η corresponding to the sparse data n are different. Specifically:

[0076] In the large dense group,

[0077] In the small sparse group,

[0078] In the large sparse group, Φ is an adaptive parameter, and its range is 0 to 1.

[0079] Among them, the large dense group and the small sparse group will utilize the values of the number m of dense data and the number n of sparse data to affect the loss factor η corresponding to the dense data m and the loss factor η corresponding to the sparse data n , so as to increase the weight of the curve data. The large sparse group can adjust the weight of the curve data by using the adaptive parameter Φ. The range of Φ is between 0 and 1. The user can adjust the adaptive parameter Φ, thereby affecting the loss function mLoss and improving the model effect.

[0080] The formula of the loss function factor of the present invention is expressed as:

[0081]

[0082] In formula (1), η is the loss function factor, m represents the number of dense data, n represents the number of sparse data, η m represents the loss factor corresponding to the dense data, and η n represents the loss factor corresponding to the sparse data.

[0083] The formula of the loss function factor η corresponding to each data group refers to Table 1.

[0084] Table 1 The formula of the loss function designed by the present invention is expressed as:

[0085]

[0086] In formula (2), mLoss represents the loss function, η is the loss function factor, M represents the total number of each group of data sets, f′ represents the predicted stress result of the test data after passing through the trained model, and f represents the stress result obtained by the test data through the real experiment.

[0087] In one embodiment of the present invention, Z-score normalization is used to normalize the predicted values and the test values, thereby obtaining evaluation metrics: mean absolute percentage error and correlation coefficient.

[0088] Z-score normalization is a method of converting data into a standard normal distribution with a mean of 0 and a standard deviation of 1. This method can eliminate the influence of dimensions, improve the performance of the model, and has a certain robustness to outliers. The formula for Z-score normalization is expressed as:

[0089]

[0090] In formula (3), X is a certain value in the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and z is the data after Z-score normalization.

[0091] In the present invention, the formula for the mean absolute percentage error of the loss function is expressed as:

[0092]

[0093] The formula for the correlation coefficient is expressed as:

[0094]

[0095] In formulas (4) and (5), MAPE represents the mean absolute percentage error of the loss function, R 2 represents the correlation coefficient, n represents the total number of samples, i represents the i-th sample, y i represents the true value of the i-th sample, represents the predicted value of the i-th sample, represents the average value of the true sample values.

[0096] MAPE is used to measure the model performance index. The smaller its value, the smaller the deviation between the predicted value and the test value, the better the model performance, and the higher the accuracy. R 2 represents the degree of explainability of the independent variable to the dependent variable, and measures the degree of fit of the data from the perspective of volatility. The closer its value is to 1, the better the model fitting degree.

[0097] Using the above normalized data, the mean absolute percentage error in the evaluation criteria is obtained: MAPE = 3.8%, and the correlation coefficient: R 2 = 0.97.

[0098] To further verify the effectiveness of the present invention, a comparative experiment was conducted. The test data set was input into the trained model for testing. Then, an MLP model (Multi-Layer Perceptron model) that uses the MSE loss function (Mean Squared Error) for prediction was trained in the same way, and gradient descent was used for optimization and iteration to obtain a training comparison graph of the predicted values and test values of the two loss functions, as shown in the appendix Figure 6 as follows.

[0099] In the appendix Figure 6 , the red line represents the prediction result obtained by the MLP model using the MSE loss function for prediction, the blue line represents the prediction result obtained by the MLP model of the method of the present invention for prediction, and the orange line represents the true value obtained from the actual experiment. As can be seen from the appendix Figure 6 , the blue line is more fitted to and closer to the true value, and the effect of the present invention is better. Therefore, the effect after processing the imbalanced samples is better.

[0100] Flow chart for the training of the Multi-Layer Perceptron (MLP) model, as shown in the appendix Figure 7 as follows.

[0101] The formula of the MSE loss function is expressed as:

[0102]

[0103] In formula (6), n represents the number of samples, y i represents the true value of the i-th sample, represents the predicted value of the i-th sample.

[0104] The present invention was compared with an MLP model that does not use an adaptive loss function to process sample data, and the data is shown in Table 2.

[0105] Table 2

[0106]

[0107] By comparing and analyzing MAPE and R 2 in Table 2, the MAPE value decreased by 2.4%, and the value of R 2 increased by 0.1.

[0108] In summary, the method for enhancing unbalanced sample data based on an adaptive loss function of the present invention uses the method of adaptive weight allocation for the unbalanced sample set to enhance the prediction accuracy of the model and retains the integrity of the data. The unbalanced data set is processed using the loss function factor η formula, and then the model is trained using the loss function mLoss to find the optimal parameters with the best evaluation indicators, optimizing the parameter control accuracy. Moreover, the method for enhancing unbalanced sample data of the present invention has strong adaptability and can be applied to various unbalanced data scenarios. It is applicable to the processing of unbalanced data models such as financial fraud detection and network intrusion recognition, and can perform model learning for different scenarios and different parameter types.

[0109] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. An imbalanced sample data augmentation method based on an adaptive loss function, characterized in that Including the steps: S1. Establish a model, obtain a data set, and divide the data set into a training data set and a test data set; S2. Preprocess the data in the training data set to obtain a preprocessed data set; S3. Conduct an importance analysis on the features of the model to obtain important features after reducing the feature dimension, and divide the preprocessed data set corresponding to the important features into dense data and sparse data; S4. Group the dense data and the sparse data in proportion, including a large data group, a small dense group, and a large sparse group; S5. Set different loss function factors according to the grouping, and design a loss function using the loss function factors; S6. Train the model according to the loss function to obtain a trained model; S7. Put the test data set into the model for prediction to obtain predicted values, and put the test data set into the trained model for testing to obtain test values; S8. Perform normalization processing on the predicted values and the test values to obtain the mean absolute percentage error of the loss function and the correlation coefficient, and evaluate the model fitting effect and accuracy.

2. The unbalanced sample data augmentation method based on an adaptive loss function according to claim 1, wherein The step S2 includes: S21. Remove blank line data; S22. Merge the feature vectors; S23. Convert the categorical feature vector into a one-hot encoded vector form.

3. An imbalanced sample data augmentation method based on an adaptive loss function according to claim 1, characterized in that The step S3 includes: Decompose the preprocessed data set corresponding to the important features along the X-axis and the Y-axis. The data set with values on both the X-axis and the Y-axis is dense data, and the data set with only X-axis or Y-axis values is sparse data.

4. A method for enhancing unbalanced sample data based on an adaptive loss function according to claim 1, characterized in that, The step S4 includes: The sample group consisting entirely of the dense data is the large dense group; the sample group with a ratio of the dense data to the sparse data of 5:1 is the small sparse group; the sample group with a ratio of the dense data to the sparse data of 1:10 is the large sparse group.

5. An imbalanced sample data augmentation method based on an adaptive loss function according to claim 1, characterized in that, In the step S5: The formula for the loss function factor is expressed as: In formula (1), η is the loss function factor, m represents the number of dense data, n represents the number of sparse data, η m represents the loss factor corresponding to the dense data, and η n represents the loss factor corresponding to the sparse data.

6. The unbalanced sample data augmentation method based on an adaptive loss function according to claim 5, wherein The loss factors η corresponding to the dense data of the large dense group, the small sparse group, and the large sparse group m , and the loss factor η corresponding to the sparse data n are different: Among the large number of dense groups, Among the small number of sparse groups, Among the large number of sparse groups, Φ is an adaptive parameter with a range of 0 to 1.

7. An imbalanced sample data augmentation method based on an adaptive loss function according to claim 5, characterized in that, In the step S5: The formula for the loss function is expressed as: In formula (2), mLoss represents the loss function, η is the loss function factor, M represents the total number of each group of data sets, f′ represents the predicted stress result of the test data after passing through the trained model, and f represents the stress result obtained by the test data through a real experiment.

8. An imbalanced sample data augmentation method based on an adaptive loss function according to claim 1, characterized in that, The step S8 includes: Perform normalization processing on the predicted values and the test values using Z-score normalization. The formula for Z-score normalization is expressed as: In formula (3), X is a certain value in the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and z is the data after Z-score normalization processing.

9. The unbalanced sample data augmentation method based on an adaptive loss function according to claim 1, characterized in that, In the step S8: The formula for the mean absolute percentage error of the loss function is expressed as: In formula (4), MAPE represents the mean absolute percentage error of the loss function, n represents the total number of samples, i represents the i-th sample, and y i represents the true value of the i-th sample, represents the predicted value of the i-th sample.

10. An imbalanced sample data augmentation method based on an adaptive loss function according to claim 1, characterized in that, In the step S8: The formula for the correlation coefficient is expressed as: In formula (5), R 2 represents the correlation coefficient, n represents the total number of samples, i represents the i-th sample, and y i represents the true value of the i-th sample, represents the predicted value of the i-th sample, represents the average value of the true sample values.