Method for Selecting Experimental Sample Points Based on Gaussian Process Regression
Through the experimental sample point selection method based on Gaussian process regression, the kernel function and error response function are determined, the training data set is generated and iteratively optimized, and the optimal sampling point collection is selected, which solves the problem of low efficiency in the selection of experimental sample point, and realizes the optimal configuration and efficiency improvement of experimental resources.
Patent Information
- Application Number
- CN202011290863.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-11-17
AI Technical Summary
In the prior art, the selection efficiency of test sample points is low, resulting in wasted experimental resources.
The experimental sample point selection method based on Gaussian process regression is used to determine the kernel function and error response function, and the training data set is generated, and iterative optimization is performed to select the optimal sample point set.
On the premise of ensuring the accuracy of the response variable test, minimize the test workload and improve the test efficiency.
Smart Images

Figure CN114510812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of systems engineering, and more particularly, to a method for selecting experimental sample points based on Gaussian process regression. Background Art
[0002] Currently, in the field of actual engineering, the verification and validation of the functional performance of products or systems is a very important part of systems engineering, which involves a large amount of functional and performance test work. For example: wind tunnel tests related to the aerodynamic layout of aircraft, thrust tests of rocket engines, electrical performance tests of antenna systems such as radars, etc. Generally speaking, the scale of these tests is huge, and the test resources such as time and funds consumed are also relatively large. Therefore, the selection of experimental sample points is particularly important. Common methods for selecting experimental sample points are as follows:
[0003] 1. Exhaustive sampling method. That is, sample points are obtained by traversing the experimental data sample space to verify the response variables one by one. This method of selecting sample points results in a large amount of redundant work and extremely low test efficiency.
[0004] 2. Linear interpolation sampling. Sample points are interpolated at a certain step size in the test data sample space to test the response variables. This method can only be applied to interpolation and not extrapolation, and the interpolation accuracy is relatively low in the local space where the data changes violently.
[0005] 3. Traditional experimental designs, such as D-optimal design, super Latin square design, etc. These experimental design methods can sample uniformly in the multi-dimensional sample space and reduce the number of sampling points to a certain extent. However, its defect is that it does not make full use of historical prior data, resulting in the distribution of sample points not being based on the data behavior, and there will also be situations of local redundancy or insufficiency.
[0006] The above methods do not optimize the allocation of test resources, consume unnecessary time and economic costs, and have obvious limitations.
[0007] As can be seen from the above, there is a problem of low efficiency in selecting experimental sample points in the prior art. Summary of the Invention
[0008] The main object of the present invention is to provide a method for selecting experimental sample points based on Gaussian process regression to solve the problem of low efficiency in selecting experimental sample points in the prior art.
[0009] To achieve the above object, the present invention provides a method for selecting test sample points based on Gaussian process regression, including: during the process of analyzing the behavior of simulation data, determining the kernel function of the Gaussian process regression model; determining the error response function; obtaining the historical data of performance tests in actual engineering problems to generate a training dataset; importing the training dataset, and using the Gaussian process regression model to determine the initial hyperparameters of the kernel function; constructing an initial experiment; performing iterative optimization of sampling using the Gaussian process regression model to select the optimal sampling point set; after the iterative optimization ends, outputting the selected optimal sampling point set.
[0010] Further, the calculation formula of the kernel function is:
[0011]
[0012] where \(x_1,x_2,\cdots,x\) n are the coordinates in the N-dimensional sample space; \(h_1,h_2,x\) 10 ,x 20 \(\cdots,x\) n0 , \(a_1,a_2,\cdots,a\) n are the hyperparameters to be optimized; \(h_1,h_2\) are used to adjust the magnitude of the correlation length; \(a_1,a_2,\cdots,a\) n are used to adjust the change trend of the correlation length; \(x\) 10 ,x 20 \(\cdots,x\) n0 are the means of the extreme values of the correlation length.
[0013] Further, the calculation formula of the error response function is:
[0014]
[0015] where \(var(x)\) is the predicted variance at point \(x\) in the N-dimensional sample space.
[0016] Further, assuming that after giving \(n\) sampling points, the error response function has \(p\) maximum values \(\varepsilon(x\) i ), \(i = 1,\cdots,p\), and assuming that the preset pointing accuracy error is \(\varepsilon_0\), the coordinates \(x\) k of the sampling points to be selected satisfy:
[0017] \(\varepsilon(x\) k ) \(\geq \varepsilon_0 + f(\varepsilon\) max - \(\varepsilon_0)\);
[0018] where \(\varepsilon\) max is the maximum maximum value of the error response function; \(f\) is the truncation coefficient, \(0 < f < 1\).
[0019] Further, the maximum maximum value \(\varepsilon\) maxThe calculation formula is as follows:
[0020] ε max = max[ε(x i )];
[0021] where i = 1,..., p.
[0022] Furthermore, the initialization experiment construction includes: initializing the hyperparameters of the kernel function as the initial hyperparameters; presetting an N max value. When the number of sampling points exceeds the N max value, the kernel function hyperparameter optimization program is started and the initial hyperparameters are optimized; setting a preset pointing precision error ε0; determining the initial sampling point set.
[0023] Furthermore, the iterative optimization includes: judging whether the number of the initial sampling point set in the initialization experiment construction is greater than N max . If it is greater, the kernel function hyperparameter optimization program is started; running the Gaussian process regression model to determine the preset specified precision error ε0 on the training data set; determining the maximum value ε(x i ) according to the error response function; judging whether the maximum value ε(x i ) satisfies max[ε(x i )] ≤ ε0. If it satisfies, the iterative optimization ends and the selected optimal sampling point set is output.
[0024] Furthermore, the iterative optimization also includes: if the number of the initial sampling point set is less than N max , then directly run the Gaussian process regression model to determine the preset specified precision error ε0 on the training data set.
[0025] Furthermore, the iterative optimization also includes: if the maximum value ε(x i ) does not satisfy max[ε(x i )] ≤ ε0, select the sampling point set {x k} that satisfies ε(x max ) ≥ ε0 + f(ε i - ε0) of the maximum value, add the sampling point set {x i} to the initial sampling point set, and repeat the iterative optimization.
[0026] Furthermore, the historical data includes historical measured data and historical simulation data.
[0027] Applying the technical solution of the present invention, during the process of analyzing the behavior of simulation data, the kernel function of the Gaussian process regression model is determined, and the error response function is determined. The training data set is generated based on the historical data of performance testing in actual engineering problems, and the Gaussian process regression model is used to determine the initial hyperparameters of the kernel function, so as to construct the initial experiment. Then, the Gaussian process regression model is used for iterative optimization of sampling, and the optimal sampling point set is selected. In this way, on the premise of ensuring the test accuracy of the response variable, by selecting the optimal sampling point set, the cost of completing the test verification work of the response variable can be minimized, achieving the effect of effectively reducing the test workload and improving the test efficiency, and solving the problem of low efficiency in selecting test sample points in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0029] Figure 1 The flowchart of the method for selecting test sample points based on Gaussian process regression of the present invention is shown;
[0030] Figure 2 The flowchart of the method for selecting test sample points based on Gaussian process regression in a specific embodiment of the present invention is shown;
[0031] Figure 3 The kernel function diagram of the azimuth pointing accuracy of the radome in a specific embodiment of the present invention is shown;
[0032] Figure 4 The kernel function diagram of the elevation pointing accuracy of the radome in a specific embodiment of the present invention is shown;
[0033] Figure 5 The sampling point set of the azimuth pointing accuracy of the radome in a specific embodiment of the present invention is shown;
[0034] Figure 6 The sampling point set of the elevation pointing accuracy of the radome in a specific embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0036] It should be pointed out that unless otherwise specified, all technical and scientific terms used in this application have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0037] In the present invention, unless otherwise specified, the orientation terms such as "upper", "lower", "top", and "bottom" generally refer to the directions shown in the drawings or to the vertical, perpendicular, or gravitational directions of the components themselves. Similarly, for the convenience of understanding and description, "inner" and "outer" refer to the inner and outer of the contours of the respective components themselves. However, the above orientation terms are not used to limit the present invention.
[0038] In order to solve the problem of low efficiency in selecting test sample points in the prior art, the present invention provides a method for selecting test sample points based on Gaussian process regression. To better elaborate on the method for selecting test sample points based on Gaussian process regression of the present invention, the method for selecting test sample points based on Gaussian process regression of the present invention will be specifically described below in application to the selection of test sample points for the pointing accuracy of radomes.
[0039] As Figures 1 to 2 shown, the method for selecting test sample points based on Gaussian process regression includes: during the process of analyzing the behavior of simulation data, determining the kernel function of the Gaussian process regression model; determining the error response function; obtaining the historical data of performance tests in actual engineering problems to generate a training data set; importing the training data set and using the Gaussian process regression model to determine the initial hyperparameters of the kernel function; conducting an initial experiment construction; using the Gaussian process regression model for iterative optimization of sampling to select the optimal sampling point set; and after the iterative optimization ends, outputting the selected optimal sampling point set.
[0040] By determining the kernel function of the Gaussian process regression model and the error response function during the process of analyzing the behavior of simulation data, generating a training data set based on the historical data of performance tests in actual engineering problems, and using the Gaussian process regression model to determine the initial hyperparameters of the kernel function, thus conducting an initial experiment construction, and then using the Gaussian process regression model for iterative optimization of sampling to select the optimal sampling point set, it is possible to minimize the cost of completing the response variable test verification work by selecting the optimal sampling point set while ensuring the test accuracy of the response variable, thereby effectively reducing the experimental workload and improving the experimental efficiency.
[0041] The kernel function of the test sample point selection method based on Gaussian process regression of the present invention is a universal elliptical contour kernel function that can adaptively adjust model parameters according to actual conditions. The correlation length defined by this kernel function is inverted bell-shaped in the N-dimensional variable space. It is a strict Gaussian function in any longitudinal section, and its transverse section is an ellipse. The kernel function of the test sample point selection method based on Gaussian process regression of the present invention can be easily extended to general situations: by increasing the dimension of the coordinates in the Gaussian terms, we can extend the function model to any high-dimensional situation; at the same time, by increasing the number of Gaussian terms and performing a certain coordinate rotation on each Gaussian term, the function model can also describe the dramatic change area of the data as a system with any discrete distribution.
[0042] Specifically, the calculation formula of the kernel function is:
[0043]
[0044] where x1,x2,...,x n are the coordinates in the N-dimensional sample space; h1,h2,x 10 ,x 20 ...,x n0 ,a1,a2,...,a n are the hyperparameters to be optimized; h1, h2 are the sizes of adjusting the association length; a1, a2, ..., a n To adjust the changing trend of the correlation length; x 10 ,x 20 ...,x n0 is the mean of the extreme values of the adjusted correlation length.
[0045] The kernel function of the test sample point selection method based on Gaussian process regression of the present invention can effectively reflect the mapping relationship between the test data and the Gaussian process regression model, and effectively ensure the model training effect and prediction accuracy.
[0046] In this embodiment, the performance analysis of the simulation data of the radome pointing accuracy is performed to determine the kernel function of the Gaussian process regression model. The test of the radome pointing accuracy has two parameters, "azimuth angle" and "pitch angle", and two test results, "azimuth pointing accuracy" and "pitch pointing accuracy", which are unrelated. Therefore, the calculation formula of the kernel function corresponding to the radome pointing accuracy is:
[0047]
[0048] Here x1, x2 are the coordinates in the parameter space, which are the azimuth and elevation angles respectively; h1, h2, x 10 ,x 20 ,a,b are the hyperparameters to be optimized.
[0049] In this embodiment, the calculation formula of the error response function is as follows:
[0050]
[0051] where var(x) is the predicted variance at point x in the N-dimensional sample space.
[0052] In this embodiment, it is assumed that after n sampling points are given, the error response function has p maximum values ε(x i ), i = 1,..., p, and it is assumed that the preset pointing accuracy error is ε0, and the coordinates x of the sampling points to be selected k satisfy:
[0053] ε(x k ) ≥ ε0 + f(ε max - ε0);
[0054] where ε max is the maximum maximum value of the error response function; f is the truncation coefficient, 0 < f < 1.
[0055] Specifically, the calculation formula for the maximum maximum value ε max of the error response function is:
[0056] ε max = max[ε(x i )];
[0057] where i = 1,..., p.
[0058] In this embodiment, historical data of the radome pointing accuracy performance test is obtained to generate a training data set. Among them, the historical data includes historical measured data and historical simulation data.
[0059] In this embodiment, the training data set of the radome pointing accuracy includes: at each scanning frequency, omnidirectional simulation data of a total of 21×121 data points with the elevation angle in the range of -38° to 82° and the azimuth angle in the range of -60° to 60°. Among them, the elevation angle takes 21 non-uniform angles, and the azimuth angle takes 121 uniform angles every 1°.
[0060] As Figures 3 to 4 shown, using the above training data set of the radome pointing accuracy, the kernel function and hyperparameters of the Gaussian process regression model are trained to obtain the kernel function images describing the azimuth pointing accuracy and the elevation pointing accuracy.
[0061] As Figure 2 shown, the initialization experiment construction includes: initializing the hyperparameters of the kernel function as the initial hyperparameters; presetting N maxValue, when the number of sampling points exceeds N max Value, the kernel function hyperparameter optimization program starts and optimizes the initial hyperparameters; set the preset pointing precision error ε0; determine the initial sampling point set. Through the initialization experiment construction, it prepares for the subsequent iterative optimization of sampling.
[0062] In this embodiment, four points on the azimuth and elevation two-dimensional space boundaries in the two-dimensional angle sample space {(xAZmin, xELmin), (xAZmax, xELmin), (xAZmin, xELmax), (xAZmax, xELmax)} are selected as the initial sampling point set.
[0063] As Figure 2 shown, after the iterative optimization starts, when the number of selected sampling points N is less than a certain value N max , we use the simulation data as the training set to fit the hyperparameters of the prediction model, and thereby calculate the prediction error of the system to obtain the coordinate distribution of the sampling points that need to be updated after iteration. When the number of sampling points N is greater than N max , if the system error has not converged within the precision range, we will use the sampling data as the training set to fit the hyperparameters of the prediction model, and calculate the prediction mean and error distribution with the newly fitted hyperparameters, so as to determine the coordinates of the next sampling point until all error values converge within the precision range. In this embodiment, the value of N max can be estimated based on past experience or simulation experiments. Usually, when the number of sampling points N is greater than N max , we have collected enough sampling data. Therefore, at this time, the hyperparameters fitted with the sampling data can already accurately describe the true data structure of the test system. In subsequent iterations, as the sampling points increase, this set of hyperparameters will be further optimized, and the prediction mean will gradually and stably approach the true data structure of the test system. Specifically, the iterative optimization includes: judging whether the number of the initial sampling point set in the initialization experiment construction is greater than N max , if it is greater, the said kernel function hyperparameter optimization program starts; run the Gaussian process regression model, determine the preset specified precision error ε0 on the training data set; determine the maximum value ε(x i ) according to the error response function; judge whether the maximum value ε(x i ) satisfies max[ε(x i )] ≤ ε0, if it satisfies, the iterative optimization ends and outputs the selected optimal sampling point set.
[0064] As Figure 2 shown, the iterative optimization also includes: if the number of the initial sampling point set is less than N max , then directly run the Gaussian process regression model and determine the preset specified precision error ε0 on the training data set.
[0065] As Figure 2 shown, the iterative optimization further includes: if the maximum value ε(x i ) does not satisfy max[ε(x i )] ≤ ε0, select the set of sampling points {x k} that satisfies ε(x max ) ≥ ε0 + f(ε i - ε0) and has the maximum value, and add the set of sampling points {x i} to the initial set of sampling points, and repeat the iterative optimization.
[0066] As Figures 5 to 6 shown, after iterative optimization, the number of sampling points for the azimuth pointing accuracy of the radome in this embodiment is reduced from 2541 to 188, and the number of sampling points for the elevation pointing accuracy of the radome is reduced from 2541 to 177. By using the method for selecting test sample points based on Gaussian process regression of the present invention to select test sample points for the radome pointing accuracy, an optimal set of sampling points is selected on the premise of ensuring the test accuracy of the response variable, greatly reducing the cost of completing the test verification work, thereby effectively reducing the test workload and improving the test efficiency.
[0067] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects: by determining the kernel function of the Gaussian process regression model and the error response function during the behavior analysis of the simulation data, generating a training data set based on the historical data of performance testing in actual engineering problems, and using the Gaussian process regression model to determine the initial hyperparameters of the kernel function, thereby constructing the initial test, and then using the Gaussian process regression model for iterative optimization of sampling to select the optimal set of sampling points. In this way, on the premise of ensuring the test accuracy of the response variable, by selecting the optimal set of sampling points, the cost of completing the test verification work of the response variable can be minimized to the greatest extent, thereby effectively reducing the test workload and improving the test efficiency.
[0068] Obviously, the above-described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] It should be noted that the terms used here are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used here, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0070] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here.
[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for selecting experimental sample points based on Gaussian process regression, characterized in that, Including: During the process of analyzing the behavior of simulation data, determine the kernel function of the Gaussian process regression model; Determine the error response function; Obtain the historical data of performance tests in actual engineering problems to generate a training dataset; Import the training dataset and use the Gaussian process regression model to determine the initial hyperparameters of the kernel function; Conduct initial experiment construction; Use the Gaussian process regression model for iterative optimization of sampling to select the optimal sampling point set; After the iterative optimization ends, output the selected optimal sampling point set; Among them, during the process of analyzing the behavior of simulation data, determining the kernel function of the Gaussian process regression model is to determine the kernel function of the Gaussian process regression model during the process of analyzing the behavior of simulation data of the radome pointing accuracy; Obtaining the historical data of performance tests in actual engineering problems is to obtain the historical data of performance tests of the radome pointing accuracy; The radome pointing accuracy includes azimuth pointing accuracy and elevation pointing accuracy; The calculation formula of the kernel function is: where x1, x2, ..., x n are coordinates in an N-dimensional sample space; h1, h2, x 10 , x 20 ..., x n0 , a1, a2, ..., a n are hyperparameters to be optimized; h1, h2 are used to adjust the magnitude of the correlation length; a1, a2, ..., a n are the changing trends of the adjusted correlation length. x 10 , x 20 ..., x n0 is the mean value of the extreme values of the adjustment correlation length.
2. The method for selecting experimental sample points based on Gaussian process regression according to claim 1, wherein The calculation formula of the error response function is: Among them, var(x) is the prediction variance at point x in the N-dimensional sample space.
3. The method for selecting test sample points based on Gaussian process regression according to claim 2, wherein Assume that after giving n sampling points, there are p maximum values ε(x i ), i = 1, ..., p, in the error response function, and assume that the preset pointing accuracy error is ε0, and the coordinates x k of the sampling points to be selected satisfy: ε(x k )≥ε0 + f(ε max - ε0); where ε max is the maximum maximum value of the error response function; f is a truncation coefficient, 0 < f < 1.
4. The method for selecting test sample points based on Gaussian process regression according to claim 3, characterized in that The maximum maximum value ε of the error response function max is calculated by the formula: ε max = max[ε(x i )]; Among them, i = 1,..., p.
5. The method for selecting test sample points based on Gaussian process regression according to claim 1, wherein The initial experiment construction includes: The initial hyperparameters of the kernel function are used as the initial hyperparameters; Preset N max value, when the number of sampling points exceeds the N max value, the kernel function hyperparameter optimization program starts and optimizes the initial hyperparameters; Set the preset pointing accuracy error ε0; Determine the initial sampling point set; The said N max is the maximum number of sampling points.
6. The method for selecting test sample points based on Gaussian process regression according to claim 1, characterized in that The iterative optimization includes: Determine whether the number of initial sampling point sets in the initialization test construction is greater than N max , if it is greater, the kernel function hyperparameter optimization program is started; Run the Gaussian process regression model to determine the preset specified accuracy error ε0 on the training dataset; Determine the maximum value ε(x i ) according to the error response function; Determine whether the maximum value ε(x i ) satisfies max[ε(x i )] ≤ ε0. If it is satisfied, the iterative optimization ends and the selected set of optimal sampling points is output; The said N max is the maximum number of sampling points.
7. The method for selecting test sample points based on Gaussian process regression according to claim 6, characterized in that The iterative optimization also includes: If the number of the initial sampling point sets is less than N max , directly run the Gaussian process regression model to determine the preset specified precision error ε0 on the training data set.
8. The method for selecting experimental sample points based on Gaussian process regression according to claim 6, characterized in that, The iterative optimization also includes: If the maximum value ε(x i ) does not satisfy max[ε(x i )] ≤ ε0, select the set of sampling points {x k} that satisfy ε(x max ) ≥ ε0 + f(ε i - ε0) of the maximum value, add the set of sampling points {x i} to the initial set of sampling points, and repeat the iterative optimization.
9. The method for selecting experimental sample points based on Gaussian process regression according to any one of claims 1 to 8, characterized in that, The historical data includes historical measured data and historical simulation data.
Citation Information
Patent Citations
Improved inertial positioning method for real estate measurement
CN108171316A
Multi-working-condition power system performance prediction method and system based on Gaussian process regression
CN111460379A