Data sampling method
By randomly selecting parameter combinations, establishing surrogate models and scoring functions, the sampling points are gradually guided to the vicinity of the performance threshold, solving the problem of insufficient sampling point selection in existing technologies, and achieving efficient observation of system performance boundaries and improvement of model accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to efficiently select sampling points to cover behaviors near performance thresholds when evaluating system performance metrics, resulting in insufficient model accuracy.
Simulations were conducted using randomly selected parameter combinations. A surrogate model and a scoring function were established to gradually guide the sampling points to near the performance threshold. The mean and standard deviation were predicted using a Gaussian process model, and the sampling point selection was optimized using the scoring function.
By conducting a limited number of simulations, the system's performance boundary characteristics can be effectively observed and understood, improving the model's accuracy and efficiency while reducing the number of sample points.
Smart Images

Figure CN121809131A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer-aided analysis, and particularly to a data sampling method. BACKGROUND
[0002] In engineering design and scientific research, it is often necessary to evaluate the performance index of a system under different parameter combinations through computer simulation. When the calculation cost of a single simulation is high, it is unrealistic to cover all parameter combinations. Therefore, researchers usually select a limited number of parameter combinations, i.e., sampling points, for simulation, and then establish an approximate model based on these sampling points and their simulation results to analyze the characteristics of the entire system. How to select sampling points is a problem that needs to be solved in this kind of method. SUMMARY
[0003] The technical problem solved by the present application is to overcome the shortcomings of the prior art and provide a data sampling method to help analysts focus on the behavior of system performance near a certain threshold value.
[0004] The technical solution of the present application is: in the first aspect, a data sampling method is provided, which is used to select parameter combinations for simulation when a limited number of simulations are performed to evaluate the performance index of a system under different parameter combinations, comprising the following steps:
[0005] S1, randomly select N parameter combinations, simulate the system to be evaluated under each parameter combination, and obtain the performance index of the system to be evaluated under each parameter combination; a total of N simulations are performed to obtain N performance indexes; the N parameter combinations and N performance indexes are denoted as a sample set; N is a positive integer, and N < total simulation times;
[0006] S2, use the sample set to establish a surrogate model; the surrogate model is used to provide a predicted mean μ(z) and a predicted standard deviation σ(z) for any parameter combination z;
[0007] S3, establish a scoring function α(z); the scoring function is used to evaluate the value score of all parameter combinations not selected by S1, when the predicted mean μ(z) of the selected parameter combination is closer to the performance index threshold and the predicted standard deviation σ(z) is larger, the value score is higher, and when the selected parameter combination is closer to the selected parameter combination, the value score is lower;
[0008] S4, use the scoring function to find the parameter combination z with the highest value score * ; use the parameter combination z * to perform a new simulation on the system to be evaluated to obtain the performance index of this simulation;
[0009] S5, add the parameter combination and performance index in S4 to the sample set and update the surrogate model;
[0010] S6, repeating steps S4 and S5 until a preset number of simulations is reached.
[0011] Further, the scoring function a(z) is specifically:
[0012] a(z)=a BE (z)-ηa P (z)
[0013]
[0014] wherein η is an adjustment coefficient; exp() represents an exponential function; y th is a performance index threshold of the system to be evaluated; κ is a hyperparameter for controlling the search intensity of the performance index threshold; min represents taking the minimum value; z k represents the selected parameter combination; and τ is a hyperparameter for controlling the exclusion radius.
[0015] Further, the surrogate model is a Gaussian process model.
[0016] Further, N≤5.
[0017] In a second aspect, the present application provides a data sampling system for selecting a parameter combination for simulation when evaluating the performance index of a system under different parameter combinations in a limited number of simulations, comprising:
[0018] a preliminary sampling module for randomly selecting N parameter combinations to form a sample set; the sample set comprises N parameter combinations and N performance indexes; the N performance indexes are the performance indexes of the system to be evaluated under each parameter combination after simulation of the system to be evaluated under each parameter combination; N is a positive integer, and N < total number of simulations;
[0019] a surrogate module comprising a surrogate model; the surrogate model is established by the sample set and is used to provide a predicted mean μ(z) and a predicted standard deviation σ(z) for any parameter combination z;
[0020] a scoring module comprising a scoring function a(z); the scoring function is used to evaluate the value score of all parameter combinations not selected by the preliminary sampling module; when the predicted mean μ(z) of the parameter combination to be selected is closer to the performance index threshold and the predicted standard deviation σ(z) is larger, the value score is higher; when the parameter combination to be selected is closer to the selected parameter combination, the value score is lower;
[0021] a value sampling module for finding the parameter combination z * with the highest value score by using the scoring function.
[0022] an updating module, configured to combine the parameters to obtain a parameter combination z * and simulate the to-be-evaluated system under the parameter combination z * After the simulation of the to-be-evaluated system under the parameter combination z
[0023] a control module, configured to cyclically call the value sampling module and the updating module until a preset simulation number is reached.
[0024] Further, the score function α(z) is specifically as follows:
[0025] α(z)=α BE (z)-ηα P (z)
[0026]
[0027] wherein η is an adjustment coefficient; exp() represents an exponential function; y th is a performance index threshold of the to-be-evaluated system; κ is a hyperparameter for controlling the search intensity of the performance index threshold; min represents taking a minimum value; z k represents a selected parameter combination; and τ is a hyperparameter for controlling a radius of exclusion.
[0028] Further, the agent model adopts a Gaussian process model.
[0029] Further, N≤5.
[0030] The present application also relates to a computer program product, characterized in that the computer program product, when executed by a processor, realizes the steps of the method as described above.
[0031] The present application also relates to a computer-readable storage medium, characterized in that a computer program in the medium, when executed by a processor, realizes the steps of the method as described above.
[0032] Compared with the prior art, the present application has the following advantages:
[0033] The method provided by the present application can help an analyst to observe and understand the performance boundary characteristics of a system to a certain extent by guiding the sampling points to the vicinity of a performance threshold in a biased manner. Compared with traditional grid sampling or random sampling, the method can achieve an equivalent boundary observation effect with fewer sample points. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 The present application is a method flowchart. DETAILED DESCRIPTION
[0035] In order to better understand the technical solutions of the present application, the specific embodiments of the present application are described below.
[0036] The data sampling method provided by the application, with reference to Figure 1 as shown, specifically includes the following steps:
[0037] (1) First, randomly select N parameter combinations, simulate the system to be evaluated under each parameter combination, and obtain the performance indicators of the system to be evaluated under each parameter combination; a total of N simulations are performed to obtain N performance indicators; the N parameter combinations and the N performance indicators are marked as a sample set;
[0038] (2) A preliminary proxy model (for example, a Gaussian process model) is established using the sample set. The proxy model can provide a predicted mean μ(z) and a predicted standard deviation σ(z) for any parameter combination z;
[0039] (3) Then, a scoring function α(z) is established to evaluate the "value" score of all unselected parameter combinations. The scoring function is calculated as follows:
[0040] α(z) = α BE (z) - ηα P (z)
[0041] where η is an adjustment coefficient. The function consists of two parts:
[0042] The first part α BE (z) is used to evaluate the degree of proximity of the parameter combination to the performance indicator threshold, and the calculation formula is as follows:
[0043]
[0044] Here, y th is the performance indicator threshold that we are interested in. This term indicates that when the predicted mean μ(z) of the proxy model approaches the threshold y th , and the predicted standard deviation σ(z) of the proxy model is large, the score of the parameter combination will be higher.
[0045] The second part α P (z) is a penalty term used to reduce the score near the selected parameter combinations to avoid excessive clustering of sampling points, and the calculation formula is as follows:
[0046]
[0047] Here, z k represents the selected parameter combination, and k is an index value.
[0048] (4) Using the above scoring function, find the parameter combination z * with the highest value score; perform a new simulation of the system to be evaluated using the parameter combination z * , and obtain the performance indicator of this simulation.
[0049] (5) Add the parameter combination and performance index in step (4) to the sample set, and update the agent model;
[0050] (6) Repeat steps (4) and (5) until a preset simulation number is reached.
[0051] In this way, the sampling points (i.e. selected parameter combinations) are gradually guided to the boundary region where the performance index threshold is located.
[0052] Two specific embodiments of the present application are given below.
[0053] Embodiment 1: Analysis of safe working area of electronic component
[0054] This embodiment is to determine whether the internal chip temperature of an electronic component will exceed a safety threshold, for example 85℃, under different working voltages (parameter 1) and environmental temperatures (parameter 2).
[0055] (1) Establish a sample set: randomly select 5 (voltage, temperature) combinations, perform simulation, and obtain 5 chip temperatures.
[0056] (2) Modeling: based on the 5 data points, a Gaussian process model is established.
[0057] (3) Sampling iteration:
[0058] Using the aforementioned scoring function α(z), evaluate in the entire (voltage, temperature) parameter space, where the performance index threshold y th is set to 85.
[0059] Find a point with the highest score, assume it is (voltage = 4.8V, temperature = 60℃), and perform a new simulation at this point.
[0060] Add the new data obtained (the 6th point) to the sample set, and update the Gaussian process model.
[0061] (4) Loop and stop: repeat step 3 until the total simulation number reaches a preset 20 times.
[0062] After the sampling is completed, the 20 simulation points will be mainly distributed around the boundary line where the chip temperature is equal to 85℃, thereby providing data reference for analysts to analyze the safe working area boundary of the component.
[0063] Embodiment 2: Boundary delineation of attackable area of guidance system
[0064] This example is to depict the Launch Acceptable Region (LAR) of a certain guidance system for a specific target. It needs to determine whether the final miss distance can be less than 20 meters (i.e. the interception is successful) under different target initial velocities (parameter 1) and initial interception angles (parameter 2).
[0065] (1) Establish the sample set: randomly select 4 combinations of (target velocity, interception angle) to perform guidance simulation and obtain 4 initial miss distance data.
[0066] (2) Modeling: based on the 4 initial data points, a Gaussian process model is established to predict the miss distance under different initial conditions.
[0067] (3) Sampling iteration:
[0068] Set the performance index as the miss distance, and the performance index threshold y th is 20 meters.
[0069] Use the scoring function a(z) to calculate and find the next most valuable simulation point. This function will tend to select new simulation parameters in the area where the predicted miss distance is close to 20 meters and the model uncertainty is high.
[0070] Run a guidance simulation at the calculated new parameter point to obtain new miss distance data.
[0071] Update the sample set and the Gaussian process model.
[0072] (4) Loop and stop: repeat this process until the preset 30 simulation budget is completed.
[0073] After the sampling stops, the 30 sample points obtained can roughly outline the contour line of the miss distance of 20 meters. This line is the boundary of the launch acceptable region of the guidance system in this scenario, thereby providing a basis for analyzing the operational effectiveness of the system and the use method.
[0074] It can be understood that the present application is described by examples, and those skilled in the art know that various changes or equivalent replacements can be made to these features and examples without departing from the spirit and scope of the present application. In addition, under the guidance of the present application, these features and examples can be modified to adapt to specific conditions without departing from the spirit and scope of the present application. Therefore, the present application is not limited by the specific examples disclosed herein, and the embodiments falling within the scope of the claims of the present application are within the scope of protection of the present application.
[0075] The contents not described in detail in the specification of the present application are known to those skilled in the art.
Claims
1. A data sampling method, used to select the parameter combination for simulation when evaluating the performance indicators of a system under different parameter combinations through a finite number of simulations, characterized in that, Includes the following steps: S1. Randomly select N parameter combinations, and simulate the system to be evaluated under each parameter combination to obtain the performance index of the system to be evaluated under each parameter combination; perform a total of N simulations to obtain N performance indexes; the N parameter combinations and N performance indexes are referred to as the sample set; N is a positive integer, and N < the total number of simulations; S2. Using the sample set, establish a surrogate model; the surrogate model is used to provide the predicted mean μ(z) and predicted standard deviation σ(z) for any combination of parameters z; S3. Establish a scoring function α(z); the scoring function is used to evaluate the value score of all parameter combinations not selected by S1. When the predicted mean μ(z) of the selected parameter combination is closer to the performance index threshold and the prediction standard deviation σ(z) is larger, the value score is higher. When the selected parameter combination is closer to the already selected parameter combination, the value score is lower. S4. Using the scoring function, find the parameter combination z with the highest value score. * Using parameter combination z * A new simulation is performed on the system to be evaluated to obtain the performance indicators of the simulation. S5. Add the parameter combinations and performance indicators from S4 to the sample set and update the agent model; S6. Repeat steps S4 and S5 until the preset number of simulations is reached.
2. The data sampling method according to claim 1, characterized in that: The scoring function α(z) is as follows: α(z)=α BE (z)-ηα P (z) In the formula, η is the adjustment coefficient; exp() represents the exponential function, y th κ represents the performance index threshold of the system to be evaluated; κ is a hyperparameter that controls the search intensity of the performance index threshold. min represents taking the minimum value, z k This represents the selected parameter combination, where τ is the hyperparameter controlling the repulsion radius.
3. The data sampling method according to claim 1, characterized in that: The proxy model adopts a Gaussian process model.
4. The data sampling method according to claim 1, characterized in that: N≤5。 5. A data sampling system, used to select the parameter combination for simulation when evaluating the system's performance under different parameter combinations through a finite number of simulations, characterized in that, include: The preliminary sampling module is used to randomly select N parameter combinations to form a sample set; The sample set includes N parameter combinations and N performance indicators; the N performance indicators are the performance indicators of the system to be evaluated under each parameter combination after simulation; N is a positive integer, N < the total number of simulations; The proxy module contains a proxy model; the proxy model is established using the sample set and is used to provide the predicted mean μ(z) and predicted standard deviation σ(z) for any combination of parameters z. The scoring module contains a scoring function α(z). The scoring function is used to evaluate the value score of all parameter combinations that were not selected by the initial sampling module. The higher the value score is, the closer the predicted mean μ(z) of the parameter combination to be selected is to the performance index threshold and the larger the prediction standard deviation σ(z) is. The lower the value score is, the closer the parameter combination to be selected is to the parameter combination that has already been selected. The value sampling module is used to find the parameter combination z with the highest value score using the scoring function. * ; The update module is used to combine parameters z. * and in parameter combination z * After simulating the system to be evaluated, the performance indicators obtained from this simulation are added to the sample set to update the proxy model. The control module is used to repeatedly call the value sampling module and the update module until the preset number of simulations is reached.
6. The data sampling system according to claim 5, characterized in that: The scoring function α(z) is as follows: α(z)=α BE (z)-ηα P (z) In the formula, η is the adjustment coefficient; exp() represents the exponential function, y th κ represents the performance index threshold of the system to be evaluated; κ is a hyperparameter that controls the search intensity of the performance index threshold. min represents taking the minimum value, z k This represents the selected parameter combination, where τ is the hyperparameter controlling the repulsion radius.
7. The data sampling system according to claim 5, characterized in that: The proxy model adopts a Gaussian process model.
8. The data sampling system according to claim 5, characterized in that: N≤5。 9. A computer program product, characterized in that: When the computer program product is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that: When the computer program within the medium is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.