Improved sampling method and system for risk factor scene generation and reduction

By introducing optimization algorithms and mixed sampling strategies into the Latin hypercube sampling method, the problem of insufficient sample uniformity and diversity in the prior art is solved, and the accuracy and efficiency of risk assessment and prediction are improved.

CN120146574APending Publication Date: 2025-06-13COMPREHENSIVE SERVICE CENT OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510240717.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing Latin hypercube sampling method has shortcomings in sample uniformity, sampling efficiency and sample diversity, which affects the accuracy of risk assessment and prediction.

Method used

By introducing optimization algorithms and mixed sampling strategies, the location of initial sample points is optimized, ensuring uniform distribution of samples in multidimensional space, and enhancing sample diversity through Monte Carlo samples.

Benefits of technology

It improves the uniformity and diversity of sample distribution, generates more accurate and efficient risk factor scenarios, improves the accuracy of risk assessment and prediction, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146574A_ABST
    Figure CN120146574A_ABST
Patent Text Reader

Abstract

The invention discloses an improved sampling method and system for risk factor scene generation and reduction, and belongs to the technical field of risk data processing. The system comprises a scene sampling module, a sampling optimization module, a curve generation module, computer equipment and a readable storage medium in which a computer program is stored. Based on the computer program, the method comprises the steps of sample point initialization, sample point optimization, mixed sampling, complexity optimization and risk curve generation. According to the method, the optimization algorithm and the mixed sampling strategy are introduced, so that the sample distribution uniformity, the sampling efficiency and the sample diversity are improved, and a more accurate and efficient risk factor scene is generated. The method not only can better represent complex risk factors, but also can effectively reduce redundant scenes, is combined with prediction data, and improves the accuracy of risk assessment and prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of risk data processing, and particularly to an improved sampling method and system for risk factor scenario generation and reduction. Background Art

[0002] In modern risk management and predictive analysis, generating and reducing risk factor scenarios is a key step. The traditional Latin Hypercube Sampling (LHS) method is widely used because it can effectively generate high-dimensional sample data. However, the existing LHS methods have certain deficiencies in sample uniformity, sampling efficiency, and sample diversity, which affect the accuracy of risk assessment and prediction.

[0003] The traditional LHS method ensures the representativeness of samples by evenly distributing sample points in each dimension. However, due to the possible uneven distance between sample points, the generated samples are unevenly distributed in high-dimensional space. This uneven distribution may lead to the underestimation or neglect of some important risk factors, thereby affecting the accuracy of the prediction model. In addition, the existing LHS methods have a high computational complexity and low efficiency when dealing with large-scale data, making it difficult to meet the needs of practical applications. Summary of the Invention

[0004] The purpose of the present invention is to provide an improved sampling method and system for risk factor scenario generation and reduction. By introducing an optimization algorithm and a hybrid sampling strategy, the uniformity, sampling efficiency, and sample diversity of the sample distribution are improved, so as to generate more accurate and efficient risk factor scenarios. It can not only better represent complex risk factors, but also effectively reduce redundant scenarios, combined with prediction data, to improve the accuracy of risk assessment and prediction, so as to solve the deficiencies of the existing technology.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: An improved sampling method for risk factor scenario generation and reduction, comprising the following steps: S1: Initialize sample points: Generate initial sample points through the traditional Latin Hypercube Sampling method to prepare for subsequent optimization and hybrid sampling; S2: Optimize sample points: Use the genetic algorithm to optimize the initial sample points in S1; S3: Hybrid sampling: Introduce a certain proportion of Monte Carlo samples into the samples optimized in S2, set the proportion of Monte Carlo samples, and generate the corresponding number of Monte Carlo sample points to form the final hybrid sample set; S4: Complexity optimization: Adopt optimized matrix operations to reduce redundant calculations and memory consumption; S5: Generate a risk curve: Divide the mixed sample points into several clusters through a clustering analysis method, calculate the center and weight of each cluster, and then draw a risk curve based on the data of the cluster centers.

[0006] Further, the specific method in S1 is as follows: S101: Evenly distribute the samples in each dimension through the traditional Latin hypercube sampling method to ensure the representativeness of the samples; S102: Set the number of samples and generate standardized sample points in each dimension; S103: Just map the standardized sample points to the specific risk factor range.

[0007] Further, the optimization process in S2 includes defining an objective function and minimizing the uniformity index of the sample points by this function. By continuously iterating and adjusting the positions of the sample points, finally, the optimized sample points are obtained.

[0008] Further, the objective function and the function for minimizing the uniformity index of the sample points in S2 are as follows: Define the objective function using the minimum distance uniformity index, and the specific mathematical modeling is as follows: Suppose there are sample points, and each sample point is represented as ( ), where is a -dimensional vector representing the values on risk factor dimensions. Let represent the Euclidean distance between sample points and : ; Among them, represents the value of sample point on the th dimension; The minimum distance objective function is defined as the sum of the minimum values of the distances from all sample points to their nearest neighbor sample points: ; The goal is to minimize this function value to ensure the uniform distribution of sample points in the multi-dimensional space.

[0009] The present invention provides another technical solution: An improved sampling system for risk factor scenario generation and reduction, including: A scenario sampling module, based on the improved Latin hypercube sampling method, captures the changes in risk factors by evenly distributing sample points in the multi-dimensional space, is used to generate initial sample points, and maps them to the actual risk factor range, laying a foundation for subsequent optimization and mixed sampling; The sampling optimization module introduces a certain proportion of Monte Carlo samples and utilizes the genetic algorithm to further optimize the initial sample points. The optimization objective is to minimize the distance non-uniformity between sample points, which is used to ensure uniform distribution in the multi-dimensional space and improve the representativeness and coverage of sample points. The curve generation module is a tool for simulating and analyzing the characteristics of risk factors. The curve generation module utilizes the optimized sample point set and combines interpolation and clustering analysis methods to generate risk curves.

[0010] Furthermore, it also includes a computer device. The computer device is built-in with a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, the processor is used to implement the steps of the sampling method when executing the computer program.

[0011] Furthermore, it also includes a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the sampling method.

[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. For the improved sampling method and system for risk factor scenario generation and reduction of the present invention, the improved LHS adjusts the positions of sample points through an optimization algorithm to ensure uniform distribution in the entire sample space. This uniformity can better represent the changes of risk factors in the real scenario, improving the representativeness and credibility of sampling.

[0013] 2. For the improved sampling method and system for risk factor scenario generation and reduction of the present invention, the improved LHS method can effectively reduce unnecessary calculation and analysis work. The optimized sample point set is more compact and representative, significantly improving the efficiency of subsequent analysis and model establishment. It not only saves time and resources but also reduces the operation cost, especially in the analysis of complex models that require a large number of sample points.

[0014] 3. For the improved sampling method and system for risk factor scenario generation and reduction of the present invention, since the improved LHS can better capture the distribution characteristics of sample points in the multi-dimensional space, it can improve the prediction accuracy of the model. In the application of risk assessment and prediction, the uniform distribution and representativeness of sample points directly affect the accuracy and reliability of the model, thus providing more credible analysis results for decision-makers. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a schematic flowchart of the method of the embodiment of the present invention; Figure 2 is a schematic diagram of generating sample data by initial Latin hypercube sampling of the embodiment of the present invention; Figure 3 The final sample distribution diagram generated by the hybrid sampling strategy of the embodiment of the present invention; Figure 4 The humidity risk curve graph of the embodiment of the present invention; Figure 5 The temperature risk curve graph of the embodiment of the present invention; Figure 6 The matching graph of the predicted value and the corresponding scenario of the embodiment of the present invention. Specific implementation manners

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Please refer to Figure 1 , an improved sampling method for risk factor scenario generation and reduction provided by the embodiment of the present invention includes the following steps: S1: Initialize sample points: Generate initial sample points through the traditional Latin hypercube sampling method. The LHS method evenly distributes the samples in each dimension to ensure the representativeness of the samples; set the number of samples and generate standardized sample points in each dimension; these standardized sample points will then be mapped to the specific risk factor ranges, such as humidity and temperature in the case of the present invention. In this way, the initial humidity and temperature sample data can be obtained, preparing for subsequent optimization and hybrid sampling; S2: Optimize sample points: Use the genetic algorithm to optimize the initial LHS sample points. The goal of optimization is to ensure that the sample points are evenly distributed in each dimension and reduce the non-uniformity of the distances between the sample points. The optimization process includes defining an objective function, which minimizes the uniformity index of the sample points. By continuously iterating and adjusting the positions of the sample points, the optimized sample points are finally obtained. These sample points are more evenly distributed in the multi-dimensional space, thereby improving the representativeness and coverage of the samples; In this step, the objective function and the function for minimizing the uniformity index of the sample points are as follows: In the optimization process, the definition of the objective function is a key step. The objective function is used to measure the uniformity of the distribution of the sample points and optimize the sample distribution by minimizing the value of this function. Specifically, the objective function can be defined as the uniformity index of the distances between the sample points. The uniformity index aims to measure the degree of uniform distribution of the sample points in the multi-dimensional space. Commonly used uniformity indexes include the minimum distance, the maximum distance, and the average distance, etc. In the embodiment of the present invention, the minimum distance uniformity index is adopted to define the objective function; the specific mathematical modeling is as follows: Suppose there are sample points, and each sample point is represented as ( ), where is a -dimensional vector representing the values on risk factor dimensions. Let represent the Euclidean distance between sample points and : ; where represents the value of sample point on the -th dimension.

[0018] The minimum distance objective function is defined as the sum of the minimum distances from all sample points to their nearest neighbor sample points: ; The goal is to minimize this function value to ensure that the distribution of sample points in the multi-dimensional space is as uniform as possible; S3: Hybrid Sampling: To further enhance the diversity of samples, a certain proportion of Monte Carlo samples (Monte Carlo Sampling) are introduced into the optimized LHS samples, the proportion of Monte Carlo samples is set, and the corresponding number of Monte Carlo sample points are generated to form the final hybrid sample set. The introduction of the hybrid sampling strategy makes the samples not only have uniformity but also higher diversity, and can more comprehensively represent complex risk factors; S4: Complexity Optimization: Optimized matrix operations are adopted to reduce redundant calculations and memory consumption; efficient algorithms and data structures are used for matrix operations to ensure that the distance matrix between sample points can be calculated quickly and accurately; S5: Generate Risk Curve: The hybrid sample points are divided into several clusters through cluster analysis methods, the centers and weights of each cluster are calculated, and then the risk curve is plotted based on the data of the cluster centers.

[0019] In the embodiment of the present invention, to verify the effectiveness of the established model, an example is used for calculation and analysis. In the embodiment of the present invention, matlab is used for verification operations: Step 1: Initialize sample points: Initial sample points are generated through the traditional Latin hypercube sampling method; the number of samples taken in this embodiment is 100, and the risk factors selected are humidity and temperature, where the range of humidity is [0, 100], and the range of temperature is [-20, 50]. The specific selection range is adjusted according to the specific area. The generated sample points are as shown in Figure 2 .

[0020] Step 2: Optimize the sample points. The objective function adopted during the optimization process is as follows: ; The number of iterations of the genetic algorithm is set to 1000; Step 3: Hybrid sampling: Introduce a certain proportion of Monte Carlo samples into the optimized LHS samples; mc_ratio = 0.1; mc_samples = rand(round(n_samples * mc_ratio), 2); In the hybrid sampling, the proportion of Monte Carlo samples is 10% of the total number of samples. Specifically, it multiplies the total number of samples n_samples by mc_ratio, and then uses the rand function to generate random samples of the corresponding proportion. The mixed data samples are as Figure 3 shown.

[0021] Step 4: Perform cyclic iteration reduction in 100 scenarios until it is reduced to 10 typical scenarios and ends. Finally, obtain the typical scenario temperature and humidity risk curves as Figure 5 and Figure 6 , and at the same time, for better understanding in the follow-up, in this embodiment, on the original basis, the prediction values are matched with the risk scenarios, so that the range of risks can be made more refined and more detailed measures can be taken.

[0022] To further better explain the above method, an embodiment of the present invention also provides an improved sampling system for risk factor scenario generation and reduction, including: A scenario sampling module, based on the improved Latin hypercube sampling method, captures the changes of risk factors by evenly distributing sample points in a multi-dimensional space, is used to generate initial sample points, and maps them to the actual risk factor range, laying a foundation for subsequent optimization and hybrid sampling; A sampling optimization module, introduces a certain proportion of Monte Carlo samples and uses the genetic algorithm to further optimize the initial sample points; the optimization goal is to minimize the non-uniformity of the distances between sample points, which is used to ensure uniform distribution in the multi-dimensional space and improve the representativeness and coverage of sample points; A curve generation module, a tool for simulating and analyzing the characteristics of risk factors. The curve generation module uses the optimized sample point set and combines interpolation and clustering analysis methods to generate risk curves.

[0023] In the above system, it also includes a computer device. The computer device is built-in with a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, the processor is used to implement the steps of the sampling method when executing the computer program.

[0024] In the above system, a computer-readable storage medium is further included, on which a computer program is stored. When the computer program is executed by a processor, the steps of the sampling method are implemented.

[0025] Among them, the computer program for implementing the steps of the above sampling method is as follows: % Number of samples n_samples = 100; % Generate Latin Hypercube samples lhs_samples = lhsdesign(n_samples, 2); % Map the samples to the specific risk factor range humidity_range = [0, 100]; temperature_range = [-20, 50]; humidity_samples = lhs_samples(:, 1) *(humidity_range(2) - humidity_range(1)) + humidity_range(1); temperature_samples = lhs_samples(:, 2) *(temperature_range(2) -temperature_range(1)) + temperature_range(1); % Combine the sample data samples = [humidity_samples, temperature_samples]; % Plot the initial sample distribution diagram ( Figure 1 ) figure; scatter(samples(:, 1), samples(:, 2), 'b', 'filled'); xlabel('Humidity (%)'); ylabel('Temperature (°C)'); title('Initial LHS Samples'); % Genetic algorithm parameters options = optimoptions('ga', 'PopulationSize', n_samples, 'MaxGenerations',500, 'Display', 'off'); % Objective function: Minimize the uniformity index of sample points objective = @(x) -sum(min(pdist2(x, x), [], 2)); % Optimize sample points lb = [humidity_range(1), temperature_range(1)]; ub = [humidity_range(2), temperature_range(2)]; optimized_samples = ga(objective, 2, [], [], [], [], lb, ub, [],options); % Plot the distribution of optimized samples ( Figure 2 ) figure; scatter(optimized_samples(:, 1), optimized_samples(:, 2), 'r', 'filled'); xlabel('Humidity (%)'); ylabel('Temperature (°C)'); title('Optimized LHS Samples'); % Monte Carlo sample ratio mc_ratio = 0.1; mc_samples = rand(round(n_samples * mc_ratio), 2); mc_samples(:, 1) = mc_samples(:, 1) * (humidity_range(2) - humidity_range(1)) + humidity_range(1); mc_samples(:, 2) = mc_samples(:, 2) * (temperature_range(2) - temperature_range(1)) + temperature_range(1); % Mixed sample data final_samples = [optimized_samples; mc_samples]; % Plot the distribution of the mixed samples ( Figure 3 ) figure; scatter(final_samples(:, 1), final_samples(:, 2), 'g', 'filled'); xlabel('Humidity (%)'); ylabel('Temperature (°C)'); title('Mixed LHS and Monte Carlo Samples'); % Efficiently calculate the distance matrix between samples distance_matrix = pdist2(final_samples, final_samples); % K-means clustering n_clusters = 10; [idx, cluster_centers] = kmeans(final_samples, n_clusters); % Calculate the weights of each cluster counts = histcounts(idx, n_clusters); weights = counts / size(final_samples, 1); % Plot the humidity risk curve ( Figure 4 ) figure; scatter(1:n_clusters, cluster_centers(:, 1), 'b', 'filled'); hold on; xq = linspace(1, n_clusters, 100); vq = interp1(1:n_clusters, cluster_centers(:, 1), xq,'spline'); plot(xq, vq, 'b'); xlabel('Scenario'); ylabel('Humidity (%)'); title('Humidity Risk Curve'); hold off; % Plot the temperature risk curve ( Figure 5 ) figure; scatter(1:n_clusters, cluster_centers(:, 2), 'b', 'filled'); hold on; xq = linspace(1, n_clusters, 100); vq = interp1(1:n_clusters, cluster_centers(:, 2), xq,'spline'); plot(xq, vq, 'b'); xlabel('Scenario'); ylabel('Temperature (°C)'); title('Temperature Risk Curve'); hold off; % Define the prediction data (assuming 10 time points of prediction data) predictions = 50, 10; 60, 20; 70, 30; 80, 40; 90, 50; 40, 15; 55, 25; 65, 35; 75, 45; 85, 5; ; % Match prediction data to the nearest scenario matched_scenarios = zeros(size(predictions, 1), 1); for i = 1:size(predictions, 1) distances = sqrt(sum((cluster_centers - predictions(i, :)).^2,2)); [~, matched_scenarios(i)] = min(distances); end % Plot predicted values against corresponding scenarios figure; subplot(2, 1, 1); % The first subplot, humidity predictions vs. scenarios scatter(1:size(predictions, 1), predictions(:, 1), 'r', 'filled'); hold on; scatter(matched_scenarios, predictions(:, 1), 'b', 'filled'); xlabel('Scenario'); ylabel('Humidity (%)'); title('Predicted Humidity vs. Scenarios'); legend('Predicted value', 'Matched scenario'); subplot(2, 1, 2); % The second subplot, temperature predictions vs. scenarios scatter(1:size(predictions, 1), predictions(:, 2), 'r', 'filled'); hold on; scatter(matched_scenarios, predictions(:, 2), 'b', 'filled'); xlabel('Scenario'); ylabel('Temperature (°C)'); title('Predicted Temperature vs. Scenarios'); legend('Predicted Value', 'Matched Scenarios'); sgtitle('Matching of Predicted Value and Corresponding Scenarios'); % If you need to save the graph, you can use the following command % saveas(gcf, 'predictions_vs_scenarios.png').

[0026] In summary, an improved sampling method and system for risk factor scenario generation and reduction provided by the present invention has the advantage of ensuring uniform distribution within the entire sample space, improving the representativeness and credibility of sampling, effectively reducing unnecessary calculation and analysis work. The optimized sample point set is more compact and representative, significantly improving the efficiency of subsequent analysis and model establishment. This not only saves time and resources but also reduces the computing cost, especially in complex model analysis that requires a large number of sample points. Secondly, since the improved LHS can better capture the distribution characteristics of sample points in the multi-dimensional space, it can improve the prediction accuracy of the model. In the application of risk assessment and prediction, the uniform distribution and representativeness of sample points directly affect the accuracy and reliability of the model, thus providing more credible analysis results for decision-makers. In addition, the improved LHS method is not only applicable to risk analysis in a single field but can also be extended to the sampling requirements of various multi-dimensional problems. Whether it is engineering design, environmental monitoring, or financial risk management, this method can provide a unified and effective sample point generation scheme. Its versatility makes the improved LHS an ideal choice for dealing with complex systems and variable factors.

[0027] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An improved sampling method for risk factor scenario generation and reduction, characterized in that: The following steps are involved: S1: Initialize sample points: Generate initial sample points through the traditional Latin hypercube sampling method to prepare for subsequent optimization and mixed sampling; S2: Optimize sample points: Use genetic algorithm to optimize the initial sample points in S1; S3: Mixed sampling: introduce a certain proportion of Monte Carlo samples into the samples optimized in S2, set the proportion of Monte Carlo samples, and generate a corresponding number of Monte Carlo sample points to form the final mixed sample set; S4: Complexity optimization: Use optimized matrix operations to reduce redundant calculations and memory consumption; S5: Generate risk curve: Divide the mixed sample points into several clusters through cluster analysis method, calculate the center and weight of each cluster, and then draw the risk curve based on the data of the cluster center.

2. The improved sampling method for risk factor scenario generation and reduction as claimed in claim 1, characterized in that: The specific method in S1 is as follows: S101: The samples are evenly distributed in each dimension through the traditional Latin hypercube sampling method to ensure the representativeness of the samples; S102: setting the number of samples and generating standardized sample points in each dimension; S103: Map the standardized sample points to the specific risk factor range.

3. The improved sampling method for risk factor scenario generation and reduction according to claim 1, characterized in that: The optimization process in S2 includes defining the objective function and minimizing the uniformity index of the sample points by the function. By continuously iterating and adjusting the positions of the sample points, the optimized sample points are finally obtained.

4. The improved sampling method for risk factor scenario generation and reduction as claimed in claim 3, characterized in that: The uniformity index of the objective function and the function minimization sample points in S2 is as follows: The minimum distance uniformity index is used to define the objective function, and the specific mathematical modeling is as follows: Assume that sample points, each sample point is represented by ( ),in is a dimensional vector, represented by The value of the risk factor dimension is Represents sample points and The Euclidean distance between: ; in, Represents sample points In the The value of the dimension; The minimum distance objective function is defined as the minimum sum of the distances from all sample points to their nearest neighboring sample points: ; The goal is to minimize the value of this function to ensure that the sample points are evenly distributed in the multidimensional space.

5. An improved sampling system for risk factor scenario generation and reduction, characterized in that: include: The scenario sampling module, based on the improved Latin hypercube sampling method, captures the changes in risk factors by evenly distributing sample points in multidimensional space. It is used to generate initial sample points and map them to the actual risk factor range, laying the foundation for subsequent optimization and mixed sampling. The sampling optimization module introduces a certain proportion of Monte Carlo samples and uses genetic algorithms to further optimize the initial sample points; The optimization goal is to minimize the distance inhomogeneity between sample points, which is used to ensure uniform distribution in multidimensional space and improve the representativeness and coverage of sample points; The curve generation module is a tool for simulating and analyzing the characteristics of risk factors. The curve generation module uses the optimized sample point set combined with interpolation and cluster analysis methods to generate risk curves.

6. An improved sampling system for risk factor scenario generation and reduction as claimed in claim 5, characterized in that: It also includes a computer device, which has a built-in memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the sampling method when executing the computer program.

7. An improved sampling system for risk factor scenario generation and reduction as claimed in claim 6, characterized in that: It also includes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the sampling method are implemented.