Multi-dimensional index power distribution system modeling method fusing random forest and particle swarm optimization

By integrating the XGBoost method with random forest and particle swarm optimization, a multi-dimensional index evaluation system is constructed, which solves the problems of insufficient dimensional coverage and subjective weights in the flexibility assessment of power distribution systems, improves the scientificity and accuracy of the assessment, and adapts to complex disturbance environments.

CN122066286APending Publication Date: 2026-05-19STATE GRID ANHUI ELECTRIC POWER CO LTD BOZHOU POWER SUPPLY CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID ANHUI ELECTRIC POWER CO LTD BOZHOU POWER SUPPLY CO
Filing Date
2025-12-31
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for assessing the flexibility of power distribution systems suffer from insufficient dimensional coverage, highly subjective index weights, and limited model performance, making it difficult to meet the needs for rapid and high-precision assessment under complex disturbance environments.

Method used

We employ the XGBoost method, which integrates random forest and particle swarm optimization, to construct a multi-dimensional index evaluation system. We extract index weights using the random forest algorithm and optimize the hyperparameters of the XGBoost model using an improved particle swarm optimization algorithm, thereby enhancing the model's prediction accuracy and stability.

Benefits of technology

It achieves scientific rigor, accuracy, and robustness in power distribution system flexibility assessment, enhances the model's adaptability and generalization ability, and meets practical engineering needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066286A_ABST
    Figure CN122066286A_ABST
Patent Text Reader

Abstract

According to the multi-dimensional index power distribution system modeling method fusing the random forest and the particle swarm optimization, a two-dimensional flexibility evaluation system containing six key indexes is constructed, and the operation adjusting capacity and stability of a power distribution system under multi-source disturbance are systematically described. Feature weights are extracted through a random forest algorithm, and objectivity and robustness of an evaluation model are improved; and in combination with an XGBoost modeling framework which introduces an improved particle swarm optimization algorithm, global efficient optimization of model hyper-parameters is realized, and the prediction precision and generalization ability are effectively improved. In addition, through typical daily clustering, normalization processing and a fusion scoring mechanism, the ability of the model to adapt to multi-scene operation conditions is enhanced. Experimental results show that the provided method is superior to a traditional method in the aspects of flexibility prediction precision, model stability, feature interpretation and the like, and the application value of the method in intelligent evaluation and flexibility quantification of the power distribution system is verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This research relates to the field of power distribution network operation and intelligent dispatching technology, specifically to a multi-dimensional index power distribution system flexibility modeling method that integrates random forest and particle swarm optimization algorithms. Background Technology

[0002] As a crucial link connecting power sources and loads in the power system, the distribution network's operational stability and regulation capabilities play a decisive role in the overall safety and reliability of the power system. With the large-scale integration of high-proportion renewable energy sources and increasingly complex user load behavior, the operational uncertainties and control difficulties faced by the distribution system have significantly increased. Traditional dispatching and assessment methods are no longer sufficient to meet the demands for flexible and efficient operation. In environments with multi-source disturbances, how to scientifically assess the flexibility of the distribution system has become a key technical issue for achieving flexible dispatching and efficient operation.

[0003] Distribution system flexibility is a crucial indicator of its ability to cope with external disturbances such as fluctuations in renewable energy output and dramatic load changes. It encompasses multiple dimensions, including system regulation capacity, operational performance, and resource coordination capabilities. In recent years, scholars both domestically and internationally have conducted extensive research on distribution network flexibility modeling and assessment, proposing various methods based on mathematical programming, simulation calculations, and machine learning. For example, some studies have constructed multi-dimensional index systems to quantitatively assess system flexibility from aspects such as absorption capacity, voltage stability, and load regulation; others have introduced multi-objective optimization methods to jointly optimize flexibility resource allocation and operational strategies; and some studies have also utilized deep learning or ensemble learning methods to train and predict flexibility scoring models.

[0004] However, existing flexibility assessment methods still suffer from problems such as insufficient dimensional coverage, strong subjectivity in weight allocation, and weak model generalization ability, making it difficult to comprehensively and accurately characterize the response capability of power distribution systems under complex disturbance environments. Furthermore, model performance largely depends on parameter settings and feature selection strategies; traditional fixed-weight or manual parameter tuning methods are inefficient and poorly adaptable, failing to meet the demands of practical engineering for fast and high-precision assessment tools.

[0005] To address the aforementioned issues, this invention proposes a multi-dimensional index-based modeling method for distribution system flexibility that integrates random forest and PSO-optimized XGBoost. This method constructs a flexibility evaluation system with six indices across two dimensions, comprehensively covering system regulation and operational capabilities. It utilizes the random forest algorithm to extract importance weights for the indices, overcoming the limitations of subjective assignment. Furthermore, it employs an improved particle swarm optimization algorithm to fine-tune the hyperparameters of the XGBoost regression model, enhancing both prediction accuracy and model stability and convergence efficiency. This method provides an effective approach to achieving objectivity, accuracy, and practicality in distribution system flexibility modeling, and has promising engineering application prospects. Summary of the Invention

[0006] The purpose of this invention is to address the problems of incomplete dimensional coverage, strong subjectivity in indicator weights, and limited model performance in existing methods for assessing the flexibility of power distribution systems. To address these issues, this invention proposes an XGBoost multidimensional flexibility modeling method that integrates random forest and particle swarm optimization. This method significantly improves the scientific rigor, accuracy, and robustness of power distribution system flexibility modeling by constructing a comprehensive multidimensional indicator evaluation framework and combining machine learning and intelligent optimization techniques.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This invention provides a multi-dimensional index-based power distribution system modeling method that integrates random forest and particle swarm optimization algorithms. The method includes the following steps: Step 1: Construct a flexibility index evaluation system, which includes a flexibility evaluation framework with six key indicators. This framework comprehensively considers multiple dimensions such as renewable energy absorption capacity, fluctuation response level, voltage stability, network transmission efficiency, and power fluctuation tolerance, reflecting the overall operational flexibility of the power distribution system from multiple perspectives.

[0008] Step 2: Collect data from typical operating scenarios. Based on historical wind power, photovoltaic, and load data, K-means clustering is used to extract several representative typical daily scenarios to enhance the diversity and representativeness of the model training data, and the probability weight of each scenario is calculated. Indicator normalization and label construction are then performed. To address the issues of inconsistent dimensions and evaluation directions in the flexibility indicators, forward and reverse normalization methods are used to preprocess the data, and a fixed-weighting method is combined to construct comprehensive scoring labels as the model's learning objective.

[0009] Step 3: Extract the importance of indicator features. Utilize a random forest regression model to calculate the relative contribution of each indicator to the flexibility score, generating normalized feature weights to overcome the bias problem of subjective weighting and realize a data-driven weight evaluation mechanism.

[0010] Step 4: Construct and optimize the XGBoost regression model. Using the key features selected by random forest as input, build a flexible XGBoost prediction model. Introduce an improved particle swarm optimization (PSO) algorithm to jointly tune the model's hyperparameters, improving the model's prediction accuracy and generalization ability.

[0011] Step 5: Design an improved particle swarm optimization strategy. Global search capability is enhanced through exponentially decaying inertial weights, population stability is maintained by combining an average update mechanism of individual and group extrema, and a local perturbation mechanism is embedded to avoid getting trapped in local optima.

[0012] Step 6: Complete model training and performance evaluation. Evaluate the model performance on both the training and test sets. Verify the effectiveness and superiority of the proposed modeling method by comparing metrics such as mean squared error, scoring error, and ROC curves. Output the final flexibility evaluation results. Furthermore, the evaluation system in step 1 of the present invention is as follows: The renewable energy absorption rate refers to the proportion of renewable energy generation that is absorbed and consumed by the power system within a certain period, relative to its total power generation or total load. This proportion reflects the utilization rate of renewable energy in the overall power supply structure. The formula for calculating the renewable energy absorption rate is as follows:

[0013] in, Indicates the renewable energy consumption rate. This represents the amount of renewable energy generated that is actually absorbed by the power grid. This indicates the maximum load power of the distribution network.

[0014] The renewable energy fluctuation support level refers to the power system's ability to maintain stable grid operation when it accommodates a large amount of fluctuating renewable energy (such as wind and solar power). This indicator reflects the power system's ability to cope with fluctuations in renewable energy generation, ensuring that the grid's frequency, stability, and load balance are not affected by severe fluctuations in renewable energy generation. The formula for calculating the renewable energy fluctuation support level is as follows:

[0015]

[0016]

[0017] in: For time period The actual interaction efficiency of flexible resources This represents the total number of nodes connected to the flexible resources. and Time periods The output power of flexible resources and the output power of the previous period; and Time periods The output power of new energy sources and the output power of the previous period.

[0018] The peak-to-valley ratio variation rate refers to the rate at which the peak-to-valley ratio in an electricity load curve changes over time, and is used to measure the degree of electricity load fluctuation. The peak-to-valley ratio describes the ratio of peak load to valley load in a power system, and is typically used to assess the smoothness of the load curve and the dispatching capacity of the power system. The peak-to-valley ratio variation rate reflects the severity of load fluctuations; the larger the variation rate, the more frequent and larger the fluctuations in electricity load between peak and valley values. The formula for calculating the peak-to-valley ratio variation rate is as follows:

[0019] Network loss refers to the power loss that occurs during power transmission and distribution due to factors such as resistance and conductor inductance. It is an unavoidable phenomenon in power grid operation and affects the efficiency of the power system. The formula for calculating network loss is as follows:

[0020] in It is the power grid loss (power loss, in MW or KW). It is the current on the i-th line. It is the resistance on the i-th line. It represents the total number of transmission lines in the system.

[0021] Average voltage deviation measures the degree of deviation of the voltage at each node of a system from the nominal voltage, reflecting the stability of voltage quality. The formula for calculating average voltage deviation is as follows:

[0022] in It is the average voltage deviation (percentage); It is the first The voltage of each node; It is the nominal voltage; It represents the total number of nodes.

[0023] Maximum permissible volatility is the rate at which the power grid can withstand fluctuations in load or renewable energy generation. It measures the grid's tolerance for rapid load changes or intermittent renewable energy generation.

[0024]

[0025] in It is the maximum permissible volatility (in MW / minute or MW / hour). It is the maximum power change that the power grid can withstand. It is a time interval.

[0026] Furthermore, the typical scenario daily generation and positive / negative normalization in step 2 of the present invention are as follows: Step 2-1: Typical Scenario Daily Generation Before clustering, the data is preprocessed to remove missing or non-numerical data, resulting in cleaned load data. Subsequently, the K-means clustering algorithm was used to divide the load curves into K classes, each class... The center curve (i.e., typical day) is denoted as The calculation formula is as follows:

[0027] in, This represents the set of all days that are classified into the i-th class. This is the load curve for day d.

[0028] To measure the representativeness of each type of typical day, the probability of occurrence of each type, pi, is further calculated.

[0029] Step 2-2: Midpoint-guided mutation rate update strategy.

[0030] Considering the inconsistent dimensions and evaluation directions of various flexibility indicators, the raw data is first normalized. For positive indicators where larger values ​​indicate better evaluation results, such as renewable energy absorption rate, volatility support level, peak-to-valley ratio change rate, and maximum support volatility, maximum-minimum standardization is adopted, as shown in the following formula.

[0031]

[0032] For negative indicators where "the smaller the value, the better," such as network loss voltage deviation, inverse normalization is used, as shown in the following formula:

[0033] After obtaining all normalized indicators, a fixed-weighted composite score label is constructed. As the supervised learning objective, let the normalized index vector be denoted as . The commonly used flexible weights correspond to fixed weights. Then it can be calculated using the following formula.

[0034]

[0035] Furthermore, the importance of the extracted indicator features in step 3 of the present invention is as follows: Step 3-1: Decision Tree Construction and Feature Splitting Each internal node of the decision tree processes a certain feature. j Set threshold The splitting is performed using the minimum mean square error (MSE) as the splitting criterion:

[0036] Where m is the number of samples in the current node. and The number of samples in the left and right child nodes after the split. and This represents the minimum mean square error of the corresponding child node.

[0037] Step 3-2: Calculation of feature importance Feature Importance (FI) in Random Forest is based on the sum of the impurity reductions brought by a feature across all decision trees:

[0038] in, Let b be the number of internal nodes of the b-th tree. Let be the sample weights of node t. The amount of impurity reduction at node t due to splitting based on feature j.

[0039] For regression trees, impurity is usually measured by MSE, and the contribution of feature j to the importance of node t is:

[0040] Step 3-3: Feature Importance Normalization Normalize the importance values ​​of all features to obtain the objective weights of each indicator:

[0041] Where n is the total number of features. This weight reflects the relative contribution of each indicator to the flexibility score.

[0042] Furthermore, the optimized XGBoost regression model in step 4 of the present invention is as follows: Step 4-1: Gradient Boosting Optimization XGBoost employs an additive training strategy, adding a new tree at each iteration.

[0043] Solve by minimizing the following approximate objective function.

[0044] in and These are the first and second derivatives of the loss function, respectively.

[0045] Step 4-2, Tree Structure Optimization For a given tree structure q(x) (which maps the input to leaf nodes), the optimal leaf node weights are:

[0046] in, This represents the set of samples assigned to leaf node j.

[0047] At this point, the optimal value of the objective function is

[0048] Step 4-3: Feature Splitting and Gain Calculation XGBoost uses a greedy algorithm to find the optimal split point, where the split gain of feature j at split point k is...

[0049] in, and This is the sample set of the left and right child nodes after the split. The sample set of the parent node.

[0050] Step 4-4: Improve the PSO optimization algorithm This invention introduces the Particle Swarm Optimization (PSO) algorithm to jointly search for key hyperparameters on top of XGBoost. XGBoost is an ensemble learning method based on gradient boosting decision trees, and its performance is highly dependent on the settings of several hyperparameters, such as the number of base learners. Maximum tree depth Learning rate Subsample sampling rate and feature column sampling rate Traditional grid search or random search is inefficient in high-dimensional parameter spaces, prone to getting trapped in local optima, and computationally expensive. PSO has strong global optimization capabilities and fast convergence speed, making it suitable for optimization problems in continuous parameter spaces.

[0051] Furthermore, the improved particle swarm optimization strategy in step 5 of the present invention is as follows: Step 5-1: Speed ​​and Position Midpoint Update The individual extreme value (pbest) and the group extreme value (gbest) in the velocity and position update formula are averaged to guide the particle to move towards a more balanced central position, thereby enhancing the global search capability. The specific formula is as follows;

[0052]

[0053] in, and It is a learning factor. and It is a random number in the range (0,1). For the first The historical best position found by an individual particle is called the individual optimal solution. The historical best position found for the entire particle swarm is called the swarm optimal solution. For the first During the first iteration The position of each particle; For inertial weights: Step 5-2, Exponentially Decreasing Dynamic Inertia Weights The specific formula is as follows:

[0054] in, and The upper and lower limits of the inertia weight are set to 0.9 and 0.2, respectively.

[0055] This allows the particles to maintain their exploratory nature in the early stages of the search process while achieving robust convergence in the later stages, significantly improving their performance and stability in XGBoost hyperparameter optimization.

[0056] In summary, this invention proposes a multi-dimensional index-based power distribution system modeling method that integrates random forest and particle swarm optimization algorithms. By constructing a two-dimensional flexibility evaluation system comprising six key indicators, it systematically characterizes the operational adjustment capability and stability of the power distribution system under multi-source disturbances. Feature weights are extracted using the random forest algorithm, improving the objectivity and robustness of the evaluation model. Combined with the XGBoost modeling framework incorporating an improved particle swarm optimization algorithm, global and efficient optimization of model hyperparameters is achieved, effectively improving prediction accuracy and generalization ability. Furthermore, typical daily clustering, normalization processing, and a fusion scoring mechanism enhance the model's ability to adapt to various operating conditions. Experimental results show that the proposed method outperforms traditional methods in terms of flexibility prediction accuracy, model stability, and feature interpretability, validating its application value in intelligent evaluation and flexibility quantification of power distribution systems.

[0057] The beneficial effects of this invention are as follows: 1. Improved accuracy of flexibility assessment: By constructing a six-indicator evaluation system that includes both regulation capability and operational performance, the system comprehensively reflects the power distribution system's ability to cope with multi-source disturbances, thus enhancing the scientific nature of the assessment results.

[0058] 2. Objectification of feature weights: The random forest algorithm is introduced to extract the importance weights of the indicators, avoiding the bias caused by subjective assignment and improving the interpretability and robustness of the scoring model.

[0059] 3. Enhanced adaptability and generalization ability: By introducing a multi-typical operation scenario generation mechanism and a fusion weight strategy, the model has good adaptability and promotion ability under different load and new energy access conditions, meeting the actual needs of engineering. Attached Figure Description

[0060] Figure 1 This is a flowchart illustrating a multi-dimensional index power distribution system flexibility modeling method that integrates random forest and PSO-optimized XGBoost according to an embodiment of the present invention. Figure 2 This is a diagram showing the various flexibility parameters under different schemes in the embodiments of the present invention. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] like Figure 1 As shown in the figure, this embodiment of the invention provides a multi-dimensional index-based power distribution system flexibility modeling method that integrates random forest and PSO-optimized XGBoost. The method includes the following steps: Step 1: Construct a flexibility index evaluation system, which includes a flexibility evaluation framework with six key indicators. This framework comprehensively considers multiple dimensions such as renewable energy absorption capacity, fluctuation response level, voltage stability, network transmission efficiency, and power fluctuation tolerance, reflecting the overall operational flexibility of the power distribution system from multiple perspectives.

[0063] The evaluation system in step 1 of this invention is as follows: The renewable energy absorption rate refers to the proportion of renewable energy generation that is absorbed and consumed by the power system within a certain period, relative to its total power generation or total load. This proportion reflects the utilization rate of renewable energy in the overall power supply structure. The formula for calculating the renewable energy absorption rate is as follows:

[0064] in, Indicates the renewable energy consumption rate. This represents the amount of renewable energy generated that is actually absorbed by the power grid. This indicates the maximum load power of the distribution network.

[0065] The renewable energy fluctuation support level refers to the power system's ability to maintain stable grid operation when it accommodates a large amount of fluctuating renewable energy (such as wind and solar power). This indicator reflects the power system's ability to cope with fluctuations in renewable energy generation, ensuring that the grid's frequency, stability, and load balance are not affected by severe fluctuations in renewable energy generation. The formula for calculating the renewable energy fluctuation support level is as follows:

[0066]

[0067]

[0068] in: For time period The actual interaction efficiency of flexible resources This represents the total number of nodes connected to the flexible resources. and Time periods The output power of flexible resources and the output power of the previous period; and Time periods The output power of new energy sources and the output power of the previous period.

[0069] The peak-to-valley ratio variation rate refers to the rate at which the peak-to-valley ratio in an electricity load curve changes over time, and is used to measure the degree of electricity load fluctuation. The peak-to-valley ratio describes the ratio of peak load to valley load in a power system, and is typically used to assess the smoothness of the load curve and the dispatching capacity of the power system. The peak-to-valley ratio variation rate reflects the severity of load fluctuations; the larger the variation rate, the more frequent and larger the fluctuations in electricity load between peak and valley values. The formula for calculating the peak-to-valley ratio variation rate is as follows:

[0070] Network loss refers to the power loss that occurs during power transmission and distribution due to factors such as resistance and conductor inductance. It is an unavoidable phenomenon in power grid operation and affects the efficiency of the power system. The formula for calculating network loss is as follows:

[0071] in It is the power grid loss (power loss, in MW or KW). It is the current on the i-th line. It is the resistance on the i-th line. It represents the total number of transmission lines in the system.

[0072] Average voltage deviation measures the degree of deviation of the voltage at each node of a system from the nominal voltage, reflecting the stability of voltage quality. The formula for calculating average voltage deviation is as follows:

[0073] in It is the average voltage deviation (percentage); It is the first The voltage of each node; It is the nominal voltage; It represents the total number of nodes.

[0074] Maximum permissible volatility is the rate at which the power grid can withstand fluctuations in load or renewable energy generation. It measures the grid's tolerance for rapid load changes or intermittent renewable energy generation.

[0075]

[0076] in It is the maximum permissible volatility (in MW / minute or MW / hour). It is the maximum power change that the power grid can withstand. It is a time interval.

[0077] Step 2: Collect data from typical operating scenarios. Based on historical wind power, photovoltaic, and load data, K-means clustering is used to extract several representative typical daily scenarios to enhance the diversity and representativeness of the model training data, and the probability weight of each scenario is calculated. Indicator normalization and label construction are then performed. To address the issues of inconsistent dimensions and evaluation directions in the flexibility indicators, forward and reverse normalization methods are used to preprocess the data, and a fixed-weighting method is combined to construct comprehensive scoring labels as the model's learning objective.

[0078] Step 2-1: Typical Scenario Daily Generation Step 2-2: Midpoint-guided mutation rate update strategy.

[0079] Step 3: Extract the importance of indicator features. Utilize a random forest regression model to calculate the relative contribution of each indicator to the flexibility score, generating normalized feature weights to overcome the bias problem of subjective weighting and realize a data-driven weight evaluation mechanism.

[0080] Step 3-1: Decision Tree Construction and Feature Splitting Step 3-2: Calculation of feature importance Step 3-3: Feature Importance Normalization Step 4: Construct and optimize the XGBoost regression model. Using the key features selected by random forest as input, build a flexible XGBoost prediction model. Introduce an improved particle swarm optimization (PSO) algorithm to jointly tune the model's hyperparameters, improving the model's prediction accuracy and generalization ability.

[0081] Step 4-1: Gradient Boosting Optimization Step 4-2, Tree Structure Optimization Step 4-3: Feature Splitting and Gain Calculation Step 4-4: Improve the PSO optimization algorithm Step 5: Design an improved particle swarm optimization strategy. Global search capability is enhanced through exponentially decaying inertial weights, population stability is maintained by combining an average update mechanism of individual and group extrema, and a local perturbation mechanism is embedded to avoid getting trapped in local optima.

[0082] Step 5-1: Speed ​​and Position Midpoint Update Step 5-2, Exponentially Decreasing Dynamic Inertia Weights In this embodiment, step 2-1 can be implemented by the following method: Before clustering, the data is preprocessed to remove missing or non-numerical data, resulting in cleaned load data. Subsequently, the K-means clustering algorithm was used to divide the load curves into K classes, each class... The center curve (i.e., typical day) is denoted as The calculation formula is as follows:

[0083] in, This represents the set of all days that are classified into the i-th class. This is the load curve for day d.

[0084] To measure the representativeness of each type of typical day, the probability of occurrence of each type, pi, is further calculated.

[0085] In this embodiment, step 2-2 can be implemented by the following method: Considering the inconsistent dimensions and evaluation directions of various flexibility indicators, the raw data is first normalized. For positive indicators where larger values ​​indicate better evaluation results, such as renewable energy absorption rate, volatility support level, peak-to-valley ratio change rate, and maximum support volatility, maximum-minimum standardization is adopted, as shown in the following formula.

[0086]

[0087] For negative indicators where "the smaller the value, the better," such as network loss voltage deviation, inverse normalization is used, as shown in the following formula:

[0088] After obtaining all normalized indicators, a fixed-weighted composite score label is constructed. As the supervised learning objective, let the normalized index vector be denoted as . The commonly used flexible weights correspond to fixed weights. Then it can be calculated using the following formula.

[0089]

[0090] In this embodiment, step 3-1 can be implemented by the following method: Each internal node of the decision tree processes a certain feature. j Set threshold The splitting is performed using the minimum mean square error (MSE) as the splitting criterion:

[0091] Where m is the number of samples in the current node. and The number of samples in the left and right child nodes after the split. and This represents the minimum mean square error of the corresponding child node.

[0092] In this embodiment, step 3-2 can be implemented by the following method: Feature Importance (FI) in Random Forest is based on the sum of the impurity reductions brought by a feature across all decision trees:

[0093] in, Let b be the number of internal nodes of the b-th tree. Let be the sample weights of node t. The amount of impurity reduction at node t due to splitting based on feature j.

[0094] For regression trees, impurity is usually measured by MSE, and the contribution of feature j to the importance of node t is:

[0095] In this embodiment, step 3-3 can be implemented by the following method: Normalize the importance values ​​of all features to obtain the objective weights of each indicator:

[0096] Where n is the total number of features. This weight reflects the relative contribution of each indicator to the flexibility score.

[0097] In this embodiment, step 4-1 can be implemented by the following method: XGBoost employs an additive training strategy, adding a new tree at each iteration.

[0098] Solve by minimizing the following approximate objective function.

[0099] in and These are the first and second derivatives of the loss function, respectively.

[0100] In this embodiment, step 4-2 can be implemented by the following method: For a given tree structure q(x) (which maps the input to leaf nodes), the optimal leaf node weights are:

[0101] in, This represents the set of samples assigned to leaf node j.

[0102] At this point, the optimal value of the objective function is

[0103] In this embodiment, step 4-3 can be implemented by the following method: XGBoost uses a greedy algorithm to find the optimal split point, where the split gain of feature j at split point k is...

[0104] in, and This is the sample set of the left and right child nodes after the split. The sample set of the parent node.

[0105] In this embodiment, step 4-4 can be implemented by the following method: This invention introduces the Particle Swarm Optimization (PSO) algorithm to jointly search for key hyperparameters on top of XGBoost. XGBoost is an ensemble learning method based on gradient boosting decision trees, and its performance is highly dependent on the settings of several hyperparameters, such as the number of base learners. Maximum tree depth Learning rate Subsample sampling rate and feature column sampling rate Traditional grid search or random search is inefficient in high-dimensional parameter spaces, prone to getting trapped in local optima, and computationally expensive. PSO has strong global optimization capabilities and fast convergence speed, making it suitable for optimization problems in continuous parameter spaces.

[0106] In this embodiment, step 5-1 can be implemented by the following method: The individual extreme value (pbest) and the group extreme value (gbest) in the velocity and position update formula are averaged to guide the particle to move towards a more balanced central position, thereby enhancing the global search capability. The specific formula is as follows;

[0107]

[0108] in, and It is a learning factor. and It is a random number in the range (0,1). For the first The historical best position found by an individual particle is called the individual optimal solution. The historical best position found for the entire particle swarm is called the swarm optimal solution. For the first During the first iteration The position of each particle; For inertial weights: In this embodiment, step 5-2 can be implemented by the following method: The specific formula is as follows:

[0109] in, and The upper and lower limits of the inertia weight are set to 0.9 and 0.2, respectively.

[0110] This allows the particles to maintain their exploratory nature in the early stages of the search process while achieving robust convergence in the later stages, significantly improving their performance and stability in XGBoost hyperparameter optimization.

[0111] Combination Figure 2This invention proposes a multi-dimensional index-based distribution system flexibility modeling method integrating random forest and particle swarm optimization (XGBoost). It constructs a two-dimensional flexibility evaluation system comprising six key indicators, systematically characterizing the operational adjustment capability and stability of the distribution system under multi-source disturbances. Feature weights are extracted using the random forest algorithm, improving the objectivity and robustness of the evaluation model. Combined with the XGBoost modeling framework incorporating an improved particle swarm optimization algorithm, global and efficient optimization of model hyperparameters is achieved, effectively improving prediction accuracy and generalization ability. Furthermore, typical daily clustering, normalization processing, and a fusion scoring mechanism enhance the model's ability to adapt to various operating conditions. Experimental results show that the proposed method outperforms traditional methods in terms of flexibility prediction accuracy, model stability, and feature interpretability, validating its application value in intelligent evaluation and flexibility quantification of distribution systems.

[0112] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.

[0113] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.

[0114] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the mobile source emission prediction methods based on time-series feature migration in the above embodiments.

[0115] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.

[0116] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0117] It should be noted that in this invention, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0118] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0119] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-dimensional index-based power distribution system modeling method integrating random forest and particle swarm optimization algorithms, characterized in that, Includes the following steps: Step 1: Construct a flexibility index evaluation system, which includes a flexibility evaluation framework with six key indicators; This framework comprehensively considers multiple dimensions, including renewable energy acceptance capacity, fluctuation response level, voltage stability, network transmission efficiency, and power fluctuation tolerance. Step 2: Collect typical operation scenario data; Based on historical wind power, photovoltaic and load data, use K-means clustering method to extract several representative typical daily scenarios to enhance the diversity and representativeness of model training data, and calculate the occurrence probability weight of each scenario; Perform indicator normalization and label construction; To address the issues of inconsistent measurement units and evaluation directions in flexibility indicators, forward normalization and reverse normalization methods are used to preprocess the data, and a comprehensive scoring label is constructed using a fixed weighting method as the model learning objective. Step 3: Extract the importance of indicator features; The random forest regression model is used to calculate the relative contribution of each indicator to the flexibility score, generate normalized feature weights, overcome the bias problem of subjective weighting, and realize a data-driven weight evaluation mechanism. Step 4: Construct and optimize the XGBoost regression model; use the key features selected by random forest as input to build an XGBoost flexible prediction model; introduce an improved particle swarm optimization (PSO) algorithm to jointly tune the model hyperparameters, thereby improving the model's prediction accuracy and generalization ability. Step 5: Design an improved particle swarm optimization strategy; enhance global search capability by using exponentially decaying inertial weights, maintain population stability by combining the average update mechanism of individual extreme values ​​and population extreme values, and embed a local perturbation mechanism to avoid getting trapped in local optima. Step 6: Complete model training and performance evaluation; evaluate model performance on the training and test sets respectively, and verify the effectiveness and superiority of the proposed modeling method by comparing and verifying the mean squared error, scoring error and ROC curve, and output the final flexibility evaluation results.

2. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 1, characterized in that, Step 1 includes the following steps: The renewable energy absorption rate refers to the proportion of renewable energy generation that is absorbed and consumed by the power system within a certain period, relative to its total power generation or total load. This proportion reflects the utilization rate of renewable energy in the overall power supply structure. The formula for calculating the renewable energy absorption rate is as follows: in, Indicates the renewable energy consumption rate. This represents the amount of renewable energy generated that is actually absorbed by the power grid. Indicates the maximum load power of the distribution network; The renewable energy fluctuation support level refers to the power system's ability to maintain stable grid operation when it accommodates a large amount of fluctuating renewable energy. This indicator reflects the power system's ability to cope with fluctuations in renewable energy generation, ensuring that the grid's frequency, stability, and load balance are not affected by severe fluctuations in renewable energy generation. The formula for calculating the renewable energy fluctuation support level is as follows: in: For time period The actual interaction efficiency of flexible resources This represents the total number of nodes connected to the flexible resources. and Time periods The output power of flexible resources and the output power of the previous period; and Time periods The output power of new energy sources and the output power of the previous period; The peak-to-valley ratio variation rate refers to the rate at which the peak-to-valley ratio in the electricity load curve changes over time, and is used to measure the degree of fluctuation in electricity load; the formula for calculating the peak-to-valley ratio variation rate is: Network loss refers to the power loss that occurs during power transmission and distribution due to factors such as resistance and wire inductance. The formula for calculating network loss is as follows: in This refers to grid losses, also known as power losses, measured in MW or KW. It is the current on the i-th line. It is the resistance on the i-th line. It represents the total number of transmission lines in the system; Average voltage deviation measures the degree of deviation of the voltage at each node of the system from the nominal voltage, reflecting the stability of voltage quality. The formula for calculating average voltage deviation is as follows: in This is the average voltage deviation, expressed as a percentage. It is the first The voltage of each node; It is the nominal voltage; It is the total number of nodes; The maximum permissible volatility is the rate at which the power grid can withstand the maximum load or the fluctuation of new energy power generation; it is used to measure the power grid's tolerance to rapid load changes or intermittent new energy power generation. in This is the maximum permissible volatility, expressed in MW / minute or MW / hour. It is the maximum power change that the power grid can withstand. It is a time interval.

3. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 1, characterized in that, Step 2 specifically includes: Step 2-1: Typical Scenario Daily Generation Before clustering, the data is preprocessed to remove missing or non-numerical data, resulting in cleaned load data. Subsequently, the K-means clustering algorithm was used to divide the load curves into K classes, each class... The central curve, i.e., the typical diary entry, is The calculation formula is as follows: ; in, This represents the set of all days that are classified into the i-th class. The load curve for day d; To measure the representativeness of each type of typical day, the probability of occurrence of each type, pi, is further calculated. Step 2-2: Midpoint-guided mutation rate update strategy.

4. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 3, characterized in that, Step 2-2, the midpoint-guided mutation rate update strategy, specifically includes: First, the original data is normalized. After obtaining all normalized indicators, a fixed-weighted composite score label is constructed. As the supervised learning objective; let the normalized index vector be . Using weights Then, it is calculated using the following formula. 。 5. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 1, characterized in that, Step 3 specifically involves: Step 3-1: Decision Tree Construction and Feature Splitting Each internal node of the decision tree processes a certain feature. j Set threshold The splitting is performed using the minimum mean square error (MSE) as the splitting criterion: Where m is the number of samples in the current node. and The number of samples in the left and right child nodes after splitting. and This represents the minimum mean square error of the corresponding child node; Step 3-2: Calculation of feature importance Feature Importance (FI) in Random Forest is based on the sum of the impurity reductions brought by a feature across all decision trees: in, Let b be the number of internal nodes of the b-th tree. Let be the sample weights of node t. The amount of impurity reduction resulting from splitting node t based on feature j; For a regression tree, impurity is measured by MSE, and the contribution of feature j to the importance of node t is: Step 3-3: Normalize the importance values ​​of all features to obtain the objective weights of each indicator.

6. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 5, characterized in that, Step 3-3 specifically includes, Normalize the importance values ​​of all features to obtain the objective weights of each indicator: Where n is the total number of features; this weight reflects the relative contribution of each indicator to the flexibility score.

7. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 1, characterized in that, The optimized XGBoost regression model in step 4 is as follows: Step 4-1: Gradient Boosting Optimization XGBoost employs an additive training strategy, adding a new tree at each iteration. Solve by minimizing the following approximate objective function. in and These are the first and second derivatives of the loss function, respectively. Step 4-2, Tree Structure Optimization For a given tree structure q(x), which maps the input to leaf nodes, the optimal leaf node weights are: in, This represents the set of samples assigned to leaf node j; At this point, the optimal value of the objective function is Step 4-3: Feature Splitting and Gain Calculation XGBoost uses a greedy algorithm to find the optimal split point, where the split gain of feature j at split point k is... in, and The sample set of the left and right child nodes after the split. The parent node's sample set; Step 4-4: Improve the PSO optimization algorithm.

8. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 7, characterized in that, Step 4-4 specifically includes, Based on XGBoost, the Particle Swarm Optimization (PSO) algorithm is introduced to jointly search its key hyperparameters.

9. The multi-dimensional index power distribution system modeling method integrating random forest and particle swarm optimization algorithms according to claim 1, characterized in that, The improved particle swarm optimization strategy in step 5 is as follows: Step 5-1: Speed ​​and Position Midpoint Update The individual extreme value (pbest) and the group extreme value (gbest) in the velocity and position update formula are averaged to guide the particle to move towards a more balanced central position, thereby enhancing the global search capability. The specific formula is as follows; in, and It is a learning factor. and It is a random number in the range (0,1). For the first The historical best position found by an individual particle is called the individual optimal solution. The historical best position found for the entire particle swarm is called the swarm optimal solution. For the first During the first iteration The position of each particle; Inertial weights: Step 5-2, Exponentially Decreasing Dynamic Inertia Weights The specific formula is as follows: in, and These are the upper and lower limits of the inertia weight, respectively.