Photovoltaic simulation prediction method and system based on cellular automaton
By integrating multi-source data and machine learning models, combined with interpretability analysis, a suitability map is generated and a cellular automata model is applied. This solves the problems of multi-source data fusion and model interpretability in existing photovoltaic simulation and prediction technologies, and enables accurate prediction of the spatial expansion and loss of photovoltaic facilities, providing a scientific basis for photovoltaic planning.
Patent Information
- Application Number
- CN202610220754.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-24
- Publication Date
- 2026-08-04
AI Technical Summary
Existing photovoltaic simulation and prediction technologies suffer from weak multi-source data fusion capabilities, insufficient characterization of nonlinear mechanisms, poor model interpretability, low spatial accuracy, and a lack of dynamic simulation and policy interaction capabilities. As a result, the prediction results deviate significantly from the actual photovoltaic layout, making it difficult to meet the accuracy and practicality requirements of regional photovoltaic planning.
By integrating multi-source data from remote sensing, meteorology, topography, and socioeconomics, a unified spatial resolution database of photovoltaic distribution and environmental factors is constructed. A machine learning model is used to build an optimal classification model for photovoltaic expansion and loss. Combined with interpretability analysis, marginal contribution features of factors are extracted to generate a suitability map. Based on this, a parameterized cellular automata model is used to predict and simulate the spatial expansion and loss of photovoltaic facilities.
It enables accurate prediction of the spatial expansion and loss of photovoltaic facilities, provides accurate and reliable decision-making basis for regional photovoltaic layout optimization and policy formulation, improves the multi-source data fusion capability and model transparency, and dynamically restores the spatial evolution process of photovoltaics.
Smart Images

Figure CN122508949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of energy information and geographic information system technology, and in particular to a photovoltaic simulation and prediction method and system based on cellular automata. Background Technology
[0002] With the advancement of global "dual-carbon" goals and the acceleration of the clean energy transition, photovoltaic (PV) power generation, as a highly promising renewable energy source, has achieved rapid growth in installed capacity and application scale due to its advantages of being clean, low-carbon, and abundant in resources, becoming one of the core forces in optimizing the global energy structure. However, the layout of PV power generation is comprehensively influenced by multiple factors, including the natural environment, socio-economic conditions, and policy guidance. Furthermore, the expansion and withdrawal of PV facilities exhibit significant spatial heterogeneity and dynamic evolution characteristics. Accurately grasping its spatial development patterns is crucial for optimizing regional energy planning, improving PV utilization efficiency, and ensuring the stable operation of the power system. Therefore, developing scientific and efficient simulation and prediction technologies for the spatial expansion and losses of PV facilities to provide data support and theoretical basis for PV industry planning decisions has become an important research direction in the current energy field.
[0003] In existing technologies, photovoltaic-related simulation and prediction techniques mainly include physical modeling, statistical modeling, traditional machine learning models, and early cellular automata (CA) models. Among them, physical modeling is based on physical parameters such as solar radiation and module efficiency, which has low data dependence but ignores dynamic influencing factors; statistical modeling uses historical data to build time series models, which are suitable for short-term stable scenarios but are difficult to handle nonlinear and multivariate coupling problems; traditional machine learning models can capture some nonlinear relationships, but require manual feature extraction, and the black-box nature of the models leads to insufficient interpretability of the driving mechanism; early CA models in photovoltaic simulation often use subjectively set transformation rules, lack systematic integration of multi-source driving factors, and do not specifically distinguish the bidirectional evolution logic of photovoltaic expansion and loss.
[0004] Meanwhile, the above-mentioned technical solutions generally suffer from problems such as weak multi-source data fusion capabilities, insufficient accuracy in quantifying the contribution of photovoltaic change driving factors, and lack of dynamic constraints and traceability in the simulation process. As a result, the prediction results often deviate significantly from the actual photovoltaic layout, making it difficult to meet the high requirements of regional photovoltaic planning for accuracy, scientificity, and practicality.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] To address the aforementioned issues, this application provides a photovoltaic simulation and prediction method and system based on cellular automata, which can perform accurate bidirectional dynamic simulation of photovoltaic space through cellular automata and supports interactive parameter adjustment to adapt to diverse planning scenarios.
[0007] To achieve the objectives of this application, the following technical solution is provided: In a first aspect, this application provides a photovoltaic simulation and prediction method based on cellular automata, comprising: Multi-source data from remote sensing, meteorology, topography, and socioeconomic sources are cleaned, spatially registered, interpolated, supervised classification, and rasterized to construct a database of photovoltaic distribution data and environmental factors at a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data. Based on the photovoltaic distribution data, expansion, loss, and maintenance samples are constructed. An optimal classification model for photovoltaic expansion and photovoltaic loss is constructed through a machine learning model. The contribution results of each influencing factor and its marginal effect characteristics are obtained through interpretability analysis. Based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, the factor marginal contribution characteristic curve is extracted, the comprehensive suitability score of photovoltaic expansion and photovoltaic loss is calculated, and the photovoltaic expansion suitability map and photovoltaic loss suitability map are generated through spatial mapping. Using the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton, after determining the operating rules and parameter constraint set of the cellular automaton model, the candidate cell set is screened, the number of transformations is controlled, and the state of each cell is iteratively updated step by step, thereby predicting and simulating the process of photovoltaic facility spatial expansion and loss; the interactive parameters include neighborhood weight, suitability weight, and target photovoltaic area.
[0008] Secondly, this application also provides a photovoltaic simulation and prediction system based on cellular automata, used to execute the above-described photovoltaic simulation and prediction method based on cellular automata, the system comprising: The data processing module is used to clean, spatially register, interpolate, supervised classify, and rasterize multi-source data from remote sensing, meteorology, topography, and socioeconomic sources, and to construct a database of photovoltaic distribution data and environmental factors at a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data. The model building module is used to construct expansion, loss, and maintenance samples based on the photovoltaic distribution data, construct the optimal classification model of photovoltaic expansion and photovoltaic loss through machine learning model, and obtain the contribution results and marginal effect characteristics of each influencing factor through interpretability analysis. The suitability evaluation module is used to extract the factor marginal contribution feature curve based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, calculate the comprehensive suitability score of photovoltaic expansion and photovoltaic loss, and generate photovoltaic expansion suitability map and photovoltaic loss suitability map through spatial mapping. The photovoltaic prediction module is used to take the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and the photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton. After determining the operating rules and parameter constraint set of the cellular automaton model, it filters the candidate cell set, controls the number of transformations, and iteratively updates the state of each cell step by step, thereby predicting and simulating the spatial expansion and loss process of photovoltaic facilities. The interactive parameters include neighborhood weight, suitability weight, and target photovoltaic area.
[0009] The technical solution provided in this application may include the following beneficial effects: The photovoltaic simulation and prediction method and system based on cellular automata provided in this application can solve the problems of weak multi-source data fusion, insufficient characterization of nonlinear mechanisms, poor model interpretability, low spatial accuracy, and lack of dynamic simulation and policy interaction capabilities in existing technologies. It can achieve accurate prediction of the spatial expansion and loss of photovoltaic facilities, and provide accurate and reliable decision-making basis for regional photovoltaic layout optimization, policy formulation and risk prevention and control.
[0010] Specifically, by integrating and standardizing multi-source data from remote sensing, meteorology, topography, and socioeconomic sources, a unified spatial resolution database of photovoltaic distribution and environmental factors can be constructed. This achieves comprehensive coverage of the driving factors of photovoltaic changes, laying an accurate and reliable data foundation for prediction and simulation. Furthermore, by combining machine learning classification models with interpretability analysis, the contribution and marginal effect of each environmental factor to photovoltaic expansion and loss can be accurately quantified, breaking the black-box limitations of traditional prediction models and making the driving logic of photovoltaic spatial evolution more transparent and traceable. At the same time, based on the marginal contribution characteristics of factors, a comprehensive suitability score is calculated and a spatial suitability map is generated, providing scientific evolutionary constraints for cellular automata simulation. Through bidirectional state transition rules, neighborhood effects, and an interactive parameter system, the photovoltaic spatial evolution process can be dynamically restored, adapting to diverse planning scenarios.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Obviously, the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0013] Figure 1 A schematic flowchart illustrating a photovoltaic simulation and prediction method based on cellular automata provided in this application embodiment; Figure 2 A flowchart illustrating step S100 of a photovoltaic simulation prediction method based on cellular automata provided in this application embodiment; Figure 3 A flowchart illustrating step S200 of a photovoltaic simulation and prediction method based on cellular automata provided in this application embodiment; Figure 4 A flowchart illustrating step S300 of a photovoltaic simulation and prediction method based on cellular automata provided in this application embodiment; Figure 5 A flowchart illustrating step S400 of a photovoltaic simulation prediction method based on cellular automata provided in this application embodiment; Figure 6 A schematic diagram illustrating the technical framework of a photovoltaic simulation and prediction method based on cellular automata provided in this application embodiment; Figure 7 A photovoltaic simulation and prediction method based on cellular automata provided in this application provides a marginal effect curve of photovoltaic expansion; Figure 8 This is a schematic diagram of a photovoltaic simulation and prediction system based on cellular automata, provided as an embodiment of this application. Detailed Implementation
[0014] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0015] This example implementation first provides a photovoltaic simulation and prediction method based on cellular automata. (Reference) Figure 1 As shown, the photovoltaic simulation and prediction method based on cellular automata may include the following steps: Step S100: Clean, spatially register, interpolate, supervised classification, and rasterize the multi-source data from remote sensing, meteorology, topography, and socioeconomic sources to construct a database of photovoltaic distribution data and environmental factors with a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data.
[0016] Step S200: Based on the photovoltaic distribution data, construct expansion, loss and maintenance samples, construct the optimal classification model of photovoltaic expansion and photovoltaic loss through machine learning model, and obtain the contribution results and marginal effect characteristics of each influencing factor through interpretability analysis.
[0017] Step S300: Based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, extract the factor marginal contribution characteristic curve, calculate the comprehensive suitability score of photovoltaic expansion and photovoltaic loss, and generate a photovoltaic expansion suitability map and a photovoltaic loss suitability map through spatial mapping.
[0018] Step S400: Taking the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton, after determining the operating rules and parameter constraint set of the cellular automaton model, the candidate cell set is screened, the number of transformations is controlled, and the state of each cell is iteratively updated step by step to predict and simulate the process of photovoltaic facility spatial expansion and loss; the interactive parameters include neighborhood weight, suitability weight, and target photovoltaic area.
[0019] Below, we will refer to Figures 2 to 8 The steps of the photovoltaic simulation prediction method based on cellular automata described in this example embodiment will be explained in more detail.
[0020] In step S100, the remote sensing, meteorological, topographic, and socioeconomic multi-source data are cleaned, spatially registered, interpolated, supervised classification, and rasterized to construct a photovoltaic distribution data and environmental factor database with a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data.
[0021] Understandably, remote sensing, meteorological, topographic, and socioeconomic data inherently suffer from heterogeneous formats, spatial resolution differences, and inconsistent coordinate systems. Directly using these data for analysis can easily lead to biased results. Therefore, cleaning, spatial registration, interpolation, and rasterization are necessary prerequisites for eliminating data heterogeneity and achieving effective fusion of multi-source data. Unifying spatial resolution ensures that photovoltaic distribution data and environmental factor data can be correlated and matched at the same spatial scale. Dividing environmental factor data into natural and social categories allows for comprehensive basic data support for subsequent analysis of the driving mechanisms of photovoltaic expansion and loss from different dimensions such as resource endowment and the impact of human activities.
[0022] In one possible implementation, step S100 may further include the following sub-steps: In step S110, based on multispectral remote sensing images with a spatial resolution of 10m, a supervised convolutional neural network is used to identify photovoltaic modules, generate binarized photovoltaic identification results, and obtain photovoltaic basic distribution data through vectorization and spatial simplification processing.
[0023] Understandably, multispectral remote sensing images with a spatial resolution of 10m possess sufficient detail to effectively distinguish photovoltaic modules from other ground features, providing a data foundation for accurate identification of photovoltaic facilities. Using supervised convolutional neural networks (CNNs) for identification can improve the accuracy and reliability of the identification results by leveraging labeled photovoltaic module samples. Generating binarized photovoltaic identification results can clearly define the spatial boundaries between photovoltaic and non-photovoltaic areas, while vectorization and spatial simplification processes reduce data redundancy while preserving the core spatial location information of photovoltaic facilities, which is beneficial for subsequent analysis and calculation of photovoltaic distribution changes over large-scale areas.
[0024] The expression for the binarized photovoltaic identification result is as follows:
[0025] in, , , This is the result of binary photovoltaic identification. For classification threshold, The predicted probability of photovoltaics appearing. For CNN-based non-linear classification functions, For pixels The multispectral eigenvectors, For the network parameters obtained during training, For pixels band The reflectivity value, This represents the number of spectral bands.
[0026] It should be noted that this binarized photovoltaic identification result is the core basis for the subsequent construction of photovoltaic basic distribution data. Through clear value division, it can directly define the photovoltaic / non-photovoltaic spatial attributes at the pixel level, providing a clear spatial boundary basis for subsequent analysis of the expansion, loss and other changes in photovoltaic areas.
[0027] In step S120, based on remote sensing images with a spatial resolution of 0.1°, the normalized vegetation index is extracted through band operations, the surface temperature is retrieved using the generalized radiative transfer formula, and a raster dataset of natural environmental factors is constructed by combining meteorological and topographic data.
[0028] It should be noted that the 0.1° spatial resolution remote sensing image balances large-area coverage with the basic accuracy of environmental factor inversion, and is suitable for the needs of photovoltaic natural environment analysis at the regional scale. Band operation and generalized radiative transfer formula are mature land parameter inversion methods in the field of remote sensing, which can transform the spectral information of remote sensing image into NDVI and surface temperature data with actual ecological and physical significance. The combination of meteorological and topographic data is to supplement the deficiencies of single remote sensing data in terms of temporal continuity and micro-topographic details, thereby constructing a complete raster dataset of natural environmental factors covering dimensions such as resource endowment and topographic conditions.
[0029] The formula for calculating the normalized vegetation index is as follows:
[0030] in, For pixels Normalized Difference Vegetation Index (NDVI) For pixels Near-infrared reflectance, For pixels Reflectivity in the red light band; The formula for calculating the land surface temperature is:
[0031] in, For surface temperature, For pixels Brightness temperature, The effective wavelength in the thermal infrared band, Let be a constant derived from Planck's law. is the surface emissivity.
[0032] Understandably, the NDVI calculation formula, by using the ratio of the difference between near-infrared and red light reflectance to the sum of values, can accurately distinguish the degree of vegetation cover. The closer the value is to 1, the higher the vegetation cover, and the vegetation cover directly determines whether an area is suitable for the deployment of photovoltaic facilities. The surface temperature calculation formula, based on the principle of radiative transfer, converts the brightness temperature received by the sensor into the actual surface temperature. This parameter is highly correlated with the heat dissipation efficiency and photoelectric conversion efficiency of photovoltaic modules. The application of these two quantitative formulas ensures the accuracy of the calculation of natural environmental factors and provides reliable parameter support for the subsequent analysis of the natural driving mechanism of photovoltaic changes.
[0033] In step S130, the photovoltaic facilities and various vector data are spatially cleaned and registered on the GIS platform. Slope and aspect are calculated based on the digital elevation model. Kriging interpolation is used to convert discrete meteorological station data into continuous spatial meteorological data. Distance-type indicators are calculated using Euclidean distance. The spatial intensity of terminal energy consumption facilities is quantified using kernel density estimation. A social environmental factor dataset is constructed by combining land use data, GDP, population, and policy driving factors. The method for quantifying the spatial intensity of terminal energy consumption facilities is as follows: a suitable search radius is determined by kernel density values, and Ripley's K function is used to identify the dominant spatial clustering scale of terminal energy consumption facilities.
[0034] It should be noted that the GIS platform is a professional tool adapted to spatial processing of multi-source vector data. It performs spatial cleaning and registration on photovoltaic facilities and various vector data, which can eliminate positional deviations and topological errors between data, and ensure that photovoltaic facilities and social environment-related data are accurately aligned in spatial dimension. This is the basic premise for subsequent multi-factor spatial correlation analysis.
[0035] The calculation of slope and aspect based on digital elevation model is because terrain conditions are directly related to the construction difficulty, installation cost and light reception efficiency of photovoltaic facilities, and are the core basis for characterizing the terrain adaptability of photovoltaic layout. The use of Kriging interpolation to process discrete meteorological station data is to transform discontinuous station observation data into a continuous spatial meteorological field covering the entire region, so as to solve the limitation that discrete data cannot support regional scale analysis.
[0036] Distance-based indicators calculated using Euclidean distance can quantify the locational accessibility of photovoltaic facilities to infrastructure such as transportation and power grids, while kernel density estimation can accurately capture the spatial agglomeration characteristics of energy consumption demand by characterizing the spatial intensity of end-user energy consumption facilities. Together, these two constitute the core measurement dimensions of the locational advantages of photovoltaic layout. Combining land use, GDP, population, and policy factors, this approach comprehensively covers the social environmental impact dimensions of photovoltaic development from the perspectives of land use constraints, economic support, demand scale, and policy guidance, making the social environmental factor dataset more closely aligned with the actual driving logic of photovoltaic layout.
[0037] Identifying the dominant spatial agglomeration scale of end-use energy consumption facilities can accurately match the service range of photovoltaic facilities, avoid spatial mismatch between photovoltaic layout and energy consumption demand, and further enhance the pertinence and rationality of subsequent analysis of photovoltaic change driving mechanisms.
[0038] The formulas for calculating the terrain slope and aspect are as follows:
[0039]
[0040] in, For terrain slope, Due to the slope of the terrain, For elevation, For along Elevation gradient in direction, For along Elevation gradient in direction; The calculation formula for the continuous spatial meteorological data is as follows:
[0041] in, , For pixels Interpolated meteorological variable values For weather station The observed values, For pixels Assigned to weather stations Kriging weights, Represents the semi-mutation function. For weather station With weather station The distance between them For pixels With weather station The distance between them For Lagrange multipliers used to apply unbiased constraints, The number of sites participating in the interpolation; The Euclidean distance calculation formula is as follows:
[0042] in, For pixels Euclidean distance to the nearest target element For pixels plane coordinates, as elements Plane coordinates; The expression for the nuclear density value is:
[0043] in, For pixels The kernel density value at that location, Represents the kernel function. For pixels With facilities The Euclidean distance between them The KDE search radius is determined by Ripley's K function; The expression for Ripley'sK function is:
[0044] in, The value of Ripley's K function calculated at the distance. The total area of the target region. For the number of facilities, This is an indicator function that takes the value 1 when the condition is met and 0 otherwise.
[0045] In step S140, all raster layers and vector layers in the photovoltaic basic distribution data, natural environmental factor raster dataset, and social environmental factor dataset are uniformly projected onto the same coordinate reference system and resampled according to a preset target resolution to obtain photovoltaic distribution data and environmental factor database under a unified spatial resolution.
[0046] It should be noted that photovoltaic (PV) distribution data, natural environmental factor raster data, and social environmental factor data from different sources often suffer from inconsistent coordinate reference systems and heterogeneous spatial resolutions. Directly using these for subsequent analysis can easily lead to spatial misalignment and scale mismatch. Unifying the projection to the same coordinate reference system eliminates spatial offsets between data points, ensuring precise alignment of various data types in spatial dimensions. Resampling at a preset target resolution ensures that all raster and vector layers are adapted to the same spatial scale, balancing the accuracy requirements of subsequent analysis with avoiding errors in correlation calculations caused by scale differences. This unified PV distribution data and environmental factor database is the core data carrier for constructing subsequent PV change models, ensuring that all factors and PV distribution can be correlated and modeled within the same spatial framework.
[0047] In step S200, expansion, loss, and maintenance samples are constructed based on the photovoltaic distribution data. An optimal classification model for photovoltaic expansion and photovoltaic loss is constructed through a machine learning model. The contribution results of each influencing factor and its marginal effect characteristics are obtained through interpretability analysis.
[0048] Understandably, dividing the samples into photovoltaic expansion, loss, and maintenance categories is to specifically differentiate different types of photovoltaic spatial evolution, providing clear labeling criteria for subsequent classification models. Only by covering samples from different change scenarios can the model accurately learn the correlation between various environmental factors and photovoltaic expansion and loss. Using machine learning to build classification models can capture the nonlinear, multi-dimensional interaction between factors and photovoltaic changes, making it more suitable for the complex driving logic of real-world scenarios compared to traditional linear analysis methods. The core value of interpretability analysis lies in breaking through the black-box limitations of the model and clarifying the contribution and characteristics of each factor to photovoltaic changes. This result is a key basis for subsequent photovoltaic suitability evaluation. Furthermore, the diversity and representativeness of the samples directly affect the model's generalization ability and the reliability of the interpretation results.
[0049] In one possible implementation, step S200 may further include the following sub-steps: In step S210, based on the photovoltaic distribution data of different periods, the photovoltaic changes in adjacent time periods are geographically calculated, the photovoltaic change types in the target area are classified, and based on the combination of photovoltaic distribution data of previous and subsequent time periods, a unique photovoltaic change code is assigned to each cell to obtain the dependent variable dataset; the photovoltaic change types include photovoltaic expansion and photovoltaic loss.
[0050] It should be noted that the selection of adjacent time periods needs to match the actual construction and dismantling cycle characteristics of photovoltaic facilities to avoid the change signal being insignificant due to excessively short time intervals, or the interference factors from multiple periods being mixed due to excessively long time intervals. This ensures that the geographic calculation results can accurately reflect the real process of photovoltaic spatial evolution. The unique photovoltaic change code for each pixel is to transform the combination of photovoltaic states in previous and subsequent time periods into a quantitative dependent variable label. This process allows different change types such as photovoltaic expansion and loss to be clearly identified by the model, which is the core label foundation for the subsequent construction of supervised classification models. At the same time, the uniqueness of the code also ensures that the change state of each pixel will not be confused.
[0051] The photovoltaic change coding formula is as follows:
[0052] in, For pixels Photovoltaic change coding, This refers to the photovoltaic status in the previous period. The photovoltaic status in the later period; The corresponding photovoltaic change state explanation summary can be expressed as follows: .
[0053] It should be noted that this encoding method can intuitively and uniquely identify the photovoltaic change state of each pixel through simple numerical combination rules: for example, when there is no photovoltaic in the previous period and there is photovoltaic in the later period, the encoding result corresponds to the expansion type; when there is photovoltaic in the previous period and there is no photovoltaic in the later period, the encoding result corresponds to the loss type. This encoding form is concise and easy to understand, and can transform qualitative change types into numerical labels that are suitable for machine learning models. At the same time, it facilitates the classification statistics of subsequent samples and the label matching in the model training process.
[0054] In step S220, based on the natural environmental factor data and the social environmental factor data, a bivariate statistical analysis is performed to eliminate redundant and weakly correlated indicators, and to select significantly correlated environmental and socioeconomic indicators to construct an independent variable dataset.
[0055] Understandably, the zonal bivariate analysis targets different types of photovoltaic (PV) changes, such as expansion and losses, separately analyzing the distribution characteristics of various environmental and socioeconomic factors. This allows for the precise capture of factor correlation patterns corresponding to different change types, avoiding the obscuring of significant local impacts by global analysis. Eliminating redundant and weakly correlated indicators reduces the computational complexity and improves efficiency of subsequent models, while also mitigating the interference of multicollinearity on model stability. The significantly correlated indicators obtained through this selection constitute the core independent variables characterizing the driving mechanism of PV changes, providing concise and effective data support for the subsequent model to accurately learn the correlation between factors and PV changes.
[0056] Furthermore, to reduce variable redundancy and eliminate weakly correlated predictors before model training, bivariate analysis can be performed based on partition statistics. For each candidate driving factor... Compare its statistical characteristics in different photovoltaic change categories, photovoltaic change categories Internal driving factors The mean is defined as: ;in, As a factor In category The mean of the values; For pixels Factor The value of ; For category The number of pixels.
[0057] In step S230, a stratified sampling strategy is adopted to construct a training set of positive and negative samples for photovoltaic expansion and photovoltaic loss based on the dependent variable dataset and the independent variable dataset. The photovoltaic expansion and photovoltaic loss classification model is trained, and the model performance is evaluated and the model is screened by using the receiver operating characteristic curve, area under the curve, precision and recall. Bayesian hyperparameter optimization is then performed to obtain the optimal classification model for photovoltaic expansion and photovoltaic loss. The photovoltaic expansion and photovoltaic loss classification model includes any one or more of the following: Logistic regression model, K-nearest neighbor model, random forest model, extreme gradient boosting model, lightweight gradient boosting machine model, and class boosting model.
[0058] It should be noted that, in order to depict the photovoltaic expansion process, the following conditions will be met. Pixels of the specified categories are labeled as positive samples, while the remaining categories are considered negative samples. To characterize the photovoltaic loss process, the following conditions will be met: Pixels of the specified classes are labeled as positive samples, and the remaining classes are considered negative samples. Let... For binary photovoltaic expansion label: .
[0059] The training dataset is constructed using a stratified sampling strategy:
[0060] in, These are environmental feature vectors (such as NDVI, DEM, RHU, SSD, Slope, etc.). This represents the set of pixels obtained from sampling.
[0061] Multiple classification models are trained to estimate the probability of photovoltaic expansion / loss. For any given model... Its general predictive relationship can be expressed as:
[0062] Logistic Regression (LR) model: ; K-Nearest Neighbors (KNN) model: ; Random Forest (RF) model: ; Extreme Gradient Boosting (XGBoost): ; Lightweight Gradient Boosting Machine (LightGBM): ; Category Boosting (CatBoost): ; in, For pixels Predicted probability of photovoltaic expansion; For the LogisticSigmoid function; Let be the coefficient vector of LR; express A set of nearest neighbors; and As a base learner; The number of trees; This is the learning rate.
[0063] The classification performance of each model was evaluated using receiver operating characteristic (ROC) curves, and the area under the curve (AUC) was used.
[0064] in, For the model Output photovoltaic expansion prediction probability; For the model The classification function is TPR (True Positive Rate) and FPR (False Positive Rate).
[0065] Based on AUC, Precision, and Recall, the best-performing model structure is selected and further optimized. Bayesian hyperparameter optimization is applied to the selected model, with the optimization objective defined as follows:
[0066] in, For hyperparameter vectors; This is the optimal set of hyperparameters.
[0067] In step S240, the interpretability analysis of the optimal classification model of photovoltaic expansion and photovoltaic loss is performed based on the SHAP method. The marginal contribution of each driving factor in the independent variable dataset to the predicted probability of photovoltaic change is quantified, and the contribution results and marginal effect characteristics of each influencing factor are obtained.
[0068] Understandably, choosing the SHAP method for interpretability analysis not only quantifies the contribution ranking of each driving factor to the predicted probability of photovoltaic changes, clarifying which factors are core drivers and which have secondary influences, but also characterizes the marginal effects of factors in different value ranges, such as... Figure 7 The figures shown are several indicators that exhibit significant threshold effects and directional inflection points, and deviate significantly from the linear assumption.
[0069] The marginal contribution expression of the driving factor to the photovoltaic change prediction probability is as follows:
[0070] in, As a factor pixels The contribution of the predicted results; It is the set of all driving factors.
[0071] Understandably, this expression, by iterating through subsets of all driving factors, can fairly quantify the independent marginal contribution of a single factor to the pixel prediction result, avoiding interference from inter-factor interactions on the contribution of a single factor, and ensuring that the effect of each driving factor is accurately decomposed. Furthermore, this expression, by iterating through subsets of all driving factors, can fairly quantify the independent marginal contribution of a single factor to the pixel prediction result, avoiding interference from inter-factor interactions on the contribution of a single factor, and ensuring that the effect of each driving factor is accurately decomposed. The sign and magnitude of the value can directly correspond to the direction and intensity of the promoting / inhibiting effect of the factor on photovoltaic changes. This result provides a quantitative basis for distinguishing the differences in the effects of different factors on photovoltaic expansion and loss, and is a key link between model interpretation and suitability evaluation.
[0072] In step S300, based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, the factor marginal contribution feature curve is extracted, the comprehensive suitability score of photovoltaic expansion and photovoltaic loss is calculated, and a photovoltaic expansion suitability map and a photovoltaic loss suitability map are generated through spatial mapping.
[0073] It should be noted that extracting the marginal contribution features of factors is the core basis for transforming the interpretability results of the photovoltaic expansion and loss classification model into suitability score calculations. Different factors have different strengths and directions of influence on expansion and loss; therefore, it is necessary to calculate the comprehensive suitability score of both separately to avoid conflation and deviation from the actual driving logic. The suitability map generated through spatial mapping not only visually presents the suitable spatial distribution of photovoltaic expansion and loss within the region but also provides precise spatial constraints for subsequent cellular automata simulations of photovoltaic spatial evolution, serving as a key visualization carrier connecting model analysis and planning applications.
[0074] In one possible implementation, step S300 may further include the following sub-steps: In step S310, based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, the positive and negative sample training sets of photovoltaic expansion and photovoltaic loss are processed in a hierarchical manner, the marginal contribution feature curves of each influencing factor are extracted, the positive marginal contribution curves are retained by the ReLU filtering mechanism, the positive marginal contribution curves are approximated by the smoothing function, and the contribution score lookup table of each factor is constructed.
[0075] It's important to note that the stratified processing of the positive and negative sample training sets is designed to accurately extract the marginal contribution features of each factor under different scenarios, targeting different sample types related to photovoltaic expansion and loss, thus avoiding interference between contribution patterns of different change types. The ReLU filtering mechanism retains positive marginal contributions because suitability evaluation focuses on the positive factors that "promote photovoltaic expansion" and "drive photovoltaic loss." Removing negative factors eliminates irrelevant interference, making the contribution features more aligned with the core evaluation logic of suitability. Fitting the positive marginal contribution function with a smoothing function transforms discrete contribution data into a continuous response relationship, preventing jumps in contribution scores caused by small fluctuations in factor values and improving the stability of the score. Constructing a factor contribution score lookup table transforms complex marginal contribution patterns into a standardized and quickly accessible tool. Subsequent suitability score calculations can directly match the corresponding contribution score based on the factor value, improving computational efficiency and ensuring consistency of the scoring results.
[0076] The formula for calculating the positive marginal contribution is as follows:
[0077] in, , As a factor Positive marginal contribution after ReLU filtering As a factor The marginal SHAP contribution function, For pixels Factor SHAP value, For pixels Factor The value of ; The fitted positive marginal contribution function is:
[0078] in, This represents the fitted marginal contribution value; As a factor The fitted response function.
[0079] It should be noted that by retaining positive marginal contributions through the ReLU filtering mechanism, we can focus on the factor effects that promote photovoltaic expansion or loss, eliminate negative interference, and ensure that subsequent suitability scoring revolves only around favorable drivers. Fitting the positive marginal contribution function transforms discrete factor contribution data into a continuous response relationship, avoiding jumps in contribution scores due to small fluctuations in factor values. This improves the stability and universality of contribution scores and provides a continuous and reliable basis for constructing a standardized contribution score lookup table.
[0080] In step S320, the marginal contribution values obtained through the contribution score lookup table of each factor are linearly normalized to an integer range of 1-255, aggregated using geographical weights, and the comprehensive suitability score of photovoltaic expansion and photovoltaic loss is calculated.
[0081] It should be noted that linearly normalizing the marginal contribution value to an integer range of 1-255 is to eliminate the scale differences in the original contribution values of different factors, allowing the contribution of each factor to be aggregated at the same order of magnitude, and preventing a single factor from dominating the scoring result due to its large value range. Using geographical weights for aggregation takes into account the spatial heterogeneity of the importance of each influencing factor. By allocating weights that are adapted to regional characteristics, the comprehensive suitability score is made more consistent with the actual driving logic of different regions. The final calculated comprehensive suitability score can directly quantify the degree of suitability of a region for photovoltaic expansion or loss, and is the core quantitative basis for generating the photovoltaic expansion and loss suitability map.
[0082] The normalization rule formula is as follows:
[0083] in, For pixels Factor The normalized suitability score; This is the floor operator; The formula for calculating the overall suitability score of photovoltaic expansion and photovoltaic loss is as follows:
[0084] in, , For pixels The overall photovoltaic suitability score; As a factor Geographic weight; This represents the number of driving factors.
[0085] It should be noted that this normalization rule transforms the marginal contribution of factors into integer scores in the range of 1-255. This eliminates the scale difference of contribution values of different factors, making the suitability scores of each factor comparable horizontally. The rounding down operator simplifies the calculation while preserving the spatial distinguishability of the scores, ensuring that the suitability differences between pixels can be effectively identified.
[0086] The core of the comprehensive suitability score calculation method is to achieve a reasonable aggregation of the contributions of each factor through geographical weighting. This weighting is determined based on the absolute value proportion of the factor's SHAP value, and is not subjectively set. Instead, it directly inherits the results of the previous interpretability analysis, which can objectively reflect the actual contribution intensity of each factor to photovoltaic changes, making the comprehensive score more in line with the real driving logic of regional photovoltaic evolution. The pixel-level score results accurately correspond to spatial location, providing pixel-level quantitative support for the subsequent generation of photovoltaic expansion and loss suitability maps.
[0087] In step S330, the comprehensive suitability score is spatially mapped on the GIS platform to generate a photovoltaic expansion suitability map and a photovoltaic loss suitability map, respectively; the spatial resolution of the suitability map is 50m.
[0088] Understandably, the GIS platform possesses mature spatial data visualization and mapping capabilities, efficiently transforming pixel-level comprehensive suitability scores into intuitive spatial distribution layers. Choosing a 50m spatial resolution not only suits the spatial characteristics of the dispersed layout of photovoltaic (PV) facilities, accurately distinguishing the location and extent of individual PV facilities, but also avoids data redundancy and computational pressure caused by excessively high resolution, ensuring the efficiency of subsequent cellular automata simulations. Separate PV expansion and loss suitability maps are generated, clearly distinguishing the spatial areas suitable for new PV installations and those prone to PV withdrawal within the region, providing targeted spatial guidance for subsequent planning and simulation.
[0089] In step S400, the current photovoltaic distribution is taken as the initial state, the photovoltaic expansion suitability map and the photovoltaic loss suitability map are taken as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout are taken as the basic parameters of the cellular automaton. After determining the operating rules and parameter constraint set of the cellular automaton model, the candidate cell set is screened, the number of transformations is controlled, and the state of each cell is iteratively updated step by step to predict and simulate the process of photovoltaic facility spatial expansion and loss. The interactive parameters include neighborhood weight, suitability weight and target photovoltaic area.
[0090] It should be noted that using the current photovoltaic distribution as the initial state ensures that the simulation process starts from the actual photovoltaic layout in the region, preventing the simulation results from deviating from reality. Using the suitability map as an evolutionary constraint ensures that the state changes of cells conform to the driving laws obtained from the previous analysis, ensuring that the simulation is not a random evolution but conforms to the actual logic of photovoltaic changes in the region. Selecting candidate cells and controlling the number of conversions are to ensure that the simulation rhythm of photovoltaic changes matches the construction and demolition progress of actual projects, avoiding unreasonable situations of large-scale changes all at once. Gradual iterative updates can dynamically restore the temporal process of photovoltaic expansion and loss, and the final prediction simulation results can intuitively present the evolution trend of future photovoltaic spatial layout.
[0091] In one possible implementation, step S400 may include the following sub-steps: In step S410, the current photovoltaic distribution data in the photovoltaic distribution data is used as the initial state of the cellular automaton model, and the photovoltaic expansion suitability map and the photovoltaic loss suitability map are combined to obtain the initial input dataset of the cellular automaton model.
[0092] It should be noted that using the current photovoltaic distribution data as the initial state of the cellular automata provides a benchmark starting point that closely matches the actual regional situation, ensuring that the initial conditions of the simulation process are completely consistent with the current photovoltaic layout. Constructing the initial input dataset by combining the photovoltaic expansion and loss suitability map transforms the results of the preceding suitability analysis into the core constraints of the simulation, providing clear spatial rules to guide the subsequent state transitions of the cells. This initial input dataset integrates the current situation and evolutionary constraints, ensuring that the cellular automata simulation results are reasonable and closely reflect the regional characteristics, thus avoiding deviations in the simulation due to a lack of actual data support.
[0093] In step S420, a binary state transition rule for photovoltaic expansion and loss is set, an interactive parameter system is introduced to calculate the neighborhood photovoltaic density, and the comprehensive transition potential is calculated through suitability score, neighborhood photovoltaic density and random perturbation to obtain the operating rules and parameter constraint set of the cellular automata model.
[0094] It should be noted that the binary state transition rule clarifies the core change logic of the cell and is the basic framework for cellular automata to simulate the photovoltaic expansion and loss process. The introduction of an interactive parameter system allows users to adjust parameters according to different planning scenarios, improving the scenario adaptability and flexibility of the simulation. The calculation of neighborhood photovoltaic density is consistent with the actual spatial agglomeration characteristics of photovoltaic facilities. In reality, photovoltaics are often concentrated in the periphery of suitable areas, making the simulation results more consistent with the real spatial distribution pattern. The construction of comprehensive transfer potential integrates suitability, neighborhood effect and random perturbation. It not only relies on the suitability analysis of the preceding analysis to ensure the scientific nature of the simulation, but also takes into account the uncertainty of photovoltaic layout in reality through random perturbation, making the simulation results more practical reference value.
[0095] The formula for calculating the neighborhood photovoltaic density is as follows:
[0096] in, , For neighborhood photovoltaic density, For pixels In the Photovoltaic state in the next iteration. For pixels The neighborhood window is centered. The number of pixels in the neighborhood; The formula for calculating the comprehensive transfer potential is as follows:
[0097] in, In order to comprehensively transfer potential, For pixels The normalized suitability score, Neighborhood photovoltaic density; Let be a random variable that follows a uniform distribution. For suitability weight, For neighborhood weights, The weights are random, and .
[0098] It should be noted that the neighborhood photovoltaic density formula quantifies the spatial agglomeration of photovoltaic facilities by statistically analyzing the proportion of photovoltaic status in the area surrounding the central pixel. This effectively restores the characteristics of photovoltaic cluster layout in reality and avoids the occurrence of isolated and scattered unreasonable layouts in the simulation. The normalization of the number of pixels within the neighborhood window ensures that the density values under different neighborhood ranges are comparable horizontally and adapts to the simulation needs of different spatial scales.
[0099] The comprehensive transfer potential formula integrates three core factors—suitability, neighborhood effect, and random disturbance—through weight allocation: the weights α, β, and γ can be flexibly adjusted according to the actual characteristics of the region or planning objectives, allowing the simulation to adapt to diverse application needs; the random variable R(t) is used to simulate unforeseen disturbance factors in reality, avoiding overly idealized simulation results; and the constraint α+β+γ=1 ensures that the contributions of each factor are aggregated at the same order of magnitude, avoiding the excessive dominance of a single factor in the calculation of transfer potential, and ensuring the balance and rationality of the results.
[0100] In step S430, the set of candidate pixels that meet the threshold is determined by the cell suitability score. The number of conversions is controlled by setting the conversion intensity parameter for each iteration. Then, based on the operating rules and parameter constraint set of the cellular automata model, the state of each cell is iteratively updated step by step to dynamically restore the spatial change process of photovoltaic expansion and photovoltaic loss.
[0101] It should be noted that determining the candidate cell set based on cell suitability scores is to screen out cells with photovoltaic expansion or loss potential, avoiding the participation of cells without suitability support in state transitions, and fundamentally ensuring the rationality of the simulated changes. Setting a conversion intensity parameter to control the number of conversions ensures that the scale of photovoltaic changes in each iteration matches the actual engineering pace. In reality, photovoltaic projects are mostly implemented in batches and are not deployed on a large scale at once. This parameter makes the simulation results more consistent with the actual construction logic. Gradually iterating and updating the cell state can dynamically restore the temporal evolution process of photovoltaic expansion and loss, rather than directly outputting the final layout, so that the simulation presents both spatial distribution characteristics and the temporal progression logic of changes.
[0102] The candidate pixel set is as follows:
[0103] in, For the first The candidate cell set corresponding to the next iteration; The number of transformations in each iteration is:
[0104] in, The number of transformations in each iteration. To control the transformation intensity in each iteration.
[0105] It should be noted that this application implements two pixel selection strategies: one is deterministic selection based on transfer potential from high to low. ; Second, probability sampling is performed based on normalized transfer potential:
[0106] in, The set of cells selected to undergo state transition.
[0107] In step S440, when the iteration reaches the preset target photovoltaic area or the maximum number of iterations, the simulation process is terminated, and a photovoltaic spatial distribution simulation map at different time points is output to predict and simulate the spatial expansion and loss process of photovoltaic facilities.
[0108] It should be noted that after completing the preset number of iterations or reaching the target condition (target photovoltaic area), the simulation results of photovoltaic spatial layout at different stages are output to characterize the spatial distribution characteristics of future photovoltaic facility expansion and loss.
[0109] For in The selected cell's cell state is updated according to the simulation mode, and the update rule is as follows:
[0110] The iterative process continues until the target photovoltaic area is reached or the maximum number of iterations is reached, and its termination condition is as follows.
[0111]
[0112] in, Represents the area of a single pixel.
[0113] Furthermore, in this exemplary embodiment, a photovoltaic simulation and prediction system based on cellular automata is also provided for executing the aforementioned photovoltaic simulation and prediction method based on cellular automata. (Reference) Figure 8 As shown, the system may include...
[0114] The data processing module is used to clean, spatially register, interpolate, supervised classify, and rasterize multi-source data from remote sensing, meteorology, topography, and socioeconomic sources, and to construct a database of photovoltaic distribution data and environmental factors at a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data. The model building module is used to construct expansion, loss, and maintenance samples based on the photovoltaic distribution data, construct the optimal classification model of photovoltaic expansion and photovoltaic loss through machine learning model, and obtain the contribution results and marginal effect characteristics of each influencing factor through interpretability analysis. The suitability evaluation module is used to extract the factor marginal contribution feature curve based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, calculate the comprehensive suitability score of photovoltaic expansion and photovoltaic loss, and generate photovoltaic expansion suitability map and photovoltaic loss suitability map through spatial mapping. The photovoltaic prediction module is used to take the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and the photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton. After determining the operating rules and parameter constraint set of the cellular automaton model, it filters the candidate cell set, controls the number of transformations, and iteratively updates the state of each cell step by step, thereby predicting and simulating the spatial expansion and loss process of photovoltaic facilities. The interactive parameters include neighborhood weight, suitability weight, and target photovoltaic area.
[0115] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
[0116] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. This application is not limited to the exact structures described above and illustrated in the accompanying drawings, and it should not be considered that the specific implementation of this application is limited to these descriptions. For those skilled in the art, various changes and modifications made without departing from the concept of this application should be considered to fall within the protection scope of this application.
Claims
1. A photovoltaic simulation prediction method based on cellular automata, characterized in that, include: Multi-source data from remote sensing, meteorology, topography, and socioeconomic sources are cleaned, spatially registered, interpolated, supervised classification, and rasterized to construct a database of photovoltaic distribution data and environmental factors at a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data. Based on the photovoltaic distribution data, expansion, loss, and maintenance samples are constructed. An optimal classification model for photovoltaic expansion and photovoltaic loss is constructed through a machine learning model. The contribution results of each influencing factor and its marginal effect characteristics are obtained through interpretability analysis. Based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, the factor marginal contribution characteristic curve is extracted, the comprehensive suitability score of photovoltaic expansion and photovoltaic loss is calculated, and the photovoltaic expansion suitability map and photovoltaic loss suitability map are generated through spatial mapping. Using the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton, after determining the operating rules and parameter constraint set of the cellular automaton model, the candidate cell set is screened, the number of transformations is controlled, and the state of each cell is iteratively updated step by step, thereby predicting and simulating the process of photovoltaic facility spatial expansion and loss; the interactive parameters include neighborhood weight, suitability weight, and target photovoltaic area.
2. The method of claim 1, wherein the method further comprises: The steps of cleaning, spatially registering, interpolating, supervising classification, and rasterizing multi-source data from remote sensing, meteorology, topography, and socioeconomic sources to construct a database of photovoltaic distribution data and environmental factors at a unified spatial resolution include: Based on multispectral remote sensing images with a spatial resolution of 10m, a supervised convolutional neural network was used to identify photovoltaic modules, generate binarized photovoltaic identification results, and obtain basic photovoltaic distribution data through vectorization and spatial simplification. Based on remote sensing images with a spatial resolution of 0.1°, the normalized vegetation index was extracted through band operations, the land surface temperature was retrieved using the generalized radiative transfer formula, and a raster dataset of natural environmental factors was constructed by combining meteorological and topographic data. Spatial cleaning and registration of photovoltaic facilities and various vector data are performed on a GIS platform. Slope and aspect are calculated based on a digital elevation model. Kriging interpolation is used to transform discrete meteorological station data into continuous spatial meteorological data. Distance-type indicators are calculated using Euclidean distance. The spatial intensity of terminal energy consumption facilities is quantified using kernel density estimation. A social environmental factor dataset is constructed by combining land use data, GDP, population, and policy driving factors. The method for quantifying the spatial intensity of terminal energy consumption facilities is as follows: a suitable search radius is determined by kernel density values, and Ripley's K function is used to identify the dominant spatial clustering scale of terminal energy consumption facilities. All raster layers and vector layers in the photovoltaic basic distribution data, natural environmental factor raster dataset, and social environmental factor dataset are projected onto the same coordinate reference system and resampled according to a preset target resolution to obtain photovoltaic distribution data and environmental factor database under a unified spatial resolution.
3. The method of claim 2, wherein the method further comprises: The expression for the binarized photovoltaic identification result is: wherein, , , is a binary photovoltaic identification result, is a classification threshold, is a predicted probability of photovoltaic occurrence, is a CNN-based nonlinear classification function, is a multispectral feature vector of a pixel , is a trained network parameter, is a reflectance value of a band at a pixel , is a number of spectral bands; The formula for calculating the normalized vegetation index is as follows: wherein, is the normalized difference vegetation index of the pixel, is the near-infrared reflectance of the pixel, is the red band reflectance of the pixel, is the green band reflectance of the pixel. The formula for calculating the land surface temperature is: wherein, is the surface temperature, is the brightness temperature of the pixel, is the effective wavelength in the thermal infrared band, is a constant derived from Planck's law, is the surface emissivity; The formulas for calculating the terrain slope and aspect are as follows: wherein, is a terrain slope, is a terrain aspect, is an elevation, is an elevation gradient along direction, is an elevation gradient along direction; The calculation formula for the continuous spatial meteorological data is as follows: wherein, , is the interpolated weather variable value for the pixel, is the observation value of the weather station, is the kriging weight assigned to the weather station for the pixel, denotes the semi-variogram, is the distance between the weather station and the weather station, is the distance between the pixel and the weather station, is the Lagrange multiplier for imposing the unbiased constraint, is the number of stations participating in the interpolation; The Euclidean distance calculation formula is as follows: wherein is the pixel Euclidean distance to the nearest target element, is the pixel planar coordinates of the pixel planar coordinates of the element planar coordinates of the element The expression for the nuclear density value is: wherein, is the kernel density value at the pixel , represents the kernel function, is the Euclidean distance between the pixel and the facility , is the KDE search radius determined by Ripley’s K function; The expression for Ripley'sK function is: wherein, is the Ripley’s K function value calculated at distance d, is the total area of the target region, is the number of set, is an indicator function that takes the value 1 when the condition is met and 0 otherwise.
4. The method of claim 3, wherein the method further comprises: The steps of constructing expansion, loss, and maintenance samples based on the photovoltaic distribution data, building an optimal classification model for photovoltaic expansion and loss through a machine learning model, and obtaining the contribution results and marginal effect characteristics of each influencing factor through interpretability analysis include: Based on the photovoltaic distribution data from different periods, geographic calculations are performed on photovoltaic changes in adjacent time periods, photovoltaic change types are classified in the target area, and a unique photovoltaic change code is assigned to each pixel based on the combination of photovoltaic distribution data from previous and subsequent time periods to obtain a dependent variable dataset; the photovoltaic change types include photovoltaic expansion and photovoltaic loss. Based on the natural environmental factor data and the social environmental factor data, a bivariate statistical analysis was performed in the partitioned area to remove redundant and weakly correlated indicators and to select significantly correlated environmental and socioeconomic indicators to construct an independent variable dataset. A stratified sampling strategy is employed to construct a training set of positive and negative samples for photovoltaic expansion and photovoltaic loss based on the dependent variable dataset and the independent variable dataset. This training is used to train a classification model for photovoltaic expansion and photovoltaic loss. Model performance is evaluated and models are selected using receiver operating characteristic (ROC) curves, area under the curve (AUC), precision, and recall. Bayesian hyperparameter optimization is then performed to obtain the optimal classification model for photovoltaic expansion and photovoltaic loss. The photovoltaic expansion and photovoltaic loss classification model includes any one or more of the following: Logistic regression model, K-nearest neighbor model, random forest model, extreme gradient boosting model, lightweight gradient boosting machine model, and class boosting model. The interpretability analysis of the optimal classification model for photovoltaic expansion and photovoltaic loss is performed based on the SHAP method. The marginal contribution of each driving factor in the independent variable dataset to the predicted probability of photovoltaic changes is quantified, and the contribution results and marginal effect characteristics of each influencing factor are obtained.
5. The method of claim 4, wherein the method further comprises: The photovoltaic change coding formula is as follows: wherein, a photovoltaic change encoding of the pixel, a photovoltaic state of the preceding time period, a photovoltaic state of the preceding time period, a photovoltaic state of the following time period; The marginal contribution expression of the driving factor to the photovoltaic change prediction probability is as follows: wherein, is a factor for a pixel contribution to the prediction result at the pixel; is a set of all driving factors.
6. The method of claim 5, wherein the method further comprises: The steps of extracting factor marginal contribution feature curves based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, calculating the comprehensive suitability score of photovoltaic expansion and photovoltaic loss, and generating photovoltaic expansion suitability maps and photovoltaic loss suitability maps through spatial mapping include: Based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, the positive and negative sample training sets of photovoltaic expansion and photovoltaic loss are processed in a hierarchical manner, the marginal contribution feature curves of each influencing factor are extracted, the positive marginal contribution curves are retained by the ReLU filtering mechanism, the positive marginal contribution curves are approximated by the smoothing function, and a contribution score lookup table for each factor is constructed. The marginal contribution values obtained through the contribution score lookup table of each factor are linearly normalized to the integer range of 1-255, aggregated using geographical weights, and the comprehensive suitability score of photovoltaic expansion and photovoltaic loss is calculated. The comprehensive suitability score is spatially mapped on the GIS platform to generate a photovoltaic expansion suitability map and a photovoltaic loss suitability map, respectively; the spatial resolution of the suitability map is 50m.
7. The method of claim 6, wherein the method further comprises: The formula for calculating the positive marginal contribution is as follows: wherein, , is a factor the marginal contribution of the ReLU filter, is a factor the marginal SHAP contribution function of is the SHAP value of the factor at the pixel is the value of the factor at the pixel ; The fitted positive marginal contribution function is: wherein, is the fitted marginal contribution value; is the factor fitted response function; The normalization rule formula is as follows: wherein as the pixel the normalized suitability score of the pixel as the pixel as the pixel The formula for calculating the overall suitability score of photovoltaic expansion and photovoltaic loss is as follows: wherein, , is the integrated photovoltaic suitability score for the pixel ; is the geographic weight for the factor ; is the number of driving factors. 8.The method of claim 1, wherein, The steps of predicting and simulating the spatial expansion and loss process of photovoltaic facilities, using the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton, include: determining the operating rules and parameter constraint set of the cellular automaton model, screening candidate cell sets, controlling the number of transformations, and iteratively updating the state of each cell step by step. Using the current photovoltaic distribution data in the photovoltaic distribution data as the initial state of the cellular automaton model, and combining the photovoltaic expansion suitability map and the photovoltaic loss suitability map, the initial input dataset of the cellular automaton model is obtained; A binary state transition rule for photovoltaic expansion and loss is set up, an interactive parameter system is introduced to calculate the neighborhood photovoltaic density, and the comprehensive transition potential is calculated by suitability score, neighborhood photovoltaic density and random perturbation to obtain the operating rules and parameter constraint set of the cellular automata model. The set of candidate pixels that meet the threshold is determined by the cell suitability score. The number of conversions is controlled by setting the conversion intensity parameter for each iteration. Then, based on the operating rules and parameter constraint set of the cellular automata model, the state of each cell is iteratively updated step by step to dynamically restore the spatial change process of photovoltaic expansion and photovoltaic loss. When the iteration reaches the preset target photovoltaic area or the maximum number of iterations, the simulation process is terminated, and a simulation map of the photovoltaic spatial distribution at different time points is output to predict and simulate the spatial expansion and loss process of photovoltaic facilities.
9. The method of claim 8, wherein the method further comprises: The formula for calculating the neighborhood photovoltaic density is as follows: wherein, , is the neighborhood photovoltaic density, is the pixel At the iteration of the photovoltaic state, is the neighborhood window centered at the pixel , is the number of pixels in the neighborhood. The formula for calculating the comprehensive transfer potential is as follows: wherein, is a comprehensive metastatic potential, is a pixel of a normalized suitability score, is a neighborhood photovoltaic density; is a random variable subject to a uniform distribution, is a suitability weight, is a neighborhood weight, is a random weight, and ; The candidate pixel set is as follows: wherein, is the first iteration corresponds to the candidate pixel set; The number of transformations in each iteration is: wherein, is the number of transformations for each iteration, is the intensity of the transformation that controls each iteration.
10. A photovoltaic simulation prediction system based on cellular automata, characterized by, The system is used to execute the photovoltaic simulation and prediction method based on cellular automata as described in any one of claims 1 to 9, and the system comprises: The data processing module is used to clean, spatially register, interpolate, supervised classify, and rasterize multi-source data from remote sensing, meteorology, topography, and socioeconomic sources, and to construct a database of photovoltaic distribution data and environmental factors at a unified spatial resolution; the environmental factor data includes natural environmental factor data and social environmental factor data. The model building module is used to construct expansion, loss, and maintenance samples based on the photovoltaic distribution data, construct the optimal classification model of photovoltaic expansion and photovoltaic loss through machine learning model, and obtain the contribution results and marginal effect characteristics of each influencing factor through interpretability analysis. The suitability evaluation module is used to extract the factor marginal contribution feature curve based on the photovoltaic expansion and photovoltaic loss classification model and the marginal effect characteristics, calculate the comprehensive suitability score of photovoltaic expansion and photovoltaic loss, and generate photovoltaic expansion suitability map and photovoltaic loss suitability map through spatial mapping. The photovoltaic prediction module is used to take the current photovoltaic distribution as the initial state, the photovoltaic expansion suitability map and the photovoltaic loss suitability map as evolutionary constraints, and the interactive parameterization settings for photovoltaic layout as the basic parameters of the cellular automaton. After determining the operating rules and parameter constraint set of the cellular automaton model, it filters the candidate cell set, controls the number of transformations, and iteratively updates the state of each cell step by step, thereby predicting and simulating the spatial expansion and loss process of photovoltaic facilities. The interactive parameters include neighborhood weight, suitability weight, and target photovoltaic area.