Site selection model construction and site selection method for artificial chamber in compressed air energy storage system
The site selection model is constructed through the deep neural network model, which solves the problem of cumbersome location selection process and lack of objectivity caused by manual experience in the existing technology, and achieves more efficient and accurate site selection decision support.
Patent Information
- Application Number
- CN202510708113.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
The existing artificial chamber site selection methods of compressed air energy storage power plants rely on manual experience, resulting in cumbersome processes, inefficient efficiency, lack of objectivity, and difficult to meet the requirements of modern energy storage power plants for site selection accuracy and data security.
The site selection model is constructed using deep neural network model, and through feature extraction and associated data set training, an artificial chamber site selection model is generated, and regional feature information is comprehensively considered to provide scientific and objective site selection support.
It improves site selection efficiency and accuracy, provides more scientific and objective site selection decision support, and meets the needs of modern energy storage power plants.
Smart Images

Figure CN120579451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic technology, and in particular to a site selection model construction and a site selection method for an artificial chamber in a compressed air energy storage system. Background Art
[0002] The site selection for a compressed air energy storage power station is a complex, multi-faceted decision-making process. Its feasibility and feasibility are constrained by numerous factors, including geographic location, geological conditions, energy demand, infrastructure, environmental, and social impacts.
[0003] In related technologies, site selection methods primarily rely on manual data collection and empirical judgment. This approach is not only cumbersome and inefficient, but also results in fragmented and opaque data. More critically, site selection results are heavily influenced by human subjectivity and lack objectivity, making them unable to meet the stringent requirements of modern energy storage power plants for site selection efficiency, accuracy, and data security. Summary of the Invention
[0004] In view of this, the present invention provides a site selection model construction and site selection method for artificial chambers in a compressed air energy storage system to solve the problem that the site selection of artificial chambers in related technologies is not scientific and objective based on human experience.
[0005] In a first aspect, the present invention provides a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system, the method comprising: obtaining target data corresponding to different regions and site suitability data of each region, the target data being used to characterize data information related to the site selection of the artificial chamber in the region; performing feature extraction on the target data of each region to obtain target feature data of the corresponding region; associating the target feature data of each region with the site suitability data to obtain an associated data set; and using the associated data set to train a pre-constructed deep neural network model until the model accuracy meets preset conditions, thereby obtaining a site selection model for the artificial chamber.
[0006] The method for constructing a site selection model for an artificial chamber in a compressed air energy storage system provided by the present invention performs feature extraction on target information data of each region to obtain target feature data of the corresponding region; associates the target feature data of each region with site suitability data to obtain a correlated data set; and uses the correlated data set to train a pre-constructed deep neural network model until the model accuracy meets preset conditions to obtain a site selection model for the artificial chamber. The site selection model for the artificial chamber can predict the site suitability of the corresponding region based on the target feature information of different regions, thereby determining a site selection plan. The target information data in the region is comprehensively considered during site selection, providing more scientific, objective and efficient support for site selection decisions.
[0007] In an optional embodiment, the step of performing feature extraction on the target information data of each region to obtain feature data for the region includes: classifying the target information data of each region according to different data structures to obtain multiple classification data of corresponding partitions; performing feature extraction on each classification data of each partition to obtain first feature data corresponding to different classification data of the corresponding partition; and fusing the multiple first feature data of each partition to obtain target feature data of the corresponding partition.
[0008] In an optional embodiment, the target data corresponding to different regions are obtained by the following steps: obtaining initial data corresponding to different regions; and preprocessing the initial data of each partition to obtain target data for the corresponding partition.
[0009] In a second aspect, the present invention provides a method for site selection of an artificial chamber in a compressed air energy storage system, the method comprising: acquiring data of a target area; performing feature extraction on the data of the target area to obtain target feature data of the target area; inputting the target feature data of the target area into a pre-constructed site selection model of an artificial chamber, so that the site selection model of the artificial chamber outputs first site selection suitability data of the target area, and the site selection model of the artificial chamber is determined by the method for constructing a site selection model of an artificial chamber in a compressed air energy storage system according to the first aspect or any corresponding embodiment thereof.
[0010] The site selection method for an artificial chamber in a compressed air energy storage system provided by the present invention inputs the target characteristic data of the target area into a pre-constructed site selection model for the artificial chamber, so that the site selection model for the artificial chamber outputs first site selection suitability data for the target area. Based on the first site selection suitability data, it can be determined whether the target area is suitable as a site selection area. The target information data in the area is comprehensively considered during site selection, providing more scientific, objective and efficient support for site selection decisions.
[0011] In an optional embodiment, the method further includes: obtaining preset indicator data of the target area; determining second site suitability data of the target area based on the preset indicator data; and evaluating the site suitability of the target area using the first site suitability data and the second site suitability data to obtain an evaluation result.
[0012] In an optional embodiment, the method further includes: if the target area is determined to be the site selection area based on the evaluation results, obtaining a membership matrix of the target area, the membership matrix being used to characterize the evaluation information corresponding to different evaluation factors; determining a weight vector of each evaluation factor in the membership matrix using a preset analysis method; and determining a comprehensive result of the target area based on the weight vector of each evaluation factor and the membership matrix.
[0013] In the third aspect, the present invention provides a device for constructing a site selection model for an artificial chamber in a compressed air energy storage system, the device comprising: a first acquisition module for acquiring target data corresponding to different regions and site selection suitability data of each region, the target data data being used to characterize data information related to the site selection of the artificial chamber in the region; a first extraction module for performing feature extraction on the target data of each region to obtain target feature data of the corresponding region; an association module for associating the target feature data of each region with the site selection suitability data to obtain an associated data set; a training module for using the associated data set to train a pre-constructed deep neural network model until the model accuracy meets preset conditions, thereby obtaining a site selection model for the artificial chamber.
[0014] In a fourth aspect, the present invention provides a site selection device for an artificial chamber in a compressed air energy storage system, the device comprising: a second acquisition module for acquiring data of a target area; a second extraction module for performing feature extraction on the data of the target area to obtain target feature data of the target area; a first determination module for inputting the target feature data of the target area into a pre-constructed site selection model of an artificial chamber, so that the site selection model of the artificial chamber outputs first site selection suitability data of the target area. The site selection model of the artificial chamber is determined by the site selection model construction method for an artificial chamber in a compressed air energy storage system according to the first aspect or any corresponding embodiment thereof.
[0015] In a fifth aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the computer instructions to thereby execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to the first aspect or any corresponding embodiment thereof, or to execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to the second aspect or any corresponding embodiment thereof.
[0016] In a sixth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to the first aspect or any corresponding embodiment thereof, or to execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to the second aspect or any corresponding embodiment thereof.
[0017] In the seventh aspect, the present invention provides a computer program product comprising computer instructions, the computer instructions being used to enable a computer to execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to the first aspect or any corresponding embodiment thereof, or to execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to the second aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 1 is a flow chart of a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention;
[0020] Figure 2 is a flow chart of a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to another embodiment of the present invention;
[0021] Figure 3 Schematic diagram of a method for selecting a site for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention
[0022] Figure 4 1 is a structural block diagram of a device for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention;
[0023] Figure 5 1 is a structural block diagram of a site selection device for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention;
[0024] Figure 6 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0026] In related technologies, site selection methods primarily rely on manual data collection and empirical judgment. This approach is not only cumbersome and inefficient, but also results in fragmented and opaque data. More critically, site selection results are heavily influenced by human subjectivity and lack objectivity, making them unable to meet the stringent requirements of modern energy storage power plants for site selection efficiency, accuracy, and data security.
[0027] In view of this, the embodiment of the present application provides a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system, which can be applied to a server to realize the construction of a site selection model for an artificial chamber. The method provided in the embodiment of the present application performs feature extraction on the target information data of each region to obtain the target feature data of the corresponding region; associates the target feature data of each region with the site selection suitability data to obtain a related data set; uses the related data set to train a pre-built deep neural network model until the model accuracy meets the preset conditions, thereby obtaining a site selection model for the artificial chamber. The site selection model for the artificial chamber can predict the site selection suitability of the corresponding region based on the target feature information of different regions, thereby determining a site selection plan. The target information data in the region is comprehensively considered during site selection, providing more scientific, objective and efficient support for site selection decisions.
[0028] According to an embodiment of the present invention, an embodiment of a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.
[0029] In this embodiment, a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system is provided, which can be used for the above-mentioned server. Figure 1 FIG. 1 is a flow chart of a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0030] Step S101: acquiring target data corresponding to different regions and site suitability data of each region. The target data is used to represent information related to site selection of artificial chambers in the region.
[0031] For example, the area may be an area that requires an assessment of the suitability of artificial chamber site selection. Different areas have different locations, and the target data may include but are not limited to meteorological, hydrological, topographic and geological, resource, power system and other related data.
[0032] Step S102 : extracting features from the target data of each region to obtain target feature data of the corresponding region.
[0033] For example, in the embodiment of the present application, target feature data is extracted from the target data according to a preset feature extraction method. The embodiment of the present application does not limit the preset feature area method, and those skilled in the art can determine it according to needs.
[0034] Step S103 : Associating the target feature data of each region with the site suitability data to obtain an associated data set.
[0035] For example, the embodiment of the present application does not limit the manner of associating target feature data with site suitability data, and those skilled in the art can determine it according to needs.
[0036] Step S104: Use the associated data set to train the pre-built deep neural network model until the model accuracy meets the preset conditions, thereby obtaining the site selection model of the artificial chamber.
[0037] For example, the preset conditions can be determined based on experience, and the specific content of the preset conditions is not limited in the embodiments of the present application. In the embodiments of the present application, the pre-built deep neural network model adopts a hybrid architecture that combines a multi-layer perceptron (MLP) with a convolutional neural network (CNN) and a recurrent neural network (RNN). The MLP is used to process the integrated feature vector after fusion. It contains multiple fully connected layers and introduces nonlinear factors through nonlinear activation functions (such as ReLU functions) to enhance the expressive power of the model. The CNN part is mainly used to further extract and process the local spatial correlation information that may exist in the fused features. For example, when processing features related to geographic space, CNN can capture the correlation patterns between features of different geographic locations. RNN or its variants, such as the long short-term memory network (LSTM) and the gated recurrent unit (GRU), are used to process the time series information in the fused features to ensure that the model can effectively utilize the dynamic change laws of historical data, such as the impact of the time evolution trend of power load curves and water level change series on site selection decisions. For the MLP, the number of neurons in each fully connected layer is appropriately set based on the dimension and complexity of the input fusion features. For example, the first fully connected layer can be set to 256 neurons, with subsequent layers gradually reducing the number to compress the feature representation and extract higher-level features. For the CNN, the convolution kernel size, stride, and number of convolution layers are appropriately set based on the spatial structure of the fused features. For example, a 3x3 convolution kernel with a stride of 1 and two to three convolution layers are used to extract local spatial features. For the RNN, the number of neurons in the hidden layer and the time step are determined based on the length and periodicity of the time series. For example, if power load curve data has distinct daily and weekly cycles, the corresponding time step can be set to capture this periodicity, and an appropriate number of hidden neurons can be used to process these time series characteristics. Furthermore, to prevent overfitting, a dropout layer is introduced after the fully connected and convolutional layers to randomly drop a certain percentage of neuronal connections, enhancing the model's generalization ability.
[0038] In an embodiment of the present application, an Adaptive Moment Estimation (Adam) optimization algorithm is selected to train the gradient first-order moment estimation and second-order moment estimation of the model to dynamically adjust the learning rate of each parameter, which can be modeled during the training process. The Adam algorithm can converge to a better solution faster according to the parameters, and adaptively adjust the learning rate of different parameters, which is suitable for processing complex deep neural network model training. At the same time, a suitable learning rate decay strategy is set to gradually reduce the learning rate as the training proceeds, such as by using methods such as exponential decay or cosine annealing decay, so that the model can adjust the parameters more finely in the later stage of training, further improving the performance of the model. During the training process, by monitoring the loss value and evaluation indicators (such as accuracy, root mean square error, etc.) on the validation set, the hyperparameters and training strategies of the model are adjusted in a timely manner to ensure that the performance of the model continues to improve until the expected performance standards are reached.
[0039] Furthermore, the associated dataset is divided into training, validation, and test sets according to a certain ratio. Typically, 70% of the data is used as the training set to train the parameters of the deep neural network model; 20% of the data is used as the validation set to adjust the model's hyperparameters and evaluate the model's performance during training to prevent overfitting; and the remaining 10% of the data is used as the test set for the final evaluation of the model's generalization ability and accuracy. When dividing the data, ensure that the distribution of different data types in each dataset is similar to that of the original dataset to ensure the reliability of the model under different data conditions. For example, the different value ranges of numerical data, the proportion of each category of categorical data, the topic distribution of text data, the geographical area coverage of image data, and the time period characteristics of sequence data should all remain relatively consistent in the training, validation, and test sets.
[0040] The constructed deep neural network model is trained using the training set. During training, the mini-batch gradient descent method is used to divide the training set into several small batches. Each iteration uses a small batch of data to calculate gradients and update model parameters. For example, setting the mini-batch size to 32 or 64 can leverage the parallel computing power of the GPU to accelerate the training process and help the model converge stably. After each training epoch, the model loss and evaluation metrics (such as accuracy and root mean square error) are evaluated on the validation set. Based on the results, the model hyperparameters (such as the learning rate, dropout rate, and number of convolution kernels) are adjusted. If the loss on the validation set no longer decreases or the evaluation metrics no longer improve, the model may be overfitting. In this case, an early stopping strategy can be implemented to stop training and save the current best-performing model parameters.
[0041] During training, continuously monitor the model's performance on the validation set. Use visualization tools, such as plotting loss curves and evaluation metric curves, to visually observe the model's training process and performance trends. If the model's performance on the validation set fluctuates significantly or shows signs of overfitting, various adjustments can be made. For example, increase the amount of training data and expand the dataset through data augmentation techniques (such as rotating, flipping, and cropping image data; performing synonym replacement, randomly inserting or deleting words, etc.); adjust the model's complexity, such as reducing the number of layers or neurons in the neural network; and optimize hyperparameters, such as adjusting the initial value and decay strategy of the learning rate and increasing the strength of regularization. Simultaneously, analyze and improve feature extraction and fusion methods for different data types. If features of certain data types are found to be ineffective in model training or to cause performance degradation, re-examine the preprocessing and feature extraction process, optimize the feature representation, or adjust the weight distribution during the fusion process. After model training is complete, use the test set to fully evaluate the final model. Calculate the model's accuracy, recall, F1 value (for classification tasks) or root mean square error, mean absolute error and other indicators (for regression tasks) on the test set to quantify the model's performance. At the same time, conduct a detailed error analysis to see on which data samples the model has incorrect predictions or large errors, and analyze the relationship between these errors and data characteristics, model structure or training process. For example, if it is found that the model frequently makes mistakes in site selection predictions in certain specific geographical areas, it may be necessary to further check the completeness and accuracy of the data in that area, or consider whether the model needs to be specially adjusted or optimized based on the characteristics of the area. Based on the evaluation results and error analysis, make final improvements and improvements to the model to ensure that the model can provide reliable and accurate site selection recommendations for compressed air energy storage artificial chambers in practical applications.
[0042] In the embodiment of the present application, k-fold validation can also be used to train and validate the model. The specific steps are as follows:
[0043] (1) Dataset partitioning. The correlated dataset is randomly divided into K non-overlapping subsets. The data distribution of each subset is consistent with the original dataset (such as the range of numerical data, the proportion of categorical data, the geographical coverage of image data, etc.). Example: If K = 5, then 4 subsets are taken as training sets each time, and the remaining subset is used as the validation set, and a total of K independent training and validation are performed.
[0044] (2) Fold-by-fold training and validation. In the training phase, in the kth iteration, the training set after merging the 1st to k-1st and k+1st to Kth subsets is used as input to the deep neural network model (MLP+CNN+RNN hybrid architecture). The model parameters are updated using the Adam optimization algorithm. The goal is to minimize the loss function (cross entropy for classification tasks and mean square error for regression tasks). Validation phase: Use the kth subset as the validation set to calculate the evaluation indicators of the model on the validation set (such as classification accuracy and regression root mean square error). At the same time, record the contribution of different data type features to the model decision (observed through the attention mechanism weight, such as whether the feature weight of geological data is higher than that of meteorological data).
[0045] (3) Evaluation of feature fusion effect. Compare the fluctuations of model performance in each verification result and analyze the stability of the fusion features. For example, if the model's dependence on the "groundwater level change" related features in a certain fold verification is abnormally high, it is necessary to check whether there is data anomaly or feature extraction bias in this subset. For the key components of the fusion features (such as the weights assigned by the attention mechanism and the terrain feature vectors extracted by CNN), use visualization tools (such as feature importance heat maps) to observe their changes in different folds to ensure that the core information of multi-source data (such as geological stability and grid access feasibility) is stably captured during the verification process.
[0046] The method for constructing a site selection model for an artificial chamber in a compressed air energy storage system provided in this embodiment performs feature extraction on target data of each region to obtain target feature data of the corresponding region; associates the target feature data of each region with site suitability data to obtain a correlated data set; and uses the correlated data set to train a pre-constructed deep neural network model until the model accuracy meets preset conditions, thereby obtaining a site selection model for the artificial chamber. The site selection model for the artificial chamber can predict the site suitability of the corresponding region based on the target feature information of different regions, thereby determining a site selection plan. The target data in the region are comprehensively considered during site selection, providing more scientific, objective and efficient support for site selection decisions.
[0047] In this embodiment, a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system is provided, which can be used for the above-mentioned server. Figure 2 FIG. 1 is a flow chart of a method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0048] Step S201: Obtain target data corresponding to different regions and site suitability data of each region. The target data is used to represent the information related to the site selection of artificial chambers in the region. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0049] In some optional implementations, the target data corresponding to different regions are obtained by the following steps:
[0050] Step a1: Obtain initial data corresponding to different regions.
[0051] For example, in the embodiment of the present application, the initial data may cover meteorological data, hydrological data, topographic and geological data, resource data, power system data and other related data of the corresponding area, totaling 5 major categories and 22 subcategories. For specific data, please refer to Table 1 below.
[0052] Table 1
[0053]
[0054]
[0055] Step a2: pre-processing the initial data of each partition to obtain the target data of the corresponding partition.
[0056] For example, in the embodiments of the present application, different types of data in the initial data have different corresponding data structures. To better preprocess the initial data, different preprocessing operations are performed on data of different data structures. Specific information of different data structures can be shown in Table 2 below.
[0057] Table 2
[0058]
[0059]
[0060]
[0061] In the embodiment of the present application, different pre-processing methods are used for data with different data structures.
[0062] 1. For numerical data, the preprocessing steps are as follows:
[0063] 1) Advanced data cleaning.
[0064] Unit and format unification: Check the recording units of each data, which must be unified into a commonly used and appropriate unit; for example, rainfall and temperature data must be unified into mm and degrees Celsius to ensure consistency in data format. Logical relationship check: Verify the logical rationality between data, for example, the installed power generation capacity should always be greater than or equal to the actual power generation; changes in groundwater levels should comply with local geological conditions and a reasonable fluctuation range under the surrounding environment. If there are unreasonable high or low values, check whether the data source is accurate. Data integrity check: Confirm whether each numerical data is complete in time and space dimensions, such as historical temperature data at different monitoring stations, whether there is data missing for a certain time period or certain areas, if so, mark it and consider how to handle it later.
[0065] 2) Fill in missing values.
[0066] Statistical imputation and model-based imputation methods are used: missing data such as temperature and power load are filled through mean and median imputation, and data related to land area and groundwater level are filled through regression models and machine learning models.
[0067] 3) Outlier detection and correction.
[0068] Statistical method (Z-score method): For data that conforms to a normal distribution or an approximately normal distribution, such as historical rainfall data in normal years, the Z-score value of each data point is calculated. When the Z-score is greater than 3 or less than -3, the data point can be determined as an outlier. For rainfall data determined to be abnormal, further verification is required to determine whether it is caused by a failure of the measuring instrument, extreme weather, or other reasons. If it is an instrument failure, corrections should be made based on data from surrounding stations or reasonable historical data for the same period. If it is an extreme weather situation, special markings can be made so that its particularity can be considered in subsequent analysis. Box plot method: In power load data, a box plot is used to determine the range of outliers. Data points that are less than Q1-1.5*IQR (interquartile range, IQR=Q3-Q1) or greater than Q3+1.5*IQR are considered outliers. For abnormal power load values, if it is determined that the erroneous data is caused by temporary failures of large-scale industrial equipment or sudden failures of the power grid, it can be deleted directly; if it is believed to be caused by special activities (such as a sudden increase in power load during a large-scale sports event), it can be smoothed according to the activity pattern and the load conditions in the previous and next time periods, such as taking the weighted average of the loads in the previous and next time periods to replace the abnormal value. Model-based methods: For example, using the isolation forest algorithm, numerical data is input into the algorithm, which can automatically identify abnormal values in the data. For data such as water level changes that are affected by multiple factors, the isolation forest algorithm can find abnormal water level change values that are significantly different from the distribution of most data based on the overall distribution characteristics of the data, and then make corrections based on the actual situation, combined with local hydrogeological conditions, recent precipitation and water use conditions and other factors.
[0069] 4) Data normalization processing.
[0070] Minimum-to-maximum normalization: If you want to combine different types of meteorological data (such as historical rainfall and temperature) and power-related data (such as installed power generation capacity and power load levels) for comprehensive analysis, minimum-maximum normalization can be used. For example, consider daily rainfall data for a region over several years, with the minimum value set to 0 and the maximum value set to 100 (in millimeters). For a single day with 50 mm of rainfall, the normalized value becomes . Similarly, other data are normalized to the interval [0, 1] based on their respective minimum and maximum values. This allows data of different dimensions and ranges to be processed on the same scale by the neural network. Z-score normalization: When analyzing data such as earthquake intensity values and elevation values, which may better conform to a normal distribution or to eliminate the influence of dimension on data characteristics, Z-score normalization can be used. Taking earthquake intensity values as an example, first calculate their mean and standard deviation, then normalize each earthquake intensity data point. This allows earthquake intensity data from different regions and time periods to be compared and analyzed on the same scale, helping the neural network to more accurately discover features and patterns in the data.
[0071] 2. For categorical data, the preprocessing process is as follows:
[0072] 1) Advanced data cleaning.
[0073] Duplicate category processing: In land resource data, if there are multiple different land types or land use status categories recorded for the same piece of land, it is necessary to determine the only accurate category based on the latest official land planning documents, field surveys, etc., and remove duplicate and inconsistent records. Logical consistency check: Check whether the logical relationship between categories conforms to the actual situation. For example, in power generation types, there should be no power generation type combinations that do not conform to the principles of energy conversion or do not exist in reality. For the classification of geological structures, rock types and soil properties, it is necessary to ensure that they conform to the logic of geological professionals. If unreasonable category labels are found, they need to be corrected based on professional knowledge.
[0074] 2) Fill in missing values.
[0075] Mode imputation: For categorical data such as rainfall seasonality, if data for a particular region is missing, the mode category of the rainfall seasonality for surrounding regions or for the same period in history is calculated and used to fill in the missing values. For example, if most surrounding areas have a distinct "summer rain" pattern, the missing rainfall seasonality for that region can be imputed as "summer rain" type.
[0076] 3) Outlier detection and correction.
[0077] Rule-based method: judge outliers according to pre-set category rules. For example, in water quality category data, if there are abnormal category labels that do not conform to the established standard classifications such as Class I, Class II, and Class III (such as the appearance of non-standard expressions such as "super special grade"), they are judged as outliers, and need to be reclassified or corrected based on the actual detection indicators of the water samples. Frequency-based method: For categories with extremely low frequency of occurrence, their rationality needs to be further verified. For example, in the power load type, if an extremely rare load type label appears, and it is obviously inconsistent with the industrial structure of the area, residents' electricity usage habits, etc., it is necessary to check whether it is a data entry error or an erroneous labeling of a special case, and then make corrections based on the actual situation.
[0078] 4) Data normalization processing.
[0079] One-Hot Encoding: This converts categorical data into numeric vectors for easier processing by neural networks. For example, if land types are classified as "construction land," "agricultural land," and "forest land," one-hot encoding can represent "construction land" as [1, 0, 0], "agricultural land" as [0, 1, 0], and "forest land" as [0, 0, 1]. Similarly, other categorical data, such as power generation type and load type, can be one-hot encoded accordingly. This allows for clear distinctions between different categories in the vector space, making it easier for neural networks to learn the distinct characteristics between them.
[0080] 3. For text data, the preprocessing process is as follows:
[0081] 1) Advanced data cleaning.
[0082] Format Standardization: Standardize the format of texts. For example, for natural disaster records, standardize the time format to "XX-XX-XXXX," and use standard place names. Legal and regulatory texts should be typeset according to a unified legal clause format, with consistent chapter and article numbering. Texts related to historical sites, protected areas, and natural heritage sites should use standardized fonts, sizes, and paragraph formats to facilitate subsequent text analysis and processing. Content Accuracy Verification: Verify the accuracy of key information in texts through careful manual review and comparison with authoritative historical data, official documents, and professional archaeological research results. For example, in natural disaster records, verify the accuracy of information such as the specific time, location, affected area, and extent of damage. For descriptions of historical sites, check whether the site's name, age, architectural style, and historical significance are consistent with verified findings. Text Deduplication and Merging: Texts with duplicated content will be merged based on their importance and relevance. For example, there may be multiple introductory texts written by different people for the same historical site. The parts that repeatedly describe the basic situation of the site can be merged, while retaining the valuable personalized information about the unique research findings, historical stories, etc. of the site in different texts, removing redundant and repetitive content, and making the text more concise and information-rich.
[0083] 2) Fill in missing values.
[0084] Methods based on knowledge graphs or external knowledge bases: If key information is missing in the text, such as the lack of information on the specific construction period of a historical site in a historical site text, external knowledge bases such as historical knowledge graphs, professional archaeological databases, and local historical records can be used to find and supplement the missing information based on related clues such as the name of the site, geographical location, and related historical figures. Manual supplementation when necessary: For some situations where it is difficult to obtain supplementary information through automatic methods, such as the lack of detailed descriptions of the disaster in some little-known natural disasters, professionals can be arranged to consult local archival materials, interview eyewitnesses, or consult experts in related fields to collect accurate information and then supplement and improve it.
[0085] 3) Outlier detection and correction.
[0086] Semantic analysis and logical check: Use natural language processing techniques to perform semantic analysis on the text and check whether the text content conforms to normal logic and language expression habits. For example, in legal and regulatory texts, if there is a situation where a certain clause conflicts with other relevant clauses and the logic is inconsistent, determine the abnormal content through semantic understanding and legal logic sorting, and make corrections based on legal professional knowledge and legislative intent; for the descriptive text of historical sites, if there are expressions that do not conform to the architectural style and cultural background of the historical period in which the site is located, it is determined as an outlier and corrected by referring to authoritative archaeological research results. Method based on text similarity and clustering: Perform clustering analysis on the text data, and focus on checking the texts that are too different from most other texts in terms of theme content, description style, key information, etc. For example, in the introduction texts of protected areas and natural heritages, if there is a text that mainly describes unrelated commercial development content while other texts revolve around themes such as ecological protection and natural landscapes, then this text may be an outlier and be corrected according to the correct theme direction or directly deleted.
[0087] 4) Data normalization processing.
[0088] Word vector representation: Use the word vector model to convert the words in the text into vector form, enabling the text to be quantitatively processed in the vector space. For example, for the text records of natural disasters, by training the word vector model, convert words such as "earthquake", "flood", "mudslide" into corresponding vectors respectively. In this way, the semantic association degree between words can be measured by calculating indicators such as the distance and similarity between vectors, facilitating subsequent operations such as text classification, information extraction, and semantic similarity matching, and laying a foundation for the neural network to process text data. Text standardization: Remove stop words such as "de", "shi", "zai" in the text that have little impact on semantic understanding, and at the same time perform word form reduction on words (for example, restore "running" to "run"), making the text data more standardized and concise, reducing the complexity and redundancy of the text data, and improving the efficiency of the neural network in extracting and analyzing text features.
[0089] IV. For image data, the preprocessing process is as follows:
[0090] 1) High-order data cleaning.
[0091] Image Quality Optimization: For high-resolution topographic maps (such as satellite remote sensing images and aerial photography), check image clarity, color accuracy, contrast, and other quality indicators. If the image is blurry, use image sharpening algorithms to improve clarity. For images with color deviation or poor contrast, use methods such as white balance adjustment and histogram equalization to correct color and enhance contrast, ensuring a clearer and more accurate representation of topographic features and other relevant information. Image Cropping and Stitching: Crop images according to actual needs to remove irrelevant edge areas, blank areas, or areas containing interfering information. For example, satellite remote sensing images may include large expanses of ocean, but the focus of research is on land topography. In this case, crop out the ocean portion and retain only the land-related image content. If the image is composed of multiple partial images (such as a large-area terrain image taken multiple times by a drone), ensure smooth and seamless transitions between the stitched images. Use professional image stitching algorithms to optimize the stitching and ensure image integrity and coherence.
[0092] 2) Fill in missing values.
[0093] Interpolation algorithm filling (for small area missing): For small area missing and damaged parts in the image (such as a small amount of noise in satellite remote sensing images, small occluded areas, etc.), interpolation algorithms can be used to fill them. For example, the bilinear interpolation algorithm estimates the value of the missing pixel through a certain weighted calculation based on the known values of the pixels around the missing pixel, thereby making the image more complete and reducing the impact of small area missing on image feature extraction and subsequent analysis. Image restoration based on deep learning (for large area missing): When there are large areas of damaged or missing areas in the image (such as large areas of wear and tear on some historical map images due to age and improper preservation), the deep learning image restoration model is used to learn the texture, structure and other feature information of other complete areas in the image, and generate content that matches the surrounding area to fill the missing part, so as to restore the integrity and original appearance of the image as much as possible.
[0094] 3) Outlier detection and correction.
[0095] Methods based on pixel value distribution calculate statistical distribution characteristics of image pixel values, such as the mean and standard deviation. Pixels with values that deviate significantly from the mean (e.g., more than three times the standard deviation) are identified as outliers. These outliers may be caused by noise due to sensor failure, transmission errors, or other factors. Filtering algorithms (such as median filtering, which replaces the outlier pixel with the median of surrounding pixels) can be used to correct them, making the pixel value distribution more balanced and preventing them from interfering with image feature extraction. Methods based on image features and region segmentation use image analysis algorithms such as edge detection and region segmentation to identify the edges and regional features of different objects in the image. If the region containing a pixel differs significantly from the surrounding area in terms of shape, texture, or color, and does not conform to the normal content logic of the image (e.g., isolated colored patches in a terrain image that do not match the surrounding terrain features), the pixel or region is identified as an outlier. These outliers can be corrected based on the characteristics of the surrounding normal areas. For example, pixels in the outlier area can be replaced by copying or fusing pixels from similar terrain areas to make them consistent with the overall image style and content.
[0096] 4) Data normalization processing.
[0097] Pixel value normalization: Normalize the pixel values of the image to a specific range, such as [0,1] or [-1,1]. For common 8-bit depth images (pixel value range is 0-255), the pixel values can be normalized to the [0,1] range through a simple linear transformation (such as dividing each pixel value by 255). This can speed up the convergence of the algorithm and improve the stability and performance of the model when performing image analysis (especially when using deep learning convolutional neural networks for feature extraction and model training). It also facilitates the comparison and processing of different image data at the same scale. Image feature normalization (if features are extracted): If various features are further extracted from the image (such as color histogram features, texture features, etc.), in order to make these features comparable, they also need to be normalized to ensure that similar features of different images are in a similar range in terms of values, so as to avoid the impact of large differences in feature values on the neural network's judgment of the importance of image features and the learning effect.
[0098] 5. For sequence data, the preprocessing process is as follows:
[0099] 1) Advanced data cleaning
[0100] Time series continuity check: Ensure that historical meteorological data (such as daily rainfall and temperature changes over many years), water level change data, and power load curves are continuous on the timeline, and check for missing dates or time points. If missing, mark them and consider using appropriate interpolation methods or other means to supplement them, depending on the specific situation, to ensure the temporal consistency of the data so that the neural network can accurately learn the patterns of data change over time. Abnormal pattern detection: Carefully observe fluctuations in the series data to identify abnormal fluctuations or jumps that clearly deviate from normal trends. For example, if an abnormally high load value far exceeds the daily peak value in the power load curve, or an abnormal water level change in water level change data that does not conform to local precipitation and water use patterns, further analysis is needed to determine whether it is caused by equipment failure, emergencies (such as floods, power system failures, etc.), or data recording errors. Anomalies should be marked and recorded for subsequent targeted processing.
[0101] 2) Missing value filling
[0102] Time series interpolation method: For missing values in sequence data, linear interpolation, spline interpolation and other methods are often used to fill in the missing values. For example, in a time series of daily rainfall, if the rainfall data for a certain day is missing, the approximate rainfall value for the missing day can be calculated using linear interpolation based on the rainfall data of the preceding and following adjacent dates, so that the sequence data remains complete in time, which facilitates the neural network to learn the continuous change pattern of the data. Model-based prediction filling: Using a time series prediction model (such as an ARIMA model, a long short-term memory network, etc.), the model is trained based on the existing complete sequence data, and then the missing values are predicted. For water level change data, if there is a long period of missing data, a suitable time series model can be constructed, taking into account relevant factors such as historical water level change trends, precipitation conditions, water use conditions, etc., to predict and fill in the missing water level values, so that the sequence can be restored to a relatively complete state, which is conducive to subsequent analysis and neural network modeling.
[0103] 3) Outlier detection and correction
[0104] Statistical detection methods: Calculate statistical characteristics of the series data, such as the mean, standard deviation, and coefficient of variation. A reasonable threshold (such as a data point deviating from the mean by more than a certain number of standard deviations) is then set to determine whether it is an outlier. For example, in a time series of temperature fluctuations, if a temperature value deviates by more than three standard deviations from the historical mean for the same period, it can be preliminarily identified as an outlier. Further verification is then conducted based on factors such as the prevailing weather conditions and the presence of unusual meteorological phenomena. If the data is erroneous, corrections can be made based on reasonable preceding and following temperature values. Graphical methods: Time series graphs: By plotting the time series data, the fluctuations of the data points can be observed to identify outliers that clearly deviate from the overall trend. Boxplots: For time series data, boxplots can be plotted by time window (e.g., monthly or quarterly) to identify outliers within each time window. Model-based methods: Use time series forecasting models (such as ARIMA and LSTM) to model the series data and predict the value at each time point. The difference between the actual observed value and the predicted value is compared; data points with excessively large differences may be considered outliers.
[0105] 4) Data normalization
[0106] The following are the steps for normalizing sequence data: Min-Max Normalization: Calculate the minimum and maximum values of the sequence data. Use the formula (X-Min) / (Max-Min) to transform each data point X to a value between 0 and 1. Z-Score Normalization: Although the Z-score method is primarily used for outlier detection, it can also be used as a data normalization method. By calculating the deviation of each data point from the mean and dividing by the standard deviation, the data can be transformed into a distribution with a mean of 0 and a standard deviation of 1. Decimal Scaling Normalization: Scale the data to a specific range by moving the decimal point. For example, if the maximum absolute value of the data is less than 100, each data point can be divided by 100 (or a larger number) to scale it to between 0 and 1 (or a smaller range). Logarithmic Transformation: For data with a long-tail distribution, a logarithmic transformation can be used to reduce the difference in the data range. For example, for power load data or financial time series data, a logarithmic transformation can make it closer to a normal distribution. Other Nonlinear Transformations: Depending on the specific circumstances of the data, other nonlinear transformation methods (such as the Box-Cox transformation) can also be considered for data normalization.
[0107] Step S202 : extracting features from the target data of each region to obtain target feature data of the corresponding region.
[0108] Specifically, the above step S202 includes:
[0109] In step S2021 , the target data of each region is classified according to the different data structures to obtain a plurality of classified data corresponding to the partitions.
[0110] For example, in the embodiment of the present application, the specific information of different data structures is shown in Table 2 above and will not be repeated here.
[0111] Step S2022: extract features from each classification data of each partition to obtain first feature data corresponding to different classification data of the corresponding partition.
[0112] Exemplarily, in the embodiment of the present application, data feature extraction includes:
[0113] (1) Numerical data feature extraction: For pre-processed numerical data, such as historical rainfall, temperature, and air pressure, correlation analysis is first performed. Methods such as the Pearson correlation coefficient or the Spearman correlation coefficient are used to find variable combinations that are closely related to the site selection of the compressed air energy storage system. For example, through analysis, it is found that historical rainfall is highly correlated with groundwater level changes, and groundwater level changes will affect the construction and operation safety of gas storage facilities. For these strongly correlated variables, dimensionality reduction methods such as principal component analysis (PCA) or linear discriminant analysis (LDA) are used for feature extraction. PCA can transform multiple related numerical variables into a few unrelated principal components. These principal components are linear combinations of the original variables and can retain most of the information of the original data. For example, when processing multiple meteorological element data, PCA may extract a principal component that comprehensively reflects the degree of climate humidity, which covers information on variables such as rainfall and evaporation. The extracted principal components or variables after dimensionality reduction are used as feature representations of numerical data for subsequent data fusion.
[0114] (2) Text data feature extraction: For pre-processed text data such as natural disaster records, laws and regulations, historical sites, protected areas and natural heritage, natural language processing tools are used to further extract deep semantic features. A text feature extraction model based on deep learning, such as the pre-trained language based on the Transformer encoder (Bidirectional Encoder Representations from Transformers, BERT) model, is used to encode the text and obtain the sentence vector representation of each text. These sentence vectors can capture the contextual semantic information of the text and can better reflect the overall meaning and semantic association of the text compared to traditional word vectors. For example, when processing legal and regulatory texts, the BERT model can accurately understand the meaning of legal terms and the logical relationship between clauses based on the context of the legal provisions, thereby generating more representative feature vectors. These sentence vectors are used as the final feature representation of the text data for the subsequent fusion process.
[0115] (3) Image data feature extraction: For image data such as high-resolution topographic maps, after image preprocessing and normalization, a convolutional neural network (CNN) is used for feature extraction. A specialized CNN architecture is constructed, such as an improved model based on the VGG (Visual Geometry Group) network or ResNet (Residual Network). Through a combination of multiple convolutional layers, pooling layers, and fully connected layers, the topographic features in the image are automatically learned, including the ups and downs of the terrain, slope changes, water system distribution, and texture features of vegetation cover. For example, the convolutional layer can detect edges and local textures in the image, the pooling layer reduces the dimensionality of the features, and the fully connected layer combines the extracted local features into a global feature vector. The feature vector output by the CNN is used as the feature representation of the image data to prepare for data fusion.
[0116] (4) Feature extraction of categorical data: For categorical data, after completing the one-hot encoding, the potential correlation features between categories are further explored. A method based on a graph neural network (GNN) is used to construct the categorical data into a graph structure, where nodes represent different categories and edges represent the relationships between categories (such as similarity, correlation, etc.). Through the message passing mechanism of GNN, nodes can aggregate information from neighboring nodes, thereby learning high-order relationship features between categories. For example, in land resource category data, GNN can be used to discover the correlation between different land types in terms of geographical distribution, utilization potential, etc., generate feature vectors containing this relationship information, and enrich the feature representation of categorical data so that it can be better integrated with other types of data.
[0117] (5) Feature extraction of sequence data: For sequence data such as historical meteorological data series, water level change series, and power load curves, after time series interpolation and outlier processing, recurrent neural networks (RNN) and their variants (such as long short-term memory networks (LSTMs) or gated recurrent units (GRUs)) are used to extract features. These networks can capture the temporal dependencies and dynamic change patterns in sequence data. For example, the LSTM network, through its unique gating mechanism, can effectively handle long-term dependency problems in long sequences and extract features such as seasonal temperature trends and daily and weekly cycle characteristics of power load. The hidden state or final output of the RNN or LSTM output is used as the feature vector of the sequence data for subsequent data fusion operations.
[0118] Step S2023: fuse the multiple first feature data of each partition to obtain the target feature data of the corresponding partition.
[0119] Exemplarily, the plurality of first feature data may be feature vectors. In the embodiment of the present application, the feature fusion step includes:
[0120] (1) Feature vector concatenation and preliminary integration: First, concatenate the numerical data feature vectors, text data sentence vectors, image data CNN output feature vectors, categorical data graph neural network generated feature vectors, and sequence data RNN or LSTM output feature vectors obtained through their respective feature extraction steps. During the concatenation process, ensure the consistency of different types of feature vectors in dimension and order to form an initial joint feature vector. For example, if the dimension of the numerical data feature vector is , and the dimension of the text data sentence vector is , and so on, the dimension of the concatenated joint feature vector is .
[0121] (2) Weight allocation based on attention mechanism: A multi-head attention mechanism is introduced to assign weights to features of different types of data. The multi-head attention mechanism can learn the correlation between features from multiple different representation subspaces, thereby more comprehensively capturing the importance differences of different data types in the site selection problem. Specifically, the joint feature vector is passed through multiple different linear transformation layers to obtain multiple query, key, and value vector groups, each corresponding to an attention head. For each attention head, the similarity score between the query and the key is calculated and normalized by the softmax function to obtain the attention weight. Then, the attention weight is weighted and summed with the corresponding value vector to obtain the output of each attention head. Finally, the outputs of multiple attention heads are spliced or weighted averaged to obtain the fused feature vector after attention weighting. For example, for the topographic image features in the site selection problem, if the correlation with the site selection decision is high in a certain attention head, then the attention weight corresponding to this part of the features will be larger, and its impact on the final result will be more significant during the fusion process.
[0122] (3) Post-processing and optimization of fused features: The fused feature vector weighted by the attention mechanism is further processed and optimized. On the one hand, the fully connected layer is used to perform nonlinear transformation on the fused features to enhance the expressive power of the model. The fully connected layer maps the input fused feature vector to a new feature space by learning a set of weight matrices and bias terms, making the relationship between features more complex and abstract, which is conducive to subsequent site selection analysis and prediction. On the other hand, the batch normalization operation is introduced to normalize the fused features, accelerate the convergence speed during model training, and improve the stability and generalization ability of the model. Batch normalization normalizes the data of each batch so that the mean and variance of the features remain relatively stable during the training process, reducing the training difficulties caused by changes in data distribution.
[0123] (4) Verification and adjustment of fusion features: In order to ensure the effectiveness and accuracy of data fusion, the cross-validation method is used to verify and adjust the fusion features. The dataset is divided into multiple combinations of training sets and validation sets. A deep neural network model based on fusion features is trained on each combination, and the performance of the model is evaluated on the validation set. By observing the impact of different data type features on model performance on different validation sets, the parameters of the attention mechanism and the structure of the fully connected layer are further adjusted to optimize the quality of the fusion features. For example, if it is found that the features of a certain type of data contribute little to the improvement of model performance in multiple validation sets, it may be necessary to re-examine its feature extraction method or adjust its weight distribution strategy in the attention mechanism.
[0124] Step S203: Associate the target feature data of each region with the site suitability data to obtain an associated data set. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0125] Step S204: Use the associated data set to train the pre-built deep neural network model until the model accuracy meets the preset conditions, and obtain the site selection model of the artificial cavern. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0126] In this embodiment, a method for selecting a site for an artificial chamber in a compressed air energy storage system is provided, which can be used for the above-mentioned server. Figure 3 FIG. 1 is a flow chart of a method for selecting a site for an artificial chamber in a compressed air energy storage system according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0127] Step S301: Acquire data of the target area.
[0128] Illustratively, the target area may be any area where suitability prediction of an artificial cavern is required, and the information data is information data related to site selection of the artificial cavern in the target area.
[0129] Step S302: extracting features from the data of the target area to obtain target feature data of the target area.
[0130] For example, the feature extraction process of the resource data is described in the above embodiment and will not be repeated here.
[0131] Step S303: Input the target characteristic data of the target area into the pre-constructed artificial chamber site selection model so that the artificial chamber site selection model outputs the first site selection suitability data of the target area. The artificial chamber site selection model is determined by the site selection model construction method of the artificial chamber in the compressed air energy storage system in the above embodiment.
[0132] Exemplarily, the site selection model for an artificial chamber will output first site selection suitability data for the target area based on the information data of the target area, and based on the first site selection suitability data, it can be determined whether the target area is suitable as the site selection area for the artificial chamber. In an embodiment of the present application, the target area data to be evaluated is processed according to the steps of data collection, preprocessing, feature extraction and fusion to obtain a feature vector consistent with the input format of the training model. The feature vector is input into a trained and verified deep neural network model, and the model will output a prediction result on whether the area is suitable for the construction of an artificial chamber for compressed air energy storage. For the classification model, the output is a probability distribution of different suitability categories, such as probability values corresponding to categories such as suitable, relatively suitable, and unsuitable; for the regression model, the output is a continuous numerical indicator, such as the expected efficiency score or stability assessment value of the energy storage system in the area. Based on the output results of the model, the site selection potential of the target area is preliminarily determined.
[0133] The embodiment of the present application provides a site selection method for an artificial chamber in a compressed air energy storage system, which inputs target feature data of a target area into a pre-constructed site selection model for an artificial chamber, so that the site selection model for the artificial chamber outputs first site selection suitability data for the target area. Based on the first site selection suitability data, it can be determined whether the target area is suitable as a site selection area. The target information data in the area is comprehensively considered during site selection, providing more scientific, objective and efficient support for site selection decisions.
[0134] In some optional embodiments, the above method further includes:
[0135] Step b1: Obtain preset indicator data of the target area.
[0136] Exemplarily, the preset indicator data can be a plurality of site selection evaluation indicators. In an embodiment of the present application, a comprehensive site selection evaluation indicator system is established, which includes, in addition to the results directly output by the model, auxiliary indicators derived from multi-source data. For example, geological stability indicators extracted from topographic and geological data (such as fault distribution density, rock strength grade, etc.), water resource impact indicators derived from hydrological data (such as groundwater level variation, interaction risk between surface water and gas storage, etc.), grid access convenience indicators calculated from power system data (such as distance from existing substations, transmission line capacity margin, etc.), and environmental impact scores comprehensively evaluated from environmental and social data (such as the degree of interference with surrounding ecosystems, satisfaction of protection needs for historical and cultural sites, etc.).
[0137] Step b2: determining second site suitability data of the target area based on preset indicator data.
[0138] Exemplarily, the second site suitability data is determined based on preset indicator data. The embodiment of the present application does not limit the calculation process of the second site suitability data, and those skilled in the art can determine it according to needs.
[0139] Step b3: using the first site suitability data and the second site suitability data to evaluate the site suitability of the target area and obtain an evaluation result.
[0140] Illustratively, in an embodiment of the present application, the first site suitability data and the second site suitability data are combined to conduct a multi-dimensional comprehensive evaluation of the site suitability of the target area to obtain an evaluation result.
[0141] Furthermore, taking into account the uncertainty of the data and the limitations of the model, an uncertainty analysis is performed on the site selection prediction results. Using methods such as Monte Carlo simulation, the uncertainty range of the input data (such as the measurement error range of meteorological data and the uncertainty of the exploration accuracy of geological data) is randomly sampled multiple times, and the model is rerun to obtain a series of prediction results. The distribution of these results is analyzed, and statistics such as the standard deviation and confidence interval of the prediction results are calculated to quantify the degree of uncertainty in the site selection prediction. For example, if the site suitability prediction results of a target area fluctuate greatly in multiple simulations, it means that there is a high degree of uncertainty in the site selection of this area, and further data collection or more detailed on-site investigation is needed.
[0142] Taking into account the uncertainty of the data and the limitations of the model, an uncertainty analysis is performed on the site selection prediction results. Using methods such as Monte Carlo simulation, the uncertainty range of the input data (such as the measurement error range of meteorological data and the uncertainty of the exploration accuracy of geological data) is randomly sampled multiple times, and the model is rerun to obtain a series of prediction results. The distribution of these results is analyzed, and statistics such as the standard deviation and confidence interval of the prediction results are calculated to quantify the degree of uncertainty in the site selection prediction. For example, if the site suitability prediction results of a target area fluctuate significantly across multiple simulations, it indicates that there is a high degree of uncertainty in the site selection of this area, and further data collection or more detailed on-site investigation is needed.
[0143] Step b4: If the target area is determined to be the site selection area based on the evaluation results, a membership matrix of the target area is obtained. The membership matrix is used to represent the evaluation information corresponding to different evaluation factors.
[0144] For example, in the embodiment of the present application, according to the site selection plan, an on-site survey is carried out, which mainly includes: geological exploration (verifying key information such as geological stability and groundwater level), topographic measurement (using drone aerial photography, laser scanning, etc. to measure the topography and landforms of the site selection area) and transportation and economic evaluation (evaluating the transportation convenience and land cost of the site selection area). The evaluation factor set U = {u1, u2, u3, u4}, where u1 represents geological stability-related factors (including stratum stability, geological structure complexity, groundwater conditions, etc.), where u2 represents terrain condition-related factors (such as terrain slope, terrain undulation, site flatness, etc.), where u3 represents transportation convenience-related factors (covering transportation routes, time, cost and transportation development planning impacts in transportation network analysis, etc.), where u4 represents economic cost-related factors (including various construction cost factors, various operating cost factors and potential economic benefits, etc.). The evaluation level set V = {v1, v2, v3, v4, v5}, corresponding to the five levels {excellent, good, medium, poor, bad}, is used to describe the pros and cons of the site selection scheme in terms of various evaluation factors. A number of cross-disciplinary experts (including geologists, surveying and mapping experts, transportation engineering experts, economic experts, etc.) were invited to form an expert group. Based on the results of the on-site survey, relevant data and their own experience, for each evaluation factor u i , for each evaluation level v j To reduce the impact of expert subjective judgment on the evaluation results, the Delphi method was used to conduct multiple rounds of expert consultation until the expert opinions converged or reached an acceptable level of consistency. The expert evaluation opinions were summarized and the membership matrix R was constructed as shown in the following formula:
[0145]
[0146] Among them, r ij Indicates the evaluation factor u i Evaluation level v j The membership matrix R reflects the fuzzy relationship between the evaluation factors and the evaluation levels.
[0147] Step b5: Determine the weight vector of each evaluation factor in the membership matrix using a preset analysis method.
[0148] For example, in the embodiment of the present application, the Analytic Hierarchy Process (AHP) is used to determine the weight vector of each evaluation factor. First, a judgment matrix A is constructed to compare the relative importance of each evaluation factor; then the maximum eigenvalue λ of the judgment matrix A is calculated. max and its corresponding eigenvector w. The eigenvector w is normalized to obtain the weight vector W = {w1, w2, w3, w4}. During the calculation process, a consistency test is performed. If the consistency ratio CR < 0.1, it is considered that the judgment matrix has satisfactory consistency and the weight vector is reasonable; otherwise, the judgment matrix needs to be readjusted until the consistency requirements are met.
[0149] Step b6: Determine the comprehensive result of the target area based on the weight vector and membership matrix of each evaluation factor.
[0150] For example, in the embodiment of the present application, a fuzzy synthesis operation is performed in, Represents the fuzzy synthesis operator, and obtains the comprehensive evaluation result vector B = {b1, b2, b3, b4, b5}.
[0151] Furthermore, a comprehensive analysis of the site selection plan is conducted based on the comprehensive evaluation results. If the comprehensive evaluation is excellent or good, it indicates that the site selection plan performs well overall in terms of geology, topography, transportation, and economy, and has high feasibility and suitability, and can be considered as a priority site selection plan. If the comprehensive evaluation is medium, it indicates that there are some areas in the site selection plan that need improvement and optimization. Further in-depth analysis of the situation of each factor is needed to clarify the direction of improvement and formulate a detailed improvement plan. If the comprehensive evaluation is poor or bad, the site selection plan has many problems and is not recommended. It is necessary to re-select the site or make major adjustments and improvements to the current site, and re-conduct surveys, assessments, and scheme designs.
[0152] In this embodiment, a device for constructing a site selection model for an artificial chamber in a compressed air energy storage system is also provided. The device is used to implement the above-mentioned embodiments and preferred implementation methods, and the details that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0153] This embodiment provides a device for constructing a site selection model for an artificial chamber in a compressed air energy storage system. Figure 4 Shown, including:
[0154] The first acquisition module 401 is used to obtain target data corresponding to different regions and site suitability data of each region. The target data is used to represent information related to the site selection of artificial chambers in the region.
[0155] The first extraction module 402 is used to extract features from the target data of each region to obtain target feature data of the corresponding region;
[0156] The association module 403 is used to associate the target feature data of each region with the site suitability data to obtain an associated data set;
[0157] The training module 404 is used to train the pre-built deep neural network model using the associated data set until the model accuracy meets the preset conditions, thereby obtaining the site selection model of the artificial chamber.
[0158] In some optional implementations, the first extraction module 402 includes:
[0159] The classification submodule is used to classify the target data of each area according to the different data structures and obtain multiple classification data of the corresponding partitions;
[0160] An extraction submodule, configured to extract features from each classification data of each partition, and obtain first feature data corresponding to different classification data of the corresponding partition;
[0161] The fusion submodule is used to fuse multiple first feature data of each partition to obtain target feature data of the corresponding partition.
[0162] In some optional implementations, the target data corresponding to different regions are obtained by the following steps:
[0163] Obtain the initial data corresponding to different regions;
[0164] The initial data of each partition is preprocessed to obtain the target data of the corresponding partition.
[0165] This embodiment provides a device for selecting a site for an artificial chamber in a compressed air energy storage system. Figure 5 Shown, including:
[0166] The second acquisition module 501 is used to acquire information data of the target area;
[0167] The second extraction module 502 is used to extract features from the data of the target area to obtain target feature data of the target area;
[0168] The first determination module 503 is used to input the target feature data of the target area into the pre-constructed artificial chamber site selection model, so that the artificial chamber site selection model outputs the first site selection suitability data of the target area. The artificial chamber site selection model is determined by the artificial chamber site selection model construction method in the compressed air energy storage system of the above embodiment.
[0169] In some optional embodiments, the above device further includes:
[0170] The third acquisition module is used to obtain preset indicator data of the target area;
[0171] A second determination module is used to determine second site suitability data of the target area based on preset indicator data;
[0172] The evaluation module is used to evaluate the site suitability of the target area using the first site suitability data and the second site suitability data to obtain an evaluation result.
[0173] In some optional embodiments, the above device further includes:
[0174] The third acquisition module is used to obtain a membership matrix of the target area if the target area is determined to be a site selection area based on the evaluation results. The membership matrix is used to represent the evaluation information corresponding to different evaluation factors.
[0175] The third determination module is used to determine the weight vector of each evaluation factor in the membership matrix using a preset analysis method;
[0176] The fourth determination module is used to determine the comprehensive result of the target area based on the weight vector and membership matrix of each evaluation factor.
[0177] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0178] The site selection model construction device for an artificial chamber in a compressed air energy storage system or the site selection device for an artificial chamber in a compressed air energy storage system in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0179] The embodiment of the present invention also provides a computer device having the above Figure 4 The site selection model construction device of the artificial chamber in the compressed air energy storage system shown in the figure or the device having the above Figure 5 A siting device for artificial chambers in compressed air energy storage systems.
[0180] See also Figure 6 , Figure 6 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of a GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0181] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0182] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0183] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0184] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0185] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0186] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0187] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0188] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for constructing a site selection model for an artificial chamber in a compressed air energy storage system, characterized in that: The method comprises: Obtaining target data corresponding to different regions and site suitability data for each region, wherein the target data is used to represent information related to site selection of artificial chambers in the region; Extracting features from the target data of each region to obtain target feature data of the corresponding region; Associating the target feature data of each region with the site suitability data to obtain an associated data set; The pre-built deep neural network model is trained using the associated data set until the model accuracy meets the preset conditions, thereby obtaining a site selection model for the artificial chamber.
2. The method according to claim 1, characterized in that The step of extracting features from the target data of each region to obtain feature data for the region includes: Classify the target data of each area according to the different data structures to obtain multiple classification data of the corresponding partitions; Perform feature extraction on each classification data of each partition to obtain first feature data corresponding to different classification data of the corresponding partition; The multiple first feature data of each partition are fused to obtain the target feature data of the corresponding partition.
3. The method according to claim 1 or 2, characterized in that The target data corresponding to the different regions are obtained by the following steps: Obtain the initial data corresponding to different regions; The initial data of each partition is pre-processed to obtain the target data of the corresponding partition.
4. A method for selecting a site for an artificial chamber in a compressed air energy storage system, characterized in that: The method comprises: Obtain data on the target area; Extracting features from the data of the target area to obtain target feature data of the target area; The target characteristic data of the target area is input into a pre-constructed site selection model of an artificial chamber so that the site selection model of the artificial chamber outputs the first site selection suitability data of the target area. The site selection model of the artificial chamber is determined by the site selection model construction method of the artificial chamber in the compressed air energy storage system according to any one of claims 1 to 3.
5. The method according to claim 4, characterized in that The method further comprises: Obtain preset indicator data for the target area; Determining second site suitability data of the target area based on the preset indicator data; The site suitability of the target area is evaluated using the first site suitability data and the second site suitability data to obtain an evaluation result.
6. The method according to claim 5, characterized in that The method further comprises: If the target area is determined to be a site selection area based on the evaluation result, a membership matrix of the target area is obtained, where the membership matrix is used to represent evaluation information corresponding to different evaluation factors; Determining the weight vector of each evaluation factor in the membership matrix using a preset analysis method; A comprehensive result of the target area is determined based on the weight vectors of the evaluation factors and the membership matrix.
7. A device for constructing a site selection model for an artificial chamber in a compressed air energy storage system, characterized in that: The device comprises: The first acquisition module is used to obtain target data corresponding to different areas and site suitability data of each area, wherein the target data is used to represent information related to the site selection of artificial chambers in the area; A first extraction module is used to extract features from the target data of each area to obtain target feature data of the corresponding area; an association module, configured to associate the target feature data of each region with the site suitability data to obtain an associated data set; The training module is used to train the pre-built deep neural network model using the associated data set until the model accuracy meets the preset conditions, thereby obtaining the site selection model of the artificial chamber.
8. A device for selecting a site for an artificial chamber in a compressed air energy storage system, characterized in that: The device comprises: The second acquisition module is used to obtain information data of the target area; A second extraction module is used to extract features from the data of the target area to obtain target feature data of the target area; The first determination module is used to input the target characteristic data of the target area into a pre-constructed artificial chamber site selection model so that the artificial chamber site selection model outputs the first site selection suitability data of the target area. The artificial chamber site selection model is determined by the site selection model construction method of the artificial chamber in the compressed air energy storage system according to any one of claims 1 to 3.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to any one of claims 1 to 3, or executes the method for selecting a site for an artificial chamber in a compressed air energy storage system according to any one of claims 4 to 6 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the method for constructing a site selection model for an artificial chamber in a compressed air energy storage system according to any one of claims 1 to 3, or to execute the method for selecting a site for an artificial chamber in a compressed air energy storage system according to any one of claims 4 to 6.