Underground water potential prediction method based on generalized width learning system
By introducing a weighted generalized width learning system (WBLS) and Chaos Sparrow Search Algorithm (CSSA) in groundwater potential prediction, data imbalance and computational complexity problems are solved, and more efficient and accurate groundwater potential prediction is achieved.
Patent Information
- Application Number
- CN202510010707.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems such as data imbalance, overfitting and computational complexity in groundwater potential prediction, resulting in low accuracy and efficiency of the model.
Weighted generalized width learning system (WBLS) based on Chaos Sparrow Optimization is introduced. By assigning different weights to training samples, especially increasing the weights of a few types of samples, the learning process is optimized and the bias caused by category imbalance is reduced. Combined with the Chaos Sparrow Search Algorithm (CSSA) to optimize the parameters of WBLS to improve the accuracy and efficiency of the model.
It significantly improves the accuracy and efficiency of groundwater potential prediction, reduces the bias and computational complexity of the model, and provides a more accurate and effective method for groundwater potential mapping.
Smart Images

Figure CN120068919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting groundwater potential, in particular to a method for predicting groundwater potential based on a generalized width learning system optimized by chaotic sparrows, belonging to the field of groundwater. Background Art
[0002] As one of the most important freshwater reservoirs on Earth, groundwater is a key lifeline for maintaining ecosystems and facilitating human adaptation to climate change. Groundwater is a fundamental resource in arid and semi-arid regions and has a significant impact on the exploitation and utilization of geological resources. Sustainable management measures for groundwater must be prioritized. Therefore, it is crucial to evaluate groundwater potential. In recent years, machine learning techniques have been widely applied in the field of groundwater potential mapping due to their ability to handle complex relationships among multiple variables. For example, Rahmati et al. (Rahmati et al., "Groundwater spring potential modelling: Comprising the capability and robustness of three different modeling approaches", Journal of Hydrology, vol. 564, pp. 152–167, 2018) used logistic regression (LR) and decision tree (DT) for GWP mapping in Khorasan Province, Iran, and found that the LR model had high accuracy. At the same time, Naghibi et al. (Naghibi et al., "A comparison between ten advanced and soft computing models for groundwater qanat potential assessment in Iran using R and GIS", Theoretical and Applied Climatology, vol. 134, no. 3-4, pp. 865–881, 2018) evaluated the performance of k-nearest neighbor (KNN), artificial neural network (ANN), boosted regression tree (BRT), and random forest (RF), and pointed out that RF performed best in dealing with complex data. In addition, Khosravi et al. (Khosravi et al., "Spatial prediction of groundwater spring potential mapping based on an adaptive neuro-fuzzy inference system and metaheuristic optimization", Hydrology and Earth System Sciences, vol. 22, no. 9, pp. 4771–4792, 2018) effectively improved the prediction accuracy of the Koohdasht-Nourabad Plain in Iran by combining the adaptive neuro-fuzzy inference system (ANFIS) with metaheuristic optimization algorithms.Although the above machine learning methods have achieved certain results in groundwater potential mapping, they also have some limitations. For example, although the LR model performs reliably in data training, its ability to handle non-linear problems is limited and it is prone to overfitting. Although the RF model can effectively process complex data and reduce overfitting and variance problems, its training process is long and the model parameters are difficult to set accurately. Linear SVM performs excellently in high-dimensional spaces and when the number of samples is less than the dimension, but its performance degrades when the dataset is large and noisy. Although ANFIS combined with optimization algorithms has good performance, its complexity and computational requirements are relatively high. Therefore, the present invention introduces a new learning system - the Generalized Learning System (BLS), which learns and predicts data by adjusting the mapping features and the number of enhancement nodes, rather than relying on increasing the number of network layers to improve performance, and uses the method of solving the pseudo-inverse instead of gradient descent to complete the prediction task. The advantages of this method are simple structure, short calculation time, and the ability to process large datasets without a large amount of feature extraction. BLS has been successfully applied in many fields such as short-term wind speed prediction, image processing, speech recognition, industrial fault diagnosis, and biomedical signal processing.
[0003] Although other mainstream optimization algorithms such as GWO, Whale Optimization Algorithm (WOA), PSO, and Differential Evolution (DE) have made certain progress in the field of groundwater, they still face performance challenges when dealing with imbalanced data distributions. Traditional machine learning algorithms usually face balanced data, but in actual situations, the data is often imbalanced, that is, the number of samples in some categories is much larger than that in other categories. This imbalance can lead to biased and inaccurate model classification results. So far, many technical methods have been proposed to alleviate the problem of dataset imbalance from different perspectives. Common methods are divided into two categories: data-level and algorithm-level. Data-level methods such as oversampling and undersampling, such as Gosain et al. (Gosain and Sardana, "Handling class imbalance problem using oversampling techniques: A review", IEEE International Conference on Computing, Communication and Automation (ICCCA), pp. 79–84, 2017), balance the data by increasing the number of minority-class samples or reducing the number of majority-class samples, but may lead to overfitting or removal of key information; algorithm-level methods such as cost-sensitive learning and ensemble learning enhance the recognition of minority classes by optimizing the algorithm, such as Tao et al. (Tao et al., "Self-adaptive cost weights-based support vector machine cost-sensitive ensemble for imbalanced data classification", Information Sciences, vol. 505, pp. 427–443, 2019), but increase the complexity and computational cost.
[0004] In view of the fast training speed and excellent prediction ability of BLS, the present invention applies it to groundwater potential research. Summary of the Invention
[0005] To solve the problems in the prior art, the present invention provides an imbalanced groundwater potential prediction method based on a chaotic sparrow optimization-based generalized width learning system, introducing a Weighted BLS (WBLS) model. By assigning different weights to training samples, especially increasing the weights of minority-class samples, the learning process is optimized, the bias caused by class imbalance is reduced, and the accuracy and fairness of the model are improved.
[0006] The present invention adopts the Chaotic Sparrow Search Algorithm (CSSA), which is a swarm intelligence optimization algorithm that simulates the foraging behavior of sparrows and improves the traditional Sparrow Search Algorithm (SSA); CSSA enhances the global search ability of the algorithm by introducing logical mapping, adaptive hyperparameters, and mutation operators, and is particularly suitable for dealing with complex optimization problems in groundwater datasets. The integration of chaos theory makes CSSA superior to traditional methods, such as the Ant Colony Algorithm and the Particle Swarm Optimization Algorithm (PSO), in terms of automated hyperparameter optimization. The introduction of CSSA significantly improves the accuracy and efficiency of the Generalized Learning System (BLS) model in groundwater potential prediction. Compared with traditional optimization algorithms such as PSO, Grey Wolf Optimization (GWO), and Genetic Algorithm (GA), CSSA not only has a faster optimization speed but also shows stronger robustness in dealing with uneven data distribution. By combining CSSA, the groundwater potential assessment model developed in the present invention demonstrates significant potential in terms of accuracy and reliability, providing important technical support for groundwater research. The present invention mainly applies and improves the Broad Learning System (BLS) to more effectively predict groundwater potential, especially to solve the common data imbalance problem, as well as outliers and noise problems in the data in groundwater potential mapping; a Weighted Broad Learning System based on CSSA (CSSA-WBLS) is developed. This new system not only inherits the advantages of WBLS in dealing with data imbalance but also optimizes parameter search through CSSA, greatly improving the accuracy and efficiency of groundwater potential prediction. CSSA-WBLS provides a new, more accurate, and effective method for groundwater potential mapping, especially showing its important value in decision support for groundwater resource management and protection.
[0007] The improvement strategies of CSSA mainly include three aspects: First, the initialization of the sparrow population position in the SSA algorithm is randomly generated, and the present invention introduces chaotic mapping to improve the quality of the population initial position; Second, in order to accelerate the convergence speed and efficiency of the SSA algorithm, two adaptive hyperparameters are used to update the position of the discoverer and the number of vigilant sparrows; Finally, since the population diversity will decrease during the iteration process, a mutation operator is added to increase the population diversity of SSA, so as to obtain the global optimal value and overcome the local optimum.
[0008] The present invention provides a method for predicting groundwater potential based on a generalized broad learning system optimized by chaotic sparrows, and the specific method includes the following steps:
[0009] S1: Data collection and data processing:
[0010] Combining machine learning models and statistical models, and using ENVI and ArcGIS software to process various data, a groundwater potential map is constructed.
[0011] Specifically, according to the existing data in the study area, remote sensing images, DEM images, data such as the boundaries of the study area, and spring data within the study area are collected; through the processing of ENVI and Arcgis software, a groundwater potential evaluation factor map of the study area is obtained, and a dataset for model training and evaluation is constructed based on the evaluation factors. This dataset consists of eleven groundwater potential evaluation factors: slope, elevation, aspect, plan curvature, profile curvature, standardized precipitation index, sediment transport index, terrain humidity index, normalized vegetation index, road distance, and river distance.
[0012] S2: Establish the BLS model
[0013] Generate feature nodes for each window. Based on the bias value and the corresponding weight matrix randomly generated from the data matrix obtained in step S1, the bias value is placed in the last column of the data matrix, and the two matrices are multiplied to obtain each feature node of the model; secondly, generate enhancement nodes to introduce non-linear factors in the network, and assign the activation function tansig to the previously generated feature nodes, so that the model can better process complex geographical information and thus have a better prediction effect on groundwater; finally, find the pseudo-inverse according to the obtained feature nodes, enhancement nodes, and target output, that is, the mapping from input to output, and then the groundwater potential of the area can be evaluated.
[0014] S3: Introduce WBLS:
[0015] Optimize the learning process by assigning different weights to the training samples, so as to be used for predicting the potential of unbalanced groundwater data.
[0016] Specifically, the newly proposed WBLS model in the present invention adopts a more direct method to calculate weights when dealing with data sample imbalance, without introducing additional penalty functions; this method directly determines the weights based on class information, thereby significantly reducing the training error. Through this strategy, the distribution of samples directly affects the weight assignment, making the majority class obtain relatively small weights, while larger weights are assigned to the minority class.
[0017] S4: Chaotic Sparrow Search Algorithm (CSSA)
[0018] To improve the global optimization ability of SSA, the Logistic mapping, adaptive hyperparameters, and mutation operator are combined with the SSA algorithm to develop an optimization algorithm called CSSA. Compared with SSA, this optimization algorithm has achieved better optimization results.
[0019] S5: Chaotic Sparrow Search Algorithm (CSSA-WBLS)
[0020] To improve the efficiency and accuracy of WBLS, the present invention uses the Chaotic Sparrow Search Algorithm (CSSA) to optimize the relevant parameters of WBLS, thereby obtaining CSSA-WBLS. The detailed description of CSSA-WBLS mainly focuses on two aspects: WBLS network construction and parameter optimization.
[0021] Advantages of the present invention:
[0022] (1) Similar to traditional BLS, WBLS first constructs feature nodes using input data, and then maps the feature nodes to enhancement nodes according to the weights calculated by the formula, and finally calculates the output weights; the newly proposed WBLS model in the present invention adopts a more direct method to calculate weights when dealing with data sample imbalance, without introducing additional penalty functions; this method directly determines weights based on class information, thereby significantly reducing training errors; through this strategy, the distribution of samples directly affects the weight allocation, making the majority class obtain relatively small weights, while larger weights are assigned to the minority class;
[0023] (2) Compared with WBLS, SSA-WBLS shows strong global search ability and fast convergence speed when dealing with high-dimensional optimization problems, but it has a tendency to fall into local optimum in the later stage of iteration; to improve the global optimization ability of the Sparrow Search Algorithm (SSA), accelerate the convergence speed, and avoid falling into local optimum, the present invention proposes a Chaotic Sparrow Search Algorithm (CSSA); the global search ability of the algorithm is enhanced by introducing improvement strategies such as logical mapping, adaptive hyperparameters, and mutation operators. Description of the drawings
[0024] Figure 1 is the BLS network model;
[0025] Figure 2 is the flowchart of BLS and WBLS networks;
[0026] Figure 3 is the flowchart of the sparrow optimization algorithm;
[0027] Figure 4 is the flowchart of the chaotic sparrow optimization algorithm CSSA-WBLS;
[0028] Figure 5 is the distribution map of groundwater spring samples;
[0029] Figure 6 is the ROC curve of the integrated model: (a) training data set; (b) validation data set;
[0030] Figure 7 is the influence factor weight map;
[0031] Figure 8Graph showing the performance evaluation results of various models;
[0032] Figure 9 It is a logical mapping graph. Detailed implementation manners
[0033] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but not to limit the present invention.
[0034] The groundwater potential evaluation method based on CSSA-WBLS provided by the present invention has a process as Figure 4 shown, and includes the following steps:
[0035] Step 1: Collect data such as remote sensing images, DEM images, the boundary of the research area, etc., and the distribution data of groundwater springs in the research area.
[0036] Step 2: Process the obtained data in ENVI and Arcgis software respectively to obtain 11 groundwater potential evaluation factor maps (slope, altitude, aspect, plan curvature, profile curvature, standardized precipitation index, sediment transport index, terrain humidity index, normalized vegetation index, road distance, river distance) and negative sample points in the research area.
[0037] Step 3: Establish the first BLS base learner
[0038] The groundwater potential evaluation factors of all r points including spring points and non-spring points used for model training constitute the input matrix R r×12 . According to the idea of weight update in Adaboost, an initial weight is assigned to all the input data and the input proof is updated using the initial weight to obtain X'. First, the updated input data is processed to obtain feature nodes, and then enhanced nodes are obtained from the feature nodes. The feature nodes and enhanced nodes together constitute the input A in the prediction process. The process from X to map features and the process from map features to enhanced nodes can be expressed by the following formulas respectively.
[0039]
[0040] n groups of map features form the feature node layer, denoted as Z n =[Z 1 , Z 2 , …, Z n , and m groups of enhanced nodes form the enhancement layer, denoted as H m =[H 1 , H 2 , …, H m ; ξ and ξ respectively represent the mapping function and the activation function. The weights W corresponding to the feature nodes and the enhancement nodes ei , and the bias β ei , are all randomly generated; the weight between the output and A is defined as W n , and the output result can be expressed as:
[0041] Y = [Z 1 ,…,Z n |H 1 ,…,H m W n = [Z n |H m W n = AW n (2)
[0042] where W n = [Z n |H m + Y = A + Y is the result that needs to be obtained through training and is the key to prediction. W n can be calculated quickly by an equation.
[0043]
[0044] λ is a scalar and is usually a constant.
[0045] Step 4: Introduce WBLS, and bring the same test samples into WBLS and SVM to obtain results. Compare with the true values through the positive prediction rate, negative prediction rate, sensitivity, specificity, accuracy, and evaluate the two groups of results;
[0046] Step 5: Introduce the Sparrow Search Algorithm (SSA)
[0047] By setting the total number of sparrow populations n, the number of discoverers, the number of sparrows aware of danger, and the maximum number of iterations iter max , the alarm threshold R 2 . Form an n×d-dimensional vector population of n sparrows to represent the positions of the sparrows, where n is the number of sparrows and represents the dimension of the variable to be optimized. Take the Mean Squared Error (MSE) as the fitness function, calculate the fitness value of each sparrow, and find the best fitness and the worst fitness values, which are defined as f g and f w . Then, according to the position update formula, the alarm value, and the safety threshold, update the positions of the discoverers, joiners, and the new positions of the sparrows aware of the danger. If the fitness value of the new position is better, update this value. Repeat the iteration to continuously update the position of the best sparrow until the number of iterations reaches iter. max , and output the optimal solution. It is expressed as follows:
[0048] S51: Parameter initialization. Set a sparrow population in the D-dimensional search space with a size of n, and use a matrix to represent the positions of the sparrows:
[0049]
[0050] At the same time, set the number of discoverers PN, the number of joiners POP - PN, the number of sparrows aware of the danger, and the maximum number of iterations iter. max , the alarm threshold R 2 , the upper bound (ub) and lower bound (lb) of the domain.
[0051] S52: Update the positions of the discoverers in each iteration using equation (5):
[0052]
[0053] Where: t is the current iteration number; j is the current dimension; is the j-th dimension position of the i-th sparrow at the t-th iteration; iter max is the maximum number of iterations; α ∈ [0, 1] is a random number; R 2 and ST are the alarm value and the safety threshold respectively. Usually, R 2 ∈ [0, 1], ST ∈ [0.5, 1]; Q is a random number following a normal distribution; L is a d-dimensional unit vector, where d is the dimension of the variable to be optimized, and each element is 1.
[0054] If R 2 < ST, that is, there are no predators nearby, and the discoverers enter the wide-area search mode. If R 2 ≥ ST, it means that there are predators near some sparrows, and all sparrows need to quickly fly to other safe areas.
[0055] S53: Update the positions of the joiners in each iteration using equation (6):
[0056]
[0057] Where: is the worst position in the entire population at the t-th iteration; is the best position of the discoverers in the (t + 1)-th iteration; A is a 1×d-dimensional matrix with the same dimension as the input, where each element is randomly assigned 1 or -1, and A+ = A T (AA T ) -1 ; n is the number of sparrows;
[0058] When it indicates that the i-th hungry joiner with a poor fitness value is most likely to starve; when it indicates that the joiner updates its position following the position of the discoverer.
[0059] S54: To be aware of danger, randomly generate the initial positions of the sentinels, accounting for 10% - 20% of the population, and update the positions of the sentinels in each iteration using the formula ():
[0060]
[0061] In the formula: is the global best position in the entire population at the t-th iteration; λ is the step size control parameter, following a normal distribution with a mean of 0 and a variance of 1; J ∈ [-1, 1] is a random number, and J represents the direction of the sparrow's movement, that is, the step size control coefficient; f i is the fitness value of the current sparrow; f g and f w are the global best fitness value and the worst fitness value; ε is defined as a particularly small constant, aiming to prevent the situation where f i - f w = 0.
[0062] When f i > f g it indicates the position change of the sparrows at the edge of the population; when f i = f g it indicates that the sparrows in the middle of the population are aware of danger and need to get closer to other sparrows.
[0063] S55: Obtain the current positions of the sparrow population. The position of each sparrow changes with the iteration of the sparrow group according to the individual historical optimal value and the global optimal value, and update the best fitness value.
[0064] S56: Repeat steps 2 to step 5 until the maximum number of iterations iter max is reached, and output the best position x best and the best fitness f g .
[0065] The proposed SSA has advantages such as stronger global search ability and faster convergence speed in high-dimensional optimization problems. However, SSA is prone to falling into local optimum in the later stage of iteration. To further improve the global optimization ability of SSA. This shortcoming provides an opportunity for the integration of other strategies with SSA, thus enriching the performance of the algorithm.
[0066] Step 6, introducing the Chaotic Sparrow Search Algorithm (CSSA)
[0067] To improve the global optimization ability of the SSA algorithm, accelerate the convergence speed, and avoid local optima, a Chaotic Sparrow Search Algorithm (CSSA) is proposed. The global search ability of the algorithm is improved by introducing improvement strategies such as logical mapping, adaptive hyperparameters, and mutation operators.
[0068] The improvement strategies of the CSSA system mainly have three aspects. First, since the initialization of the sparrow population position in the SSA algorithm is randomly generated, chaotic mapping is introduced to initialize the population position to improve the quality of the initial solution. Second, to accelerate the convergence speed and efficiency of the SSA algorithm, two adaptive hyperparameters are used to update the position of the discoverer and the number of vigilant individuals. Finally, since the population diversity will decrease during the iteration process, a mutation operator is added to increase the population diversity of the SSA, thereby obtaining the global optimal value of the SSA and overcoming local optima. The specific CSSA improvement strategies are as follows:
[0069] S61: Logical mapping
[0070] Chaotic mapping plays an important role in improving the algorithm performance. The Logistic mapping is a typical chaotic system, and the chaotic logical mapping is introduced into the SSA algorithm and described by Equation (8) as follows:
[0071] x(k + 1) = μx(k)(1 - x(k)) (8)
[0072] In the formula, μ ∈ (0, 1); x ∈ (0, 4] is the control parameter, and when 3.5699 < x ≤ 4, the system is in a chaotic state. It can be seen from Figure 9 that when x ≤ 4, the generated chaotic sequence is approximately uniformly distributed between 0 and 1, without periodicity and non-convergent. Therefore, the control parameter x should be set closer to 4.
[0073] According to the properties of logical mapping, the initial position of the sparrow population in CSSA is realized as follows:
[0074] First, normalize the initial position of the sparrow x i,j (i = 1, 2,..., N; j = 1, 2,..., D), and express it in chaotic form x i,j (0)(i = 1, 2,..., N; j = 1, 2,..., D) using Equation (9):
[0075] x i,j (0) = (x i,j - lb) / (ub - lb) (9)
[0076] Where ub represents the upper bound of x, and lb represents the lower bound of x.
[0077] Using the Logistic map, x i,j (k + 1) is obtained from Equation (10):
[0078] x i,j (k + 1) = μx i,j (k)(1 - x i,j (k)) (10)
[0079] And x i,j (k + 1) is inverse-normalized using Equation (11) to obtain the new x i,j (i = 1, 2,..., N; j = 1, 2,..., D)
[0080] x i,j = lb + x i,j (k + 1)×(ub - lb) (11)
[0081] S62: Adaptive hyper-parameters
[0082] Since the sparrow population is divided into discoverers and joiners, discoverers have more resources and determine the search direction of the population, and joiners move following the discoverers to find food. Therefore, the optimization ability of SSA mainly depends on the search range of the discoverers.
[0083] From the above formula for updating the position of the discoverer, it affects the update of the discoverer's position. And when α has a large random value, as j increases, its value range may gradually decrease from (0, 1) to approximately (0, 0.4). Therefore, an adaptive control factor (Eq. (12)) is introduced to expand the search range of the joiners as in Eq. (13):
[0084] w = w 0 ×c t (12)
[0085] where, w 0 is the initial weight, c is the adaptive factor of w, set according to the actual problem, and t is the current iteration number. Among them, the value of w 0 is 1, and c is set to 0.9 to keep the value of w small, thereby expanding the search range of the discoverer and enhancing the global search ability of the algorithm.
[0086]
[0087] In addition, in order to avoid predators during foraging, the initial positions of the vigilant individuals are randomly generated, accounting for 10% - 20% of the population. The presence of vigilant individuals can help the sparrow population better solve the SSA problem. When the number of vigilant individuals (SN) is large, it is beneficial to improve the global optimization ability of sparrows. However, when SN is small, it is beneficial to accelerate the convergence speed of the algorithm. Therefore, this paper proposes an adaptive update formula for the number of vigilant individuals, as shown in Eq. (14), which can non-linearly decrease during the iteration process.
[0088]
[0089] In the formula, SNmax is the maximum value of the number of vigilant individuals; SNmin represents the minimum value of the number of vigilant individuals; the Round function is used for rounding; t represents the current iteration number; iter max represents the maximum number of iterations.
[0090] S63: Mutation operator
[0091] During the search process of the SSA algorithm, the diversity of the sparrow population will decrease, which will cause the algorithm to converge prematurely. Therefore, in order to avoid the local optimum of SSA, a mutation operator is introduced to increase the population diversity. The specific operation is as follows:
[0092] x i = lb + (ub - lb) × rand(1, D) (15)
[0093] where i represents a random individual in each iteration process, and the position of sparrow i will be initialized by using the above mutation formula.
[0094] Example:
[0095] Guozhuang Spring is located in the western part of Shanxi Province, China, and is inland. Due to the shielding effect of the mountainous area in the eastern part of the province, the influence of the marine climate on the study area is inhibited, so it presents semi-arid climate characteristics. According to the data collected by county meteorological stations, the annual average rainfall in this area is 468.4 mm, mainly concentrated from June to August.
[0096] Step 1) Compile a groundwater inventory map
[0097] In this study, based on historical groundwater information, high-resolution remote sensing images, and a large number of field observation data, a groundwater sample distribution map was drawn. A total of 122 groundwater springs were identified and are represented as red triangles in Figure 5 . At the same time, 1220 reasonable points were selected in the study area as negative samples (points not used to predict groundwater, i.e., non-springs), and in Figure 5It is represented as a black five-pointed star in the figure. Using GIS, these spring points and non-spring points are divided into two parts, and 70% of the points are randomly selected for model training, and the remaining 30% of the points are used for model verification.
[0098] Step 2) Selection and evaluation of groundwater influencing factors
[0099] The evaluation and selection of groundwater spring influencing factors is a key link in groundwater potential evaluation. Random forest can analyze complex interdependent features, and has the robustness to handle missing and noisy data and a fast learning speed. Therefore, this study selects random forest for feature selection of spring influencing factors. Random forest provides variable importance (VI) scores during the data analysis process, which are obtained by summing the Gini indices of each evaluation factor at all nodes of the entire tree and then normalizing. A VI value greater than 0 indicates that the factor contributes to the prediction, and the higher the VI value, the greater the contribution.
[0100] The importance scores of all 11 influencing factors obtained from the random forest are shown in Table 1. The results show that the most significant groundwater influencing factors are, in order: altitude, distance to road, distance to river, and Normalized Difference Vegetation Index (NDVI). The importance scores of these factors are: altitude (1.8757), distance to road (1.0012), distance to river (0.8168), NDVI (0.8118), aspect (0.7377), Topographic Wetness Index (TWI) (0.6860), Stream Topographic Index (STI) (0.5837), Stream Power Index (SPI) (0.4876), slope (0.2122), profile curvature (0.1974), and plan curvature (0.1831). The importance scores of all considered spring influencing factors are greater than 0, indicating that the selected 11 factors all contribute significantly to the prediction results of the model.
[0101] Table 1. Influencing factors
[0102] Influencing factor VI Elevation 1.8757 NDVI 0.8118 Distance to river 0.8168 Distance to road 1.0012 Aspect 0.7377 Slope 0.2122 SPI 0.4876 STI 0.5837 TWI 0.6860 Plan curvature 0.1831 Profile curvature 0.1974
[0103] Step 3) Construct the WBLS model
[0104] Since WBLS has a considerable processing ability for imbalanced data, it has been increasingly applied in the period after its emergence. The present invention uses the WBLS method to process the imbalanced samples of groundwater to achieve accurate classification. Among them, the algorithm principle is mainly based on the fact that in the imbalanced classification problem, the number of samples in one class is significantly larger than that in other classes. In the present invention, the class with the number of samples lower than the average is regarded as the minority class, and the class with the number of samples higher than the average is regarded as the majority class. To handle this phenomenon well, we should treat different classes differently, which means that different weight values should be assigned to different classes. Among them, the weight value of the majority class is smaller, and the weight value of the minority class is larger. The proposed weighted scheme is designed as follows: This process is generated by evaluating the number of groundwater samples. To further implement this process, the result of whether groundwater exists in the data is divided into the minority class and the majority class. Based on this consideration, weights are assigned to different classes, so that the weights assigned to the minority class are large and the weights of the majority class are small. The design scheme of the weighting is shown in Equation (16):
[0105]
[0106] In the formula, V i is the dummy number of x i , W ii is the weight value obtained by synthesizing the samples of X i , |P j | is the number of samples belonging to the class P j , and j = 1, 2,..., N. Here, the dummy number V i of the input data is obtained by sorting the samples of each water source category in ascending and descending order. Finally, the descending samples of each water source category are replaced with the ascending samples. Therefore, the weights in the diagonal matrix within each water source sample category updated by the above formula are W = diag{W}. A specific example is demonstrated as follows: In a binary classification task (i.e., j = 2), there are five samples in the dataset, where there is only one sample in the minority class and the other samples belong to the majority class.
[0107] First, the samples of each class are sorted in ascending order as (1, 4), and then the samples of each class are sorted in descending order (4, 1). Finally, we replace (1, 4) with (4, 1) obtained in descending order to obtain the dummy number set (4, 1) of the samples of each class, which means that in the binary classification task, the first dummy number 4 is assigned to each sample of the minority class, and the other dummy number 1 is assigned to each sample of the majority class. Based on the above formula, we have:
[0108]
[0109] Finally, the diagonal matrix (W) is introduced using the standard BLS to obtain the desired WBLS as follows:
[0110]
[0111] W n = (λI + AW 2 A T ) -1 A T W 2 Y(19)
[0112] Similar to BLS, WBLS first constructs feature nodes using the input data, then maps the feature nodes to enhanced nodes using the weight W obtained from Equation (19), and finally calculates the output weight W n .
[0113] Compared with the traditional BLS, the present invention does not require an additional penalty function to calculate the weight when dealing with data sample imbalance. Instead, it calculates the weight according to the class information, further reducing the training error. Since the weight is calculated for the corresponding class through the above formula according to the sample distribution. Therefore, the majority class obtains a relatively small weight, while a larger weight is assigned to the minority class. Although this model solves the problem of data imbalance to a certain extent, this model still faces difficulties such as redundancy of the majority class and oversampling weight, which in turn reduces the classification accuracy of the model. Therefore, in order to overcome these difficulties, a powerful optimization technique is proposed to optimize the model
[0114] The basic structure of BLS is as Figure 1 shown
[0115] Step 4) Implementation of CSSA-WBLS
[0116] In the WBLS model, the feature nodes and enhanced nodes are the basic components in BLS. Adjusting the number of feature nodes N1, the number of feature mapping groups N2, and the number of enhanced nodes N3 can affect the complexity and learning ability of the model, and thus affect the performance of the model. However, the parameters N1, N2, and N3 are restricted by a preset range. In order to improve the efficiency and accuracy of WBLS, the present invention uses the proposed CSSA to optimize the number of feature nodes N1, the number of feature mapping groups N2, and the number of enhanced nodes N3 of the WBLS model, so as to obtain CSSA-WBLS. The present invention describes CSSA-WBLS in detail, mainly including two aspects: the construction of WBLS and parameter search. The process is as Figure 4 shown
[0117] Step 5) Model performance evaluation
[0118] Evaluation based on statistical metrics
[0119] In groundwater potential mapping, the accuracy of the assessment model is crucial for the reliability of the generated maps. However, in an imbalanced dataset, the number of majority-class samples is often much larger than that of minority-class samples, which may result in a relatively high overall precision even if the classification results of minority-class samples are poor. Therefore, this study adopted multiple metrics to evaluate the model performance on imbalanced datasets, including precision, recall, specificity, balanced accuracy, F1-score, and the area under the receiver operating characteristic curve (AUC or ROC). In addition, the present invention also uses a confusion matrix to represent the classification performance results of machine learning models.
[0120] The confusion matrix is commonly used with four elements, namely true positive (TP), true negative (TN), false positive (FP), and false negative (FN), to evaluate the performance of a classification model. TP represents the correctly classified spring samples, TN represents the correctly classified non-spring samples, FP represents the non-spring samples misclassified as springs, and FN represents the spring samples misclassified as non-springs.
[0121] Precision and recall measure the proportion of positive samples correctly identified and the proportion of correctly identified positive samples among those that should be predicted as positive samples, respectively. Specificity represents the proportion of negative samples correctly identified. The calculation formulas are as follows:
[0122]
[0123] In the imbalanced data classification task, since the classifier may be biased towards majority-class samples, using accuracy alone may not fairly evaluate the model's performance. Therefore, this study introduced balanced accuracy, which better reflects the model's recognition performance on imbalanced datasets by calculating the arithmetic mean of recall and specificity and the F1-score. The formulas are as follows:
[0124]
[0125] The F1-score is one of the important metrics for imbalanced data evaluation. It evaluates the accuracy of each category by calculating the harmonic mean of precision and recall. The higher the F1-score, the more effective the test method. The calculation formula is:
[0126]
[0127] P is the accuracy rate, and R is the recall rate.
[0128] The ROC curve evaluates the predictive ability of a model by plotting the False Positive Rate (FPR) on the X-axis and the True Positive Rate (TPR) on the Y-axis. The AUC, which is the area under the ROC curve, is an important metric for measuring the overall performance of the model. The formula is as follows:
[0129]
[0130] In the formula, P is the total number of positive samples, and N is the total number of negative samples.
[0131] Verification of the groundwater potential evaluation map:
[0132] In this study, we focused on evaluating the classification effect of imbalanced and balanced data, and calculated various evaluation metrics such as precision, recall, specificity, balanced accuracy, F1-score, and AUC. Since the deviation of the sample number distribution in the imbalanced dataset is often relatively serious, we specifically considered the metrics of balanced accuracy, F1-score, and AUC to comprehensively evaluate the performance of the model.
[0133] During the analysis process, we found that when using Support Vector Machine (SVM) and Broad Learning System (BLS) for classification, the BLS model outperformed the SVM in all metrics, especially showing higher performance in predicting imbalanced spring samples. This result highlights the potential of the BLS model in dealing with class-imbalanced data, so we decided to further optimize the BLS model.
[0134] Table 2. Index evaluation results
[0135]
[0136] As shown in Table 2 and Figure 8As shown, we compared the performance of three models: BLS, SSA-optimized BLS (SSA-BLS), and chaotic sparrow swarm-optimized BLS (CSSA-BLS). The optimized CSSA-BLS model performed the best, with a balanced accuracy of 0.8116 and an F1 value of 0.8777. Followed by SSA-BLS, with a balanced accuracy of 0.7942 and an F1 value of 0.8754, while the unoptimized BLS model had a balanced accuracy of 0.7645 and an F1 value of 0.8641. This comparison shows that the parameter optimization of the BLS model by CSSA significantly improved the prediction performance. Further, to optimize the handling of imbalanced data, we introduced the weighted BLS model (WBLS) and optimized it using SSA and CSSA. The results showed that CSSA-optimized WBLS (CSSA-WBLS) had the best performance on the test set, with a balanced accuracy of 0.8645 and an F1 value of 0.8904. In addition, the balanced accuracy of SSA-WBLS was 0.8471 and the F1 value was 0.8884, while the balanced accuracy of WBLS itself was 0.8401 and the F1 value was 0.8751.
[0137] As Figure 6 shown in the ROC-AUC curve evaluation, the CSSA-WBLS had the highest AUC value, reaching 0.8737, followed by SSA-WBLS and WBLS, with values of 0.9484 and 0.9479, respectively. These results further verified the efficiency and accuracy of the CSSA-WBLS model in handling imbalanced data.
[0138] The conclusions of this embodiment can be summarized as follows: 1. Analysis of geographical and environmental characteristics: According to the frequency ratio (FR) method, springs are mainly distributed under specific geographical and environmental conditions, including altitude (1400 - 1600 meters), steep slopes (greater than 25 degrees), planar and profile curvatures with a slightly concave shape, high stream power index (SPI greater than 15), high sediment transport index (STI greater than 8), south slope direction, medium to high terrain wetness index (TWI between 6 - 8), moderate to high vegetation coverage (NDVI between 0.59 and 0.74), and moderate distances from roads and rivers (1000 - 1500 meters from roads, 500 - 1000 meters from rivers). 2. Analysis of key influencing factors: Through variable importance (VI) values and collinearity analysis, it is confirmed that the factor with the most significant impact on groundwater springs is altitude, followed by the distance to roads, the distance to rivers, NDVI, slope aspect, TWI, STI, SPI, slope, profile curvature, and planar curvature. The importance of these factors is all greater than 0, indicating that they make significant contributions to the model's prediction of the location of groundwater springs. 3. Model performance evaluation: This study used multiple evaluation indicators such as AUC, balanced accuracy, and F1 value to compare the performance of SVM, BLS, WBLS, SSA - BLS, CSSA - BLS, SSA - WBLS, and CSSA - WBLS models. The results show that the CSSA - optimized WBLS model performs excellently in dealing with imbalanced datasets, with an AUC value as high as 0.8737, demonstrating its efficiency and accuracy in predicting the distribution of groundwater springs.
Claims
1. A groundwater potential prediction method based on a generalized width learning system, characterized in that The steps include: S1: Data collection and data processing: The machine learning model was combined with the statistical model, and various data were processed using ENVI and ArcGIS software to construct a groundwater potential map; S2: Building a BLS model Generate a feature node for each window, randomly generate a bias value and a corresponding weight matrix based on the data matrix obtained in step S1, place the bias value in the last column of the data matrix, and multiply the two matrices to obtain each feature node of the model; secondly, generate an enhanced node, introduce the enhanced node to enhance the nonlinear factors in the network, and assign the activation function tansig to the previously generated feature node, so that the model can better handle complex geographic information, thereby having a better prediction effect on groundwater; finally, according to the obtained feature nodes, enhanced nodes and target output, the pseudo-inverse is obtained, that is, the mapping from input to output, so that the groundwater potential of the region can be evaluated; S3: Introducing WBLS The learning process is optimized by assigning different weights to the training samples, thus being used for potential prediction of unbalanced groundwater data; S4: Establishing the Chaotic Sparrow Algorithm CSSA Model In order to improve the global optimization ability of SSA, the Logistic mapping, adaptive hyperparameters and mutation operators are combined with the SSA algorithm to establish CSSA. S5: Building the CSSA-WBLS model The present invention uses the chaotic sparrow algorithm CSSA to optimize the parameters of WBLS, thereby obtaining CSSA-WBLS, thereby improving the efficiency and accuracy of WBLS.
2. The groundwater potential prediction method based on the generalized width learning system according to claim 1 is characterized by: Step S1: The specific operation methods of data collection and data processing are as follows: Based on the existing data in the study area, remote sensing images, DEM images, data on the boundaries of the study area and spring data in the study area were collected; through processing with ENVI and ArcGIS software, the groundwater potential evaluation factor map of the study area was obtained, and a data set for model training and evaluation was constructed based on the evaluation factors. The data set consists of eleven groundwater potential evaluation factors: slope, altitude, aspect, plane curvature, profile curvature, standardized precipitation index, sediment transport index, terrain humidity index, normalized vegetation index, road distance, and river distance.
3. The groundwater potential prediction method based on the generalized width learning system according to claim 1 is characterized by: Step S3: The specific operation method of introducing WBLS is: determine the weights based on category information, thereby significantly reducing the training error; through this strategy, the distribution of samples directly affects the weight allocation, so that the majority class obtains a relatively small weight, while a larger weight is allocated to the minority class.
4. The groundwater potential prediction method based on the generalized width learning system according to claim 1 is characterized by: The specific operation method of establishing the CSSA-WBLS model is as follows: S1: Initialize CSSA-WBLS Initialize the population size pop, the number of discoverers PN, the number of joiners POP-PN, the number of guards SD, and the maximum number of iterations iter max , alarm threshold R2, dimension D, upper and lower bounds ub and lb of the parameter, control parameter μ, tolerance ε; Since it is necessary to select the number of feature nodes N1, the number of feature mapping groups N2 and the number of enhancement nodes N3, generate the initial value x i =(x i , N1, x i,N2 , x i,N3 )(i=1,2,...,pop), where x i,N1 , x i,N2 With, x i,N3 Respectively represent N1, N2, N3; S2: Logical mapping and transformation By the following formula, x i =(x i,N1 , x i,N2 )(i=1,2,...,pop) performs initial logical mapping and transformation; x i,j (0)=(x i,j -lb) / (ub-lb) x i,j (k+1)=μx i,j (k)(1-x) i,j (k)) x i,j =lb+x i,j (k+1)x(ub-lb); S3: Build and train the BLS model Use N1, N2 and N3 obtained according to the formula as initial parameters to construct and train the BLS model; S4: Update the number of feature nodes N1, the number of feature mapping groups N2 and the number of enhanced nodes N3 The number of feature nodes N1, the number of feature mapping groups N2 and the number of enhanced nodes N3 are updated by the following formula: Where t is the current number of iterations; j is the current dimension; is the j-th dimension position of the i-th sparrow at the t-th iteration; iter max is the maximum number of iterations; α∈[0,1] is a random number; R2 and ST are the alarm value and safety threshold, respectively, usually R2∈[0,1], ST∈[0.5,1]; Q is a random number that follows a normal distribution; L is a d-dimensional unit vector, where d is the dimension of the variable to be optimized, and each element is 1; is the worst position in the entire population at the tth iteration; is the best position of the finder in the t+1th iteration; A is a 1×d-dimensional matrix with the same dimension as the input, in which each element is randomly assigned 1 or -1, and A + =A T (AA T ) -1 ; n is the number of sparrows; is the global optimal position in the whole population at the tth iteration; λ is the step length control parameter, which obeys the normal distribution with mean 0 and variance 1; J∈[-1,1] is a random number, J represents the direction of the sparrow's movement, that is, the step length control coefficient; f i is the fitness value of the current sparrow; f g and f w is the global best fitness value and the worst fitness value; ε is defined as a very small constant in order to prevent f i -f w =0; w is the adaptive control factor; S5: Calculate fitness value The mean square error (MSE) is selected as the fitness function to evaluate CSSA-WBLS. The fitness value of each individual can be calculated by the following formula: In the formula, yi represents the actual value of sample i; represents the output value of sample i; n is the total number of samples; S6: Use the following formula to perform mutation operation x i =lb+(ub-lb)×rand(1,D) S7: Terminal conditions, get the final parameters: When CSSA-WBLS reaches the maximum number of iterations iter max When , the optimal position is output by the following formula And the best fitness f g , get the optimal N1, N2; Otherwise, repeat steps S2 to S7 until the termination condition is met; S8: Output results Use the BLS global optimal parameters selected in step S7 To construct and train WBLS, and output the final training and testing results.