Geological disaster monitoring data processing method and system based on reinforcement learning
By constructing a state representation model and reinforcement learning strategy network, the problem that traditional methods are difficult to adapt to the dynamic changes of geological disasters is solved, more efficient and accurate geological disaster monitoring data processing is achieved, and early warning capabilities are improved.
Patent Information
- Application Number
- CN202511134448.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Traditional geological disaster monitoring methods rely on fixed rules and empirical models, which are difficult to adapt to the complexity and dynamic changes of geological disasters, resulting in poor data processing effects and affecting the accuracy and timeliness of early warnings.
Construct a state representation model for geological hazard monitoring data, convert the data into a state vector recognizable by the reinforcement learning agent, initialize the reinforcement learning policy network, generate a set of processing actions through the policy evaluation and improvement module, and iteratively optimize the policy parameters according to the feedback signal to adapt to the dynamic changes of the geological hazard monitoring environment.
It has improved the accuracy and flexibility of geological disaster monitoring data processing, enhanced the intelligence level of geological disaster monitoring, and enhanced the effectiveness of early warning.
Smart Images

Figure CN120632432A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of reinforcement learning technology, and in particular to a geological disaster monitoring data processing method and system based on reinforcement learning. Background Art
[0002] In the field of geological hazard monitoring, with the continuous advancement of monitoring technology, a large amount of geological hazard monitoring data is continuously collected. However, how to efficiently and accurately process this data to achieve effective geological hazard early warning and assessment is a major challenge. Traditional methods often rely on preset fixed rules and empirical models to process monitoring data. For example, fixed filtering parameters are used to denoise the data, data features are extracted according to established feature extraction methods, and anomalies are determined based on pre-set thresholds. However, the occurrence of geological hazards is complex and uncertain. The characteristics of monitoring data vary significantly across different regions and types of geological hazards. Moreover, the geological environment changes over time, causing the distribution and characteristics of monitoring data to change accordingly. Traditional fixed rules and empirical models are difficult to adapt to these dynamic changes and cannot fully tap the potential information in the monitoring data. This leads to poor data processing results and affects the accuracy and timeliness of geological hazard early warning. Summary of the Invention
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, the present invention provides a geological disaster monitoring data processing method based on reinforcement learning, the method comprising: Constructing a state representation model for geological hazard monitoring data, wherein the state representation model is used to convert the continuously collected geological hazard monitoring data into a state vector recognizable by a reinforcement learning agent, wherein the state vector includes sequence correlation features and spatial correlation features of the geological hazard monitoring data; Initializing a reinforcement learning policy network, wherein the reinforcement learning policy network includes a policy evaluation module and a policy improvement module, wherein the policy evaluation module is used to calculate an estimated value function of the current policy, and the policy improvement module is used to adjust policy parameters based on the estimated value function; Executing a strategy selection operation based on the state vector and the reinforcement learning strategy network to generate a processing action set for geological disaster monitoring data, wherein the processing action set includes a data filtering operation, a feature enhancement operation, and an anomaly recognition operation; Obtaining a feedback signal from the geological hazard monitoring environment on the processing action set, and generating a reward signal based on the feedback signal and the estimated value of the value function, wherein the reward signal is used to evaluate the impact of the processing action set on the processing effect of the geological hazard monitoring data; Iteratively optimize the policy parameters of the reinforcement learning policy network according to the reward signal and the state vector until the value function estimate output by the policy evaluation module converges to a preset stable range.
[0004] On the other hand, the present invention also provides a geological disaster monitoring data processing system based on reinforcement learning, including a processor and a machine-readable storage medium, the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0005] Based on the above aspects, the present invention converts the continuously collected geological disaster monitoring data into a state vector containing serial correlation features and spatial correlation features that can be recognized by the reinforcement learning agent by constructing a state representation model, so that the agent can fully and accurately understand the intrinsic characteristics and mutual relationships of the monitoring data, initialize the reinforcement learning policy network including the policy evaluation module and the policy improvement module, realize the evaluation of the current policy value and the adjustment of the policy parameters based on the evaluation results, and provide the agent with the ability to autonomously optimize the policy. Based on the state vector and the policy network, a set of processing actions including operations such as data filtering, feature enhancement and anomaly recognition is generated. It can flexibly select the appropriate processing method according to the actual situation of the monitoring data, obtain environmental feedback signals and generate reward signals for evaluating the effect of the processing actions, and it iteratively optimizes the policy parameters according to the reward signal and the state vector until the estimated value of the value function converges, so that the policy network can continuously adapt to the dynamic changes of the geological disaster monitoring environment, continuously improve the effect of data processing, and thus significantly improve the accuracy, flexibility and intelligence level of geological disaster monitoring data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 It is a schematic diagram of the execution flow of the geological disaster monitoring data processing method based on reinforcement learning provided by an embodiment of the present invention.
[0007] Figure 2 Schematic diagram of exemplary hardware and software components of a geological disaster monitoring data processing system based on reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0008] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a geological disaster monitoring data processing method based on reinforcement learning provided by an embodiment of the present invention. The geological disaster monitoring data processing method based on reinforcement learning is introduced in detail below.
[0009] Step S110: Construct a state representation model for geological hazard monitoring data, which is used to convert the continuously collected geological hazard monitoring data into a state vector that can be recognized by the reinforcement learning agent. This state vector contains the serial correlation characteristics and spatial correlation characteristics of the geological hazard monitoring data.
[0010] In actual geological disaster monitoring, continuous monitoring of the geological environment generates massive amounts of continuous data. For example, in a mountainous area prone to debris flows and landslides, a large number of monitoring devices are deployed. These include high-precision displacement sensors for measuring mountain surface displacement. These sensors can accurately record the slightest movement of various parts of the mountain at different times; soil moisture sensors, which can measure the moisture content in the soil in real time, as changes in soil moisture are often closely related to disasters such as landslides; and groundwater level monitors, which monitor the rise and fall of groundwater levels. Abnormal fluctuations in groundwater levels can be a key signal of a geological disaster.
[0011] However, this raw continuous monitoring data is disorganized and cannot be directly analyzed and utilized by reinforcement learning agents. Therefore, a state representation model is needed. The core goal of this state representation model is to systematically process and transform this complex continuous data into a state vector that the reinforcement learning agent can understand and manipulate. The serial correlation characteristics in the state vector can reflect the temporal changes in geological hazard monitoring data. For example, whether mountain displacement data increases steadily over a period of time or exhibits periodic fluctuations is important for predicting the development trend of geological hazards. Spatial correlation characteristics reflect the interrelationships between data from monitoring points at different locations. For example, whether soil moisture at adjacent monitoring points affects each other, or whether changes in groundwater levels in one area trigger chain reactions in surrounding areas, can help to fully understand the spatial spread and development of geological hazards.
[0012] Step S111: performing time window division processing on the continuously collected geological disaster monitoring data to obtain a plurality of monitoring data windows with a time sequence, each monitoring data window containing geological disaster monitoring data within a preset time length.
[0013] In the mountainous and disaster-prone areas mentioned above, the collected geological disaster monitoring data is accumulated continuously in chronological order. In order to facilitate subsequent analysis and processing, these continuous data need to be divided into time windows. The determination of the preset time length requires comprehensive consideration of many factors. If the changes in geological disasters are relatively slow, such as the slow creep of the mountain, then a longer time length can be selected, such as one day or one week. Thus, each monitoring data window contains all the geological disaster monitoring data collected by each monitoring device within this day or week, such as the total displacement of the mountain during this period, the average change in soil moisture, and the overall rise and fall trend of the groundwater level.
[0014] Conversely, if a geological disaster is changing rapidly, such as a sudden mudslide or earthquake-induced landslide, a shorter timeframe, perhaps an hour or even less, is required. Each monitoring data window contains detailed monitoring data from this shorter period, enabling more timely and accurate capture of subtle changes before a geological disaster occurs. By using this time window division approach, complex continuous data is discretized into monitoring data windows with clear time boundaries, making subsequent data analysis more organized and targeted.
[0015] Step S112: extract the numerical change sequence of the geological disaster monitoring data in each monitoring data window, and calculate the correlation parameter of the numerical change sequence between adjacent monitoring data windows. The correlation parameter is used to characterize the correlation degree of the monitoring data change trend in different time windows.
[0016] For each monitoring data window, we need to extract the numerical change sequence of the geological disaster monitoring data. For example, within a monitoring data window, the displacement sensor collects data at set intervals. By sequentially extracting these chronologically arranged mountain displacement values, we form the numerical change sequence of the mountain displacement within that window. Similarly, we process data such as soil moisture and groundwater level to obtain their respective numerical change sequences.
[0017] Next, the correlation parameter of the numerical change sequences between adjacent monitoring data windows is calculated. This correlation parameter can be calculated using various methods. A common approach is based on similarity. The numerical change sequences of mountain displacement within two adjacent monitoring data windows are considered as two data sets. First, the overall magnitude of change in each data set is calculated. The magnitude of change can be measured by calculating the difference between the maximum and minimum values in the data set. The magnitude of these two magnitudes is then compared, and the direction of change in the two data sets is observed to be consistent. If the magnitude of change in mountain displacement within two adjacent windows is similar and the direction of change is the same, then their trends are highly consistent, and the correlation parameter will be relatively large. Conversely, if the magnitudes of change differ significantly or the directions of change are opposite, the correlation parameter will be relatively small. The same calculation method is used to obtain the correlation parameters for numerical change sequences of other data, such as soil moisture and groundwater level. By calculating the correlation parameter, the degree of correlation between the changing trends of monitoring data within different time windows can be effectively determined.
[0018] Step S113: performing serial correlation analysis on multiple monitoring data windows based on the above correlation parameters to generate serial correlation features of geological disaster monitoring data, which include trend continuity indicators of numerical change sequences and distribution density of mutation points.
[0019] Step S1131: Sort multiple monitoring data windows in chronological order to obtain an ordered monitoring data window sequence, and calculate the sequence correlation strength of adjacent windows in the ordered monitoring data window sequence based on the above correlation parameter. The sequence correlation strength is positively correlated with the numerical value of the correlation parameter.
[0020] After obtaining multiple monitoring data windows, they need to be sorted according to their corresponding chronological order to obtain an ordered sequence of monitoring data windows. The sorting process is like rearranging a bunch of messy time segments along a timeline, thus ensuring that subsequent analysis can be carried out according to the continuity of time.
[0021] The sequential correlation strength between adjacent windows in the ordered monitoring data window sequence is calculated based on the previously calculated correlation parameter. Since the correlation parameter reflects the degree of similarity in the monitoring data trends within adjacent windows, the sequential correlation strength is positively correlated with the value of the correlation parameter. A large correlation parameter indicates that the trends in the monitoring data within two adjacent monitoring data windows are very similar, and the corresponding sequential correlation strength is high. For example, if the correlation parameter for the mountain displacement values in two adjacent windows is large, it means that the trends in the mountain displacement during these two adjacent time periods are almost identical, perhaps both showing a continuous increase or decrease. Therefore, the sequential correlation strength between them will be correspondingly high. Conversely, when the correlation parameter is small, the data trends within adjacent windows differ significantly, and the sequential correlation strength is low. For example, if the mountain displacement in one window is increasing while it is decreasing in the adjacent window, the sequential correlation strength will be very low. By calculating the sequential correlation strength, the degree of correlation between adjacent windows in the ordered monitoring data window sequence can be further quantified.
[0022] Step S1132: Extract window segments in the ordered monitoring data window sequence whose sequence correlation strength is continuously higher than a preset correlation threshold, mark the window segments as trend continuation segments, and calculate the length ratio of the trend continuation segments as a trend continuity indicator. The length ratio is the ratio of the number of windows contained in the trend continuation segments to the total number of windows in the ordered monitoring data window sequence.
[0023] After obtaining the sequence correlation strength of adjacent windows in the ordered monitoring data window sequence, a preset correlation threshold needs to be set. The determination of this preset correlation threshold requires combining actual experience and historical data in geological hazard monitoring. Its role is to determine which windows have a correlation degree that can be considered as a continuity standard for the data change trend.
[0024] Next, we traverse the ordered monitoring data window sequence and extract window segments whose sequence correlation strength is continuously higher than the preset correlation threshold. These window segments represent that the trend of change in geological disaster monitoring data has a strong continuity over a continuous period of time. For example, for the ordered monitoring data window sequence of groundwater level monitoring data, if there is a continuous window where the sequence correlation strength between adjacent windows is higher than the preset correlation threshold, it means that the change trend of the groundwater level is continuously stable during this period of time, and may be in an upward or downward state. These window segments that meet the conditions are marked as trend continuation segments.
[0025] Then, the length ratio of the trend continuation segment is calculated as the trend continuity index. The specific calculation method is to count the number of windows contained in the trend continuation segment, then count the total number of windows in the ordered monitoring data window sequence, and divide the number of windows contained in the trend continuation segment by the total number of windows. The ratio obtained is the trend continuity index. The trend continuity index can intuitively reflect the trend continuity of geological disaster monitoring data over time. The larger the value of this indicator, the stronger the continuity of the data change trend, which is of great significance for predicting the development trend of geological disasters. If the trend continuity index is high and the monitoring data shows a trend that is not conducive to geological stability, such as a continuous increase in mountain displacement or a continuous rise in groundwater levels, then it is necessary to increase vigilance against the occurrence of geological disasters.
[0026] Step S1133: Perform first-order difference processing on the numerical change sequence of each monitoring data window to obtain a difference sequence, and determine the potential mutation point based on the point where the absolute value of the difference sequence exceeds a preset difference threshold.
[0027] For each monitoring data window, a first-order difference is performed to generate a differential sequence. First-order difference processing involves calculating the difference between two consecutive values within a numerical change sequence. For example, within a monitoring data window, a displacement sensor records mountain displacement values at different times. The difference between the displacement value at the previous moment and the displacement value at the next moment is the result of the first-order difference. Arranging these differences in chronological order forms a differential sequence.
[0028] After obtaining the difference sequence, a preset difference threshold needs to be set. The determination of this preset difference threshold should take into account the normal fluctuation range of geological hazard monitoring data and possible abnormal changes. Generally speaking, by analyzing historical monitoring data, the value range of the difference sequence when the data fluctuates normally can be found. Based on this, the threshold can be appropriately increased to ensure that potential mutation points can be accurately identified.
[0029] Potential mutation points are identified based on the difference sequence. When the absolute value of a value in the difference sequence exceeds a preset difference threshold, the time point corresponding to that value is identified as a potential mutation point. Potential mutation points may indicate an abnormal change in geological hazard monitoring data. For example, in a difference sequence of mountain displacement, if the absolute value of a difference far exceeds the preset difference threshold, this may indicate a sudden and significant displacement of the mountain at that moment, potentially signaling an impending geological disaster such as a landslide. By identifying potential mutation points, anomalies in geological hazard monitoring data can be promptly detected.
[0030] Step S1134: Count the total number of potential mutation points in the ordered monitoring data window sequence, and calculate the ratio of the total number of potential mutation points to the total number of windows in the ordered monitoring data window sequence as the mutation point distribution density.
[0031] After determining the potential mutation points within each monitoring data window, we need to count the total number of potential mutation points in the ordered monitoring data window sequence. This can be done by traversing the entire ordered monitoring data window sequence, counting the potential mutation points within each window, and then summing the number of potential mutation points across all windows to obtain the total number of potential mutation points.
[0032] Next, the ratio of the total number of potential mutation points to the total number of windows in the ordered monitoring data window sequence is calculated, and this ratio is used as the mutation point distribution density. The mutation point distribution density reflects the temporal mutation of geological hazard monitoring data. If the mutation point distribution density is large, it means that the geological hazard monitoring data has a high frequency of mutations throughout the entire monitoring period. This may indicate that the geological environment in the area is unstable and there is a high risk of geological hazards. For example, if the mutation point distribution density of data such as mountain displacement and groundwater level is relatively high over a period of time, then the area may be in a period of relatively active geological activity, and the monitoring and prevention of geological hazards needs to be strengthened. Conversely, if the mutation point distribution density is small, it means that the data is relatively stable and the possibility of geological hazards is relatively low.
[0033] Step S1135: performing feature normalization processing on the trend continuity index and the mutation point distribution density, and splicing the normalized trend continuity index and mutation point distribution density into the sequence correlation feature of the geological disaster monitoring data.
[0034] After obtaining the trend continuity index and the mutation point distribution density, in order to better integrate these two indicators and form a unified sequence correlation feature, they need to be normalized. The purpose of feature normalization is to unify feature data of different ranges and scales to make them comparable.
[0035] Trend continuity indicators and mutation point distribution density may have different value ranges. For example, the trend continuity indicator may range from 0 to 1, while the mutation point distribution density value range may vary depending on the specific circumstances of the monitored data. Through standardization, they can be mapped to a relatively uniform scale. A common standardization method is to normalize the data, that is, to convert the data to the range of 0 to 1. By calculating the difference between each indicator value and the maximum and minimum values of the indicator, the indicator values can be linearly transformed to fall within the range of 0 to 1.
[0036] The standardized trend continuity index and mutation point distribution density are then combined to form a serial correlation feature of the geological hazard monitoring data. This serial correlation feature integrates the temporal trend continuity and mutation of the geological hazard monitoring data. In the subsequent reinforcement learning process, this serial correlation feature can be input into the reinforcement learning agent as part of the state vector, helping the agent better understand the temporal characteristics of the geological hazard monitoring data and make more accurate decisions.
[0037] Step S114: Divide the geological disaster monitoring data in each monitoring data window into spatial dimensions to obtain monitoring data subsets of multiple spatial sub-regions, and calculate the correlation coefficient between the monitoring data subsets of different spatial sub-regions. The correlation coefficient is used to characterize the degree of mutual influence of geological disaster monitoring data between spatial sub-regions.
[0038] In geological disaster monitoring, in addition to analyzing the temporal characteristics of the data, its spatial correlation also needs to be considered. For the geological disaster monitoring data within each monitoring data window, it is divided according to the spatial dimension to obtain monitoring data subsets of multiple spatial sub-regions. Taking mountainous and disaster-prone areas as an example, the monitoring equipment in the area is distributed in different geographical locations. The entire monitoring area can be divided into multiple spatial sub-regions based on the division of geographical areas, such as different areas of the mountains, different locations of the valleys, etc. The data collected by the monitoring equipment in each spatial sub-region constitutes a monitoring data subset.
[0039] Calculate the correlation coefficient between the monitoring data subsets of different spatial sub-regions. There are many ways to calculate the correlation coefficient. A commonly used method is based on the calculation of covariance. Taking mountain displacement data as an example, for the monitoring data subsets of two different spatial sub-regions, first calculate the mean of the two subsets respectively. Then, for each data point in each subset, calculate its difference with the mean of the subset. Next, multiply the differences between the corresponding data points in the two subsets, and sum all the products to obtain the numerator of the covariance. Then calculate the sum of the squares of the differences between the two subsets respectively, multiply the two square sums and take the square root to obtain the denominator of the covariance. Divide the numerator by the denominator to obtain the correlation coefficient of the mountain displacement data of the two spatial sub-regions. For the monitoring data subsets of other data such as soil moisture and groundwater level, the same calculation method is used to obtain their respective correlation coefficients.
[0040] The correlation coefficient is used to characterize the degree of mutual influence between geohazard monitoring data in spatial subregions. A large and positive correlation coefficient indicates a strong positive correlation between the monitoring data trends of the two spatial subregions. That is, when data in one subregion increases, data in the other is likely to increase as well. This may indicate the existence of some mechanism of mutual influence between the two subregions, such as groundwater connectivity or geological structure continuity. A correlation coefficient close to 0 indicates no significant correlation between the data in the two subregions. A negative correlation coefficient with a large absolute value indicates a negative correlation between the data trends in the two subregions. That is, when data in one subregion increases, data in the other is likely to decrease. This may be due to differences in geological structure or other special factors. By calculating the correlation coefficient, we can fully understand the spatial relationships between geohazard monitoring data.
[0041] Step S115: Construct a spatial correlation feature based on the above correlation coefficient, perform feature splicing processing on the above spatial correlation feature and the above sequence correlation feature, and generate a state vector containing the sequence correlation feature and the spatial correlation feature. The dimension of the state vector matches the state input dimension of the reinforcement learning agent.
[0042] Based on the previously calculated correlation coefficients between the monitoring data subsets of different spatial subregions, a spatial correlation feature is constructed. A spatial correlation feature is a comprehensive representation of the spatial relationships between geological hazard monitoring data. The correlation coefficients between all different spatial subregions can be arranged in a set order to form a spatial correlation feature.
[0043] Then, the spatial correlation features are combined with the previously generated sequence correlation features for feature splicing. The goal of this splicing is to integrate the temporal and spatial features of the geological hazard monitoring data to form a more comprehensive and richer feature representation. During the splicing process, it is necessary to ensure the consistency of the order and dimensions of the two features to ensure that the spliced features accurately reflect the overall characteristics of the geological hazard monitoring data.
[0044] Finally, a state vector containing both sequential and spatial correlation features is generated. When generating the state vector, it is important to ensure that its dimensions match those of the reinforcement learning agent's state input. Reinforcement learning agents are designed with a fixed state input dimension requirement. Only when the state vector's dimensions match these requirements can the agent correctly receive and process the state vector and perform subsequent decision-making and analysis. If the state vector's dimensions do not match, the agent may not function properly or produce inaccurate decision results. By generating a state vector that matches the reinforcement learning agent's state input dimensions, the agent can fully leverage its capabilities and improve the accuracy of geological hazard monitoring and early warning.
[0045] Step S120: Initialize the reinforcement learning policy network, which includes a policy evaluation module and a policy improvement module. The policy evaluation module is used to calculate the estimated value function of the current policy, and the policy improvement module is used to adjust the policy parameters based on the estimated value function.
[0046] After completing the state vector construction of the geological hazard monitoring data, it is necessary to initialize the reinforcement learning strategy network. The reinforcement learning strategy network is the core part of the entire geological hazard monitoring data processing system, which consists of a strategy evaluation module and a strategy improvement module.
[0047] The strategy evaluation module's primary function is to calculate the estimated value function for the current strategy. This value function estimate reflects the expected return from taking different actions under the current strategy. In the context of geological hazard monitoring, actions can include different processing methods for monitoring data, such as different filter parameter settings and feature enhancement strength. Based on the current state vector and strategy, the strategy evaluation module evaluates each possible action and calculates the potential return of that action over a period of time. This return can range from improved geological hazard prediction accuracy to improved processing of abnormal data.
[0048] The policy improvement module adjusts policy parameters based on the estimated value function. If the estimated value function for an action is high, it indicates that the action is likely to achieve a good return under the current policy. The policy improvement module will adjust the policy parameters accordingly to increase the probability of the action being selected. Conversely, if the estimated value function for an action is low, the policy improvement module will decrease the probability of the action being selected. By continuously adjusting policy parameters, the policy improvement module can gradually optimize the policy and achieve higher returns.
[0049] When initializing a reinforcement learning policy network, it's important to properly configure the parameters of the policy evaluation module and policy improvement module. These parameters include the network's initial weights and learning rate. The initial weights affect the network's initial performance. Random initialization is generally recommended, but it's important to ensure that the initial weights are within a reasonable range to avoid overfitting or underfitting. The learning rate controls the speed at which policy parameters are adjusted and should be adjusted appropriately based on the specific application scenario and training conditions.
[0050] Step S121: Determine the network architecture of the reinforcement learning strategy network, which includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the same as the dimension of the state vector, and the number of neurons in the output layer is the same as the action dimension of the processing action set.
[0051] When initializing a reinforcement learning policy network, the first thing to do is to determine its network architecture. Reinforcement learning policy networks typically use a multi-layer neural network architecture, including an input layer, a hidden layer, and an output layer.
[0052] The input layer receives external input data. In this system, the number of neurons in the input layer matches the dimension of the previously generated state vector. This is because the state vector is recognizable input data for the reinforcement learning agent. Matching the number of neurons in the input layer with the dimension of the state vector ensures that the state vector is correctly input into the network. For example, if the state vector contains both sequential and spatial correlation features, and the concatenation of these features results in a state vector with a dimension of 100, then the input layer requires 100 neurons, each corresponding to a feature value in the state vector.
[0053] Hidden layers are a key component of the network for feature extraction and transformation. There can be multiple hidden layers, each containing a set number of neurons. These neurons process input data using nonlinear activation functions, mapping the input data into a higher-dimensional feature space to better capture the inherent patterns of the data. The number of neurons and layers in the hidden layer needs to be adjusted based on the specific problem and data characteristics. Generally speaking, increasing the number of hidden layers and neurons improves the network's expressive power, but also increases training complexity and time.
[0054] The number of neurons in the output layer is the same as the action dimension of the processing action set. The processing action set includes various processing actions for geological hazard monitoring data, such as data filtering operations, feature enhancement operations, and anomaly identification operations. Each action may have different parameter settings, which constitute the action dimension. The number of neurons in the output layer matches the action dimension, allowing the network to output decision information for different actions. For example, if there are five different actions in the processing action set, and each action has three parameters that need to be determined, then the action dimension of the processing action set is 15, and the output layer needs to have 15 neurons, one for each action parameter output value.
[0055] Step S122: setting the initial parameters of the value function of the strategy evaluation module. The initial parameters of the value function include a state feature weight matrix and a bias vector. The state feature weight matrix is used to perform weighted processing on different feature dimensions of the state vector.
[0056] After determining the network architecture of the reinforcement learning policy network, it is necessary to set the initial parameters of the value function of the policy evaluation module. The initial parameters of the value function include the state feature weight matrix and the bias vector.
[0057] The state feature weight matrix is used to weight the different feature dimensions of the state vector. The state vector contains information from multiple dimensions, such as the serial correlation characteristics and spatial correlation characteristics of geological hazard monitoring data. Features from different dimensions may contribute differently to the value function. For example, in geological hazard prediction, the trend continuity characteristics of mountain displacement may have a greater impact on the prediction results than the soil moisture correlation characteristics of certain spatial subregions. The state feature weight matrix can be used to assign different weights to each feature dimension, highlighting the role of important features.
[0058] The state-feature weight matrix is a two-dimensional matrix with the number of rows equal to the dimensions of the state vector, and the number of columns can be determined based on specific design requirements. Each element in the matrix represents the weight value for the corresponding feature dimension. During initialization, random initialization can be used to set the matrix elements to random values. However, these random values must be within a reasonable range to avoid excessively large or small weights that may adversely affect network performance.
[0059] The bias vector is used to adjust the output of the value function. The length of the bias vector is the same as the number of columns in the state feature weight matrix. Each element of the bias vector corresponds to an output dimension, providing a basic offset for the output value. During initialization, the elements of the bias vector can also be randomly initialized.
[0060] By reasonably setting the initial values of the state feature weight matrix and the bias vector, the policy evaluation module can effectively evaluate the state vector in the early stages of training, laying the foundation for subsequent value function calculation and policy adjustment.
[0061] Step S123: configuring a policy parameter update rule of the policy improvement module, wherein the policy parameter update rule determines the parameter adjustment direction and adjustment range based on the difference between the estimated value of the value function and the target value function.
[0062] The core task of the policy improvement module is to adjust the policy parameters based on the difference between the estimated value function and the target value function. The target value function is an ideal value function that represents the return that can be obtained under the optimal policy.
[0063] First, we need to define the target value function. This can be done in a variety of ways. In the case of geological disaster monitoring, the target value function can be determined based on historical data and actual needs. For example, the target value function could be to maximize the accuracy of geological disaster predictions or to improve the efficiency of abnormal data processing.
[0064] Next, the difference between the estimated value function and the target value function is calculated and used as the parameter adjustment error. If the parameter adjustment error is positive, it means that the estimated value function of the current policy is lower than the target value function, indicating that the performance of the current policy can be improved and the policy parameters need to be adjusted to increase the estimated value function. For example, if the current accuracy of geological disaster prediction is lower than the target accuracy, the policy improvement module will adjust the policy parameters to increase the probability of selecting actions that can improve prediction accuracy.
[0065] If the parameter adjustment error is negative, it means that the current policy's estimated value function is higher than the target value function, indicating overestimation. The policy parameters need to be adjusted to reduce the estimated value function. For example, if the current efficiency of abnormal data processing is too high, but there may be misjudgments, the policy improvement module will adjust the policy parameters to reduce the probability of selecting actions that may lead to misjudgments.
[0066] To control the magnitude of parameter adjustments, a learning rate coefficient is required. The learning rate coefficient determines the magnitude of each parameter update. If the learning rate coefficient is too large, the parameter update amplitude may be too large, causing the policy parameters to fluctuate around the optimal value and making convergence difficult. If the learning rate coefficient is too small, the parameter update speed will be very slow, and the training process will become very time-consuming. The value of the learning rate coefficient needs to be dynamically adjusted according to the training stage of the policy network. In the early stages of training, a large learning rate coefficient can be set to quickly explore the policy space. As training progresses, the learning rate coefficient can be gradually reduced to ensure that the policy parameters converge to the optimal value more stably.
[0067] Finally, the policy parameter adjustment step size is calculated based on the parameter adjustment error and the learning rate coefficient. The adjustment step size is equal to the parameter adjustment error multiplied by the learning rate coefficient. The parameter adjustment direction (increase or decrease) and the adjustment step size are combined to form the policy parameter update rule. The policy improvement module uses this policy parameter update rule to modify the policy parameters during each iteration, gradually optimizing the policy and enabling the reinforcement learning agent to take more optimal actions and achieve higher rewards.
[0068] Step S124: Construct an experience replay cache pool, which is used to store experience samples consisting of state vectors, processing action sets, reward signals and next state vectors generated during the interaction between the reinforcement learning agent and the geological disaster monitoring environment.
[0069] During the reinforcement learning process, the experience replay cache pool is used to store experience samples generated during the interaction between the reinforcement learning agent and the geological disaster monitoring environment.
[0070] An experience sample consists of a state vector, a set of processing actions, a reward signal, and a next-state vector. The state vector represents the geological hazard monitoring data received by the reinforcement learning agent at a given moment and incorporates both sequential and spatial correlation features. The set of processing actions represents the processing actions taken by the agent based on the current state vector, such as data filtering, feature enhancement, and anomaly identification. The reward signal is the geological hazard monitoring environment's feedback on the set of processing actions, reflecting the effectiveness of the action set. If the action set improves the accuracy of geological hazard predictions or effectively handles anomaly data, the reward signal will be positive. Conversely, if the action set increases prediction error or degrades anomaly data handling, the reward signal will be negative. The next-state vector represents the new state of the geological hazard monitoring environment after executing the action set, reflecting the impact of the action set on the geological hazard monitoring data.
[0071] The purpose of the experience replay cache is to provide a diverse set of experience samples during training. In reinforcement learning, the agent's learning process is based on its interaction with the environment. If the latest experience samples are used for each training session, the correlation between the data may be too high, causing the agent to fall into a local optimum. The experience replay cache allows random sampling of experience samples from the cache for training, thereby breaking the correlation between the data and improving the stability and efficiency of training.
[0072] The experience replay cache pool has a certain capacity limit. When the number of experience samples in the cache pool reaches the upper limit, a set strategy is used to replace the old experience samples. A common strategy is the first-in-first-out strategy, which means that when new experience samples enter the cache pool, the oldest experience samples are removed.
[0073] Step S125: Initialize the network weight parameters of the policy evaluation module and the policy improvement module. The network weight parameters are initialized using a random normal distribution. The initialized network weight parameters meet the preset value range constraints.
[0074] After completing the previous steps, it is necessary to initialize the network weight parameters of the policy evaluation module and the policy improvement module. The network weight parameters determine the connection strength between each neuron in the neural network and have a significant impact on the performance of the network.
[0075] A common approach is to use a random normal distribution for initialization. A random normal distribution is a probability distribution characterized by a symmetric distribution with most data concentrated near its mean. When initializing network weight parameters, values are randomly drawn from the random normal distribution as the initial values for the weight parameters.
[0076] At the same time, the network weight parameters after initialization must meet the preset numerical range constraints. The preset numerical range constraints ensure that the network weight parameters are not too large or too small. If the network weight parameters are too large, the output values of the neurons may be too large, making the network training process unstable; if the network weight parameters are too small, the output values of the neurons may be too small, which will limit the network's learning ability.
[0077] For example, a reasonable standard deviation and mean can be set to control the range of the random normal distribution. By adjusting the standard deviation and mean, the network weight parameters after initialization can be kept within a suitable range. During initialization, the connection weights of each neuron are independently and randomly initialized to ensure a certain degree of randomness and diversity in the network's initial state. This prevents all neurons from having the same initial weights, which would prevent the network from learning different features during training.
[0078] Step S130: Execute a strategy selection operation based on the state vector and the reinforcement learning strategy network to generate a set of processing actions for geological disaster monitoring data, which includes data filtering operations, feature enhancement operations, and anomaly recognition operations.
[0079] After completing the initialization of the reinforcement learning policy network, the policy selection operation is performed based on the previously generated state vector and the reinforcement learning policy network. The state vector contains the serial correlation characteristics and spatial correlation characteristics of the geological disaster monitoring data.
[0080] The reinforcement learning policy network evaluates different processing actions based on the input state vector through internal calculations and processing. This set of processing actions includes data filtering, feature enhancement, and anomaly detection. The purpose of data filtering is to remove noise and interference from monitoring data and improve data quality. For example, mountain displacement monitoring data may contain noise due to sensor errors or external interference. Appropriate filtering can remove this noise.
[0081] Feature enhancement emphasizes important features in the data, making subsequent analysis more accurate. Certain features in geological hazard monitoring data may be crucial for hazard prediction and analysis, such as the continuity of mountain displacement trends and the rate of change of groundwater levels. Feature enhancement can amplify the impact of these important features and improve the accuracy of hazard prediction.
[0082] Anomaly recognition is used to detect anomalies in data and promptly identify potential geological disaster risks. For example, when abnormal conditions such as sudden increases in mountain displacement or sharp rises in groundwater levels occur, anomaly recognition can quickly identify these anomalies and issue early warning signals.
[0083] The reinforcement learning policy network selects the optimal combination of actions based on the estimated value functions of different actions under the current state vector. During this process, the policy network considers the expected reward and risk of each action and selects those actions that maximize the reward under the current state. For example, if the current state vector indicates a high probability of a geological disaster, the policy network might choose to enhance data filtering and anomaly detection to improve geological disaster monitoring and early warning capabilities.
[0084] Step S131: Input the state vector into the input layer of the reinforcement learning strategy network, perform nonlinear feature conversion on the state vector through the hidden layer, and generate a high-dimensional strategy feature vector. The dimension of the high-dimensional strategy feature vector is higher than the dimension of the state vector.
[0085] When executing a policy selection operation, the state vector is first input into the input layer of the reinforcement learning policy network. The neurons in the input layer receive each feature value in the state vector and pass it to the hidden layer.
[0086] The hidden layer is the key component of the network for feature extraction and transformation. Neurons in the hidden layer perform nonlinear feature transformation on the input state vector. This nonlinear feature transformation is achieved using nonlinear activation functions. Common nonlinear activation functions include ReLU (rectified linear unit) and Sigmoid function.
[0087] Taking the ReLU activation function as an example, the expression for the ReLU function is f(x) = max(0, x). When the input value x is greater than 0, the output value is equal to x; when the input value x is less than or equal to 0, the output value is 0. Through the ReLU activation function, neurons in the hidden layer can perform nonlinear transformations on the input state vector, enabling the network to learn more complex features.
[0088] In the hidden layer, each neuron receives multiple eigenvalues from the input layer and performs a weighted summation of these eigenvalues based on its connection weights. The result of the weighted summation is then input into a nonlinear activation function for processing to obtain the output value of the neuron.
[0089] Through multiple hidden layers, the state vector is gradually transformed into a high-dimensional policy feature vector. The dimension of the high-dimensional policy feature vector is higher than that of the state vector. This is because the hidden layers transform the low-dimensional input features into a higher-dimensional feature space through nonlinear feature transformation. In a high-dimensional feature space, the data features are richer and more complex, better representing the inherent patterns of geological hazard monitoring data.
[0090] For example, the dimension of the state vector may be 100, and after processing by the hidden layer, the dimension of the generated high-dimensional policy feature vector may reach 500 or even higher.
[0091] Step S132: Call the strategy evaluation module to perform value function estimation processing on the above high-dimensional strategy feature vector, and calculate the expected cumulative reward value corresponding to different processing actions in the current state. The expected cumulative reward value is used to measure the long-term utility of the processing action.
[0092] After generating the high-dimensional policy feature vector, the policy evaluation module is called to perform value function estimation. The policy evaluation module calculates the expected cumulative reward value corresponding to different processing actions in the current state based on the pre-set value function and network parameters.
[0093] The expected cumulative reward value is used to measure the long-term utility of a treatment action. In the geological disaster monitoring scenario, the impact of a treatment action may not only be limited to the current moment, but also affect geological disaster monitoring and treatment for a period of time in the future. Therefore, the long-term utility of the treatment action needs to be considered.
[0094] The strategy evaluation module evaluates each possible action. For each action, it predicts the potential rewards it will bring in the future based on the high-dimensional strategy feature vector and the current strategy. These rewards can include improvements in the accuracy of geological disaster predictions, improved processing of abnormal data, and so on.
[0095] For example, when filtering data, different filter parameters may have different impacts on the quality of the current data and the accuracy of subsequent geological disaster predictions. The strategy evaluation module takes these factors into consideration and calculates the expected cumulative reward value corresponding to each filter parameter selection.
[0096] The calculation of expected cumulative reward is based on certain assumptions and models. The strategy evaluation module estimates future rewards based on historical data and experience. This calculation takes into account the interplay between different actions and the dynamic nature of the geological hazard monitoring environment. By calculating the expected cumulative reward, the reinforcement learning agent can select actions that will yield the greatest long-term rewards.
[0097] Step S133: Execute ε-greedy strategy selection based on the expected cumulative reward value and a preset exploration rate parameter, randomly selecting a processing action within a preset probability range or selecting the processing action with the maximum expected cumulative reward value. The exploration rate parameter is used to balance exploration and exploitation in strategy selection.
[0098] After obtaining the expected cumulative reward values corresponding to different processing actions, the ε-greedy strategy is used to select the processing action. The ε-greedy strategy is a strategy that balances exploration and exploitation.
[0099] Exploration involves randomly selecting actions to discover new, potentially better strategies. In geological hazard monitoring scenarios, there may be undiscovered actions or parameter settings. Exploration allows us to try out these new possibilities, potentially leading to more effective solutions.
[0100] Exploitation means choosing the action with the largest expected cumulative reward to obtain the best known reward. If the expected cumulative reward of an action is significantly higher than that of other actions, then choosing that action will yield the best results in the current situation.
[0101] The preset exploration rate parameter determines whether to explore or exploit within a preset probability range. The exploration rate parameter is usually a value between 0 and 1. For example, setting the exploration rate parameter to 0.1 means that exploration is performed with a 10% probability, and the action is randomly selected; and exploitation is performed with a 90% probability, and the action with the maximum expected cumulative reward is selected.
[0102] To implement the ε-greedy strategy selection, a random number uniformly distributed between 0 and 1 is first generated. This random number is compared with the preset exploration rate parameter.
[0103] If the generated random number is less than the exploration rate parameter, an exploration operation is performed. In the exploration operation, a processing action is randomly selected from the set of processing actions. The random selection is based on a uniform probability distribution, that is, each processing action has an equal probability of being selected.
[0104] If the generated random number is greater than or equal to the exploration rate parameter, the exploit action is executed. In the exploit action, the action with the largest expected cumulative reward value is selected from the set of actions. If there are multiple actions with the same maximum expected cumulative reward value, one of them is randomly selected.
[0105] As training progresses, the exploration rate parameter is dynamically adjusted to better balance exploration and exploitation. Typically, in the early stages of training, the exploration rate parameter is set high to allow the agent to broadly explore the policy space and discover more possible strategies. As training progresses, the exploration rate parameter is gradually reduced, allowing the agent to more readily exploit the discovered optimal strategy, improving its performance.
[0106] Step S1331: Generate a random number uniformly distributed within a preset interval, and compare the random number with a preset exploration rate parameter.
[0107] When executing the ε-greedy strategy, we first need to generate a random number that is uniformly distributed within a preset interval. The preset interval is usually [0, 1]. Uniform distribution means that the probability of each value being generated within the interval is equal.
[0108] There are many ways to generate random numbers. One approach is to use a computer's random number generator. These are typically based on pseudorandom number algorithms, which generate seemingly random values within a set range.
[0109] After generating a random number, it is compared to the preset exploration rate parameter. This comparison is a key step in determining whether to explore or exploit. If the random number is less than the exploration rate parameter, the current phase is exploration, and a random action should be selected. If the random number is greater than or equal to the exploration rate parameter, the current phase is exploitation, and the action with the highest expected cumulative reward should be selected.
[0110] For example, if the default exploration rate parameter is 0.2, if the generated random number is 0.1, since 0.1 is less than 0.2, the exploration operation is performed, and a random action is selected from the set of actions. If the generated random number is 0.3, since 0.3 is greater than 0.2, the exploitation operation is performed, and the action with the largest expected cumulative reward value is selected.
[0111] Step S1332: When the random number is less than the exploration rate parameter, an exploration operation is performed to randomly select a processing action from the processing action set. The random selection is implemented based on a uniform probability distribution.
[0112] When the generated random number is less than the exploration rate parameter, the exploration phase begins. In the exploration phase, a processing action is randomly selected from the processing action set.
[0113] Random selection is based on a uniform probability distribution. This means that each action in the set has an equal probability of being selected. For example, if there are five different actions in the set: Action A, Action B, Action C, Action D, and Action E, each action has a 1 / 5 probability of being selected during the exploration phase.
[0114] Random selection based on a uniform probability distribution ensures that the agent can try out a wide range of possibilities within the set of actions. In geological hazard monitoring scenarios, there may be some underexplored actions or parameter settings. By using random selection based on a uniform probability distribution, the agent has the opportunity to try these new possibilities and potentially discover a more optimal strategy.
[0115] For example, for data filtering operations, there may be different filtering algorithms and filtering parameters to choose from. During the exploration phase, the agent may randomly select a filtering algorithm or parameter setting that has not been tried before to observe its impact on the processing effect of geological hazard monitoring data.
[0116] Step S1333: When the random number is greater than or equal to the exploration rate parameter, the utilization operation is executed to select the processing action with the maximum expected cumulative reward value from the processing action set. If there are multiple processing actions with the same maximum expected cumulative reward value, one of the processing actions is randomly selected.
[0117] When the generated random number is greater than or equal to the exploration rate parameter, the exploitation phase begins. In the exploitation phase, the action with the largest expected cumulative reward value is selected from the action set.
[0118] After calling the policy evaluation module to calculate the expected cumulative reward values for different processing actions, these reward values are compared to find the processing action with the largest expected cumulative reward value.
[0119] If only one action has the highest expected cumulative reward, that action is chosen directly. For example, there are three actions in the action set: Action X, Action Y, and Action Z, with expected cumulative rewards of 0.7, 0.9, and 0.6, respectively. Since Action Y has the highest expected cumulative reward, it is chosen in the exploitation phase.
[0120] If there are multiple actions with the same maximum expected cumulative reward, one of them is randomly selected. For example, there are four actions in the action set: action M, action N, action P, and action Q, with expected cumulative rewards of 0.8, 0.8, 0.6, and 0.5, respectively. If actions M and N have the same maximum expected cumulative reward, one of them is randomly selected.
[0121] By selecting the action with the largest expected cumulative reward, the agent can leverage the knowledge it has learned to obtain the optimal reward in the current state. This helps improve the effectiveness and efficiency of geological hazard monitoring data processing.
[0122] Step S1334: After completing a preset number of strategy selection operations, the value of the exploration rate parameter is reduced according to a preset decay rate. The decay rate is a positive number less than 1, so that the exploration rate parameter gradually decreases as the number of training iterations increases.
[0123] In order to better balance exploration and exploitation during training, the exploration rate parameter needs to be dynamically adjusted. After completing a preset number of strategy selection operations, the exploration rate parameter value is reduced according to the preset decay rate.
[0124] The preset number of times is determined based on specific training needs and experience. For example, you can set the preset number to 100, which means that the exploration rate parameter is adjusted once every 100 strategy selection operations.
[0125] The preset decay rate is a positive number less than 1. For example, the preset decay rate is 0.9. Each time you adjust the exploration rate, multiply the current exploration rate parameter by the preset decay rate to get the new exploration rate parameter.
[0126] In the early stages of training, the exploration rate parameter is often set high, such as 0.5. This is because the agent has little knowledge of the policy space and needs to conduct extensive exploration to discover new, potentially optimal policies. As training iterations increase, the agent gradually learns some effective policies. At this point, the exploration rate parameter can be appropriately lowered to encourage the agent to utilize the optimal policy it has already discovered.
[0127] For example, the initial exploration rate parameter is 0.5, the preset decay rate is 0.9, and the preset number of times is 100. After completing 100 policy selection operations, the new exploration rate parameter is 0.5*0.9=0.45. As training continues, the exploration rate parameter will continue to decrease. For example, after completing 200 policy selection operations, the exploration rate parameter becomes 0.45*0.9=0.405.
[0128] Step S1335: After the exploration rate parameter decreases to a preset minimum threshold, the decay is stopped and the exploration rate parameter is kept at the preset minimum threshold.
[0129] When reducing the exploration rate parameter according to the preset decay rate, a preset minimum threshold needs to be set. The preset minimum threshold is to ensure that the agent still has a certain probability of exploration in the later stages of training, avoiding being completely trapped in the local optimal solution.
[0130] When the exploration rate parameter drops to the preset minimum threshold, the exploration rate parameter decay operation stops and the exploration rate parameter is maintained at the preset minimum threshold. For example, the preset minimum threshold is 0.05. Once the exploration rate parameter drops to 0.05, no matter how many subsequent strategy selection operations are completed, the exploration rate parameter will remain at 0.05.
[0131] This setting aims to ensure that, in the later stages of training, while the agent primarily utilizes the optimal strategy already discovered, it still has a chance to try new actions to discover potentially better strategies. In the geological disaster monitoring scenario, the geological environment may change, and new geological disaster types or data features may emerge. By maintaining a certain exploration probability, the agent can adapt to these changes in a timely manner, improving the effectiveness of geological disaster monitoring and response.
[0132] Step S134: determining a corresponding action parameter configuration according to the selected processing action, where the action parameter configuration includes a filter window size for a data filtering operation, an enhancement strength coefficient for a feature enhancement operation, and a recognition threshold for an anomaly recognition operation.
[0133] After completing the policy selection operation, determine the corresponding action parameter configuration based on the selected processing action. Different processing actions may require different parameters to be set.
[0134] For data filtering operations, the action parameter configuration includes the filter window size. The filter window size determines the range and accuracy of the filter. If the filter window size is set to a larger value, the filter smoothing effect will be better, but some detailed information may be lost; if the filter window size is set to a smaller value, more detailed information can be retained, but the smoothing effect may be poor. For example, when filtering mountain displacement monitoring data, if the filter window size is set to 5, it means that the current data point and the two data points before and after it will be averaged to obtain the filtered value.
[0135] For feature enhancement operations, the action parameter configuration includes an enhancement strength coefficient. The enhancement strength coefficient determines the degree of feature enhancement. If the enhancement strength coefficient is set to a large value, the feature will be significantly enhanced, but may cause data distortion; if the enhancement strength coefficient is set to a small value, the feature enhancement effect will be less obvious. For example, when enhancing a key feature in geological disaster monitoring data, if the enhancement strength coefficient is 2, the feature value will be multiplied by 2, thereby increasing its role in subsequent analysis.
[0136] For anomaly identification operations, the action parameter configuration includes a recognition threshold. The recognition threshold determines the criteria for identifying anomalies. A low recognition threshold may identify more anomalies, but may also result in false positives. A high recognition threshold may reduce false positives, but may miss some true anomalies. For example, when identifying anomalies in mountain displacement data, if the recognition threshold is set to 10, any change in the mountain displacement exceeding 10 is considered an anomaly.
[0137] Based on the selected processing action, appropriate parameter values are selected from a pre-set parameter set as the action parameter configuration. These parameter values are selected based on the learning results of the policy network and the analysis of geological disaster monitoring data.
[0138] Step S135: Combine the above processing actions with the corresponding action parameter configurations to generate a processing action set including data filtering operations, feature enhancement operations, and anomaly recognition operations. Each operation in the processing action set is executed in a preset order.
[0139] After determining the processing actions and the corresponding action parameter configurations, they are combined to generate a processing action set including data filtering operations, feature enhancement operations, and anomaly recognition operations.
[0140] Each operation is executed in a pre-set order. This order is determined by the logic and efficiency of geological hazard monitoring data processing. Typically, data filtering is performed first because it removes noise and interference from the data, improving data quality. For example, when processing mountain displacement monitoring data, the data is first filtered using a pre-set filter window size to produce smoothed data.
[0141] Then, feature enhancement is performed. For example, some important features in the filtered data are enhanced according to the set enhancement intensity coefficient.
[0142] Finally, anomaly detection is performed. This process uses the set detection threshold to identify possible anomalies in the enhanced features. For example, based on the change in mountain displacement and the set detection threshold, it can determine whether there is a mountain displacement anomaly.
[0143] By combining processing actions with corresponding action parameter configurations and executing them in a preset order, geological disaster monitoring data can be processed systematically and effectively, thereby improving the accuracy of geological disaster monitoring and early warning.
[0144] Step S140: Obtain the feedback signal of the geological disaster monitoring environment to the above-mentioned processing action set, and generate a reward signal based on the above-mentioned feedback signal and the above-mentioned value function estimation value. The reward signal is used to evaluate the impact of the processing action set on the processing effect of the geological disaster monitoring data.
[0145] After generating and executing a set of processing actions, it is necessary to obtain feedback signals from the geological hazard monitoring environment on the processing action set. The feedback signals reflect the actual processing effect of the processing action set on the geological hazard monitoring data.
[0146] Applying a set of processing actions to geological hazard monitoring data yields processed monitoring data results. These results consist of a filtered data sequence, an enhanced feature set, and anomaly identification result markers. The filtered data sequence is the result of filtering, removing noise and interference to a certain extent. The enhanced feature set is the result of feature enhancement, highlighting key features. The anomaly identification result markers clearly identify anomalies in the data.
[0147] At the same time, actual monitoring data of the geological disaster monitoring environment is collected. This data includes information on geological structural changes and environmental influencing factors. Geological structural change information reflects actual changes in the geological environment, such as mountain deformation and rock fractures; environmental influencing factor data reflects the impact of the external environment on geological disasters, such as rainfall and temperature.
[0148] The processed monitoring data is compared and analyzed with the actual monitoring data, and a matching parameter is calculated to determine the accuracy of the processing action set. For example, if the anomaly area identified in the anomaly identification result tag closely matches the anomaly area displayed in the actual geological structure change information, this indicates a high degree of accuracy in the anomaly identification operation, and the matching parameter will be large.
[0149] The policy evaluation module is called to obtain the estimated value function for the current set of processing actions. The difference between the matching parameter and the estimated value function is calculated as the error correction term. If the error correction term is positive, the actual processing effect is better than the estimated value function value; if the error correction term is negative, the actual processing effect is worse than the estimated value function value.
[0150] A reward function is constructed based on the matching parameter and the error correction term. The design of the reward function requires comprehensive consideration of both the matching parameter and the error correction term. Generally, the value of the reward signal is positively correlated with the matching parameter and negatively correlated with the absolute value of the error correction term. A larger matching parameter indicates a high accuracy of the processing action set, and the reward signal will increase accordingly. A larger absolute value of the error correction term indicates a significant deviation between the value function estimate and the actual value, and the reward signal will decrease accordingly. The reward signal can be used to evaluate the impact of the processing action set on the effectiveness of geological hazard monitoring data processing.
[0151] Step S141: Apply the above processing action set to the geological disaster monitoring data to obtain the processed monitoring data results, which include the filtered data sequence, the enhanced feature set and the abnormal recognition result mark.
[0152] The generated set of processing actions is applied to the geological hazard monitoring data. The processing action set includes data filtering operations, feature enhancement operations, and anomaly identification operations, which are performed in a preset order.
[0153] First, data filtering is performed. The geological disaster monitoring data is filtered based on the filter window size configured in the data filtering action parameters. For mountain displacement monitoring data, assuming a filter window size of 3, for each data point, the current data point and one data point before and after it are selected. The average of these three data points is calculated as the filtered value of the data point. This process is repeated for all data points in sequence to obtain a filtered data sequence. This filtered data sequence removes some noise and interference, resulting in smoother data and facilitating subsequent analysis.
[0154] Next, the feature enhancement operation is performed. Based on the enhancement intensity coefficient configured in the feature enhancement operation's action parameters, certain important features in the filtered data sequence are enhanced. For example, in geological disaster monitoring data, the rate of change of groundwater levels may be a key feature. If the enhancement intensity coefficient is 1.5, the rate of change of groundwater levels in the filtered data sequence is multiplied by 1.5 to obtain the enhanced feature set. This enhanced feature set highlights important features, allowing them to play a more important role in subsequent analysis.
[0155] Finally, anomaly identification is performed. Based on the identification threshold configured in the anomaly identification operation's action parameters, the enhanced feature set is evaluated to identify possible anomalies. For example, for mountain displacement monitoring data, if the identification threshold is 8, any change in mountain displacement in the enhanced feature set exceeding 8 will be marked as an anomaly. All identified anomalies are marked to obtain an anomaly identification result label.
[0156] By sequentially performing data filtering operations, feature enhancement operations, and anomaly recognition operations, a processed monitoring data result is obtained, which includes a filtered data sequence, an enhanced feature set, and an anomaly recognition result label.
[0157] Step S142: Collecting actual monitoring data of the geological disaster monitoring environment, where the actual monitoring data includes geological structure change information and environmental impact factor data.
[0158] The actual monitoring data of the geological disaster monitoring environment is collected for comparison and analysis with the processed monitoring data results to evaluate the accuracy of the processing action set.
[0159] Information on geological structural changes can be collected in a variety of ways. For example, geological radar can be used to monitor changes in underground geological structures. Geological radar transmits electromagnetic waves and receives reflected waves, analyzing the characteristics of the reflected waves to infer the underground geological structure. Inclinometers can also be used to monitor changes in the tilt angle of a mountain. Inclinometers record the mountain's inclination in real time, and when the mountain deforms, the tilt angle changes accordingly. These devices and methods can provide detailed information on geological structural changes, such as the direction of mountain displacement and the location of rock fractures.
[0160] Collecting data on environmental impact factors also requires a variety of monitoring equipment. For example, rain gauges can accurately record rainfall over a specific period of time, while thermometers can measure ambient temperature in real time. This data is crucial for understanding the mechanisms of geological hazards. For example, heavy rainfall can increase the weight of a mountain, increasing the risk of landslides; high temperatures can cause rocks to expand, affecting the stability of geological structures.
[0161] The collected geological structure change information and environmental impact factor data will be sorted and recorded to form actual monitoring data of the geological disaster monitoring environment.
[0162] Step S143: Compare and analyze the processed monitoring data results with the actual monitoring data, and calculate the matching parameter between the data processing results and the actual situation. The matching parameter is used to characterize the accuracy of the processing action set.
[0163] For example, first, the anomaly identification result markers are extracted from the processed monitoring data results to determine the location of the abnormal area and the degree of abnormality indicated by the anomaly identification result markers. The anomaly identification result markers clearly define the abnormalities identified in the data. By analyzing the markers, the specific location of the abnormal area and the severity of the anomaly can be determined. For example, in mountain displacement monitoring data, the anomaly identification result markers may mark certain areas where the mountain displacement exceeds the normal range. Based on the markers, the geographical location of these abnormal areas can be determined, and the degree of abnormality can be divided into mild, moderate, severe, etc. according to the size of the displacement.
[0164] The actual location and severity of abnormal areas, as indicated by geological structural change information, are then extracted from the actual monitoring data. This information reflects the actual changes in the geological environment, including the location of the actual abnormal areas and the severity of the abnormalities. For example, geological radar monitoring results may indicate rock fractures in certain areas underground, representing the actual location of the abnormal areas; the severity of the abnormality can be determined based on the extent of the rock fractures.
[0165] Calculate the spatial overlap ratio between the abnormal region indicated by the anomaly identification result marker and the actual abnormal region. The spatial overlap ratio is the ratio of the area of the overlapping region to the area of the actual abnormal region. A high spatial overlap ratio indicates that the anomaly identification operation accurately identifies the actual abnormal region; a low spatial overlap ratio indicates that the anomaly identification operation has a certain degree of error.
[0166] Calculate the level deviation between the abnormality level indicated by the abnormality identification result mark and the actual abnormality level. The level deviation is the absolute value of the difference between the two levels. If the level deviation is small, it means that the abnormality identification operation accurately judged the abnormality level; if the level deviation is large, it means that the abnormality identification operation had an error in judging the abnormality level.
[0167] A matching calculation function is constructed based on the spatial overlap ratio and the level deviation value. This function performs a weighted summation of the spatial overlap ratio and (1 minus the normalized level deviation value) to obtain a matching parameter that compares the data processing result with the actual situation. The weights in this weighted summation are pre-set based on the priority of the anomaly identification task. For example, if the accuracy of the anomaly location is highly required, a larger weight can be assigned to the spatial overlap ratio; if the accuracy of the anomaly severity is highly required, a larger weight can be assigned to (1 minus the normalized level deviation value). By constructing this matching calculation function, the matching parameter obtained between the data processing result and the actual situation can be used to intuitively reflect the accuracy of the processing action set. The closer the matching parameter is to 1, the more consistent the processing result of the geological hazard monitoring data is with the actual situation, and the higher the accuracy of the processing action set. The closer the matching parameter is to 0, the greater the discrepancy between the processing result and the actual situation, and the lower the accuracy of the processing action set.
[0168] Step S144: calling the policy evaluation module to obtain the estimated value of the value function corresponding to the current processing action set, and calculating the difference between the matching parameter and the estimated value of the value function as an error correction item.
[0169] After obtaining the matching parameters, the policy evaluation module is called to obtain the estimated value function corresponding to the current set of processing actions. The policy evaluation module has previously evaluated the value of different processing actions based on the input state vector and policy. Here, the estimated value corresponding to the currently selected processing action set is directly obtained.
[0170] The matching parameter is compared with the estimated value of the value function, and the difference between them is calculated. This difference is the error correction term. The error correction term reflects the deviation between the value estimated by the policy evaluation module and the actual treatment effect. If the error correction term is positive, it means that the actual treatment effect is worse than the policy evaluation module's previous estimate, indicating that the policy may have underestimated the value of the treatment action. If the error correction term is negative, it means that the actual treatment effect is worse than the estimated value, and the policy may have overestimated the value of the treatment action.
[0171] For example, if the matching parameter is 0.8 and the estimated value of the value function is 0.6, then the error correction term is 0.8-0.6=0.2, indicating that the actual treatment effect is better than estimated. By analyzing the error correction term, we can identify problems in the strategy and then optimize it.
[0172] Step S145: constructing a reward function based on the matching parameter and the error correction term, and calculating a reward signal through the reward function. The numerical value of the reward signal is positively correlated with the matching parameter and negatively correlated with the absolute value of the error correction term.
[0173] The reward function is constructed based on the matching parameter and the error correction term. The design of the reward function should comprehensively consider these two factors to accurately evaluate the impact of the processing action set on the processing effect of geological hazard monitoring data.
[0174] Since the matching parameter reflects the accuracy of the set of processing actions, a higher matching parameter indicates a more effective processing action. Therefore, the magnitude of the reward signal is positively correlated with the matching parameter. In other words, the larger the matching parameter, the larger the reward signal. For example, when the matching parameter increases from 0.5 to 0.8, all other conditions remaining unchanged, the reward signal will increase accordingly.
[0175] The absolute value of the error correction term reflects the accuracy of the policy evaluation. A larger absolute value indicates a greater deviation between the policy evaluation and the actual situation, which is detrimental to policy optimization. Therefore, the magnitude of the reward signal is negatively correlated with the absolute value of the error correction term. That is, the larger the absolute value of the error correction term, the smaller the reward signal. For example, increasing the error correction term from 0.1 to 0.3 will result in a decrease in the reward signal.
[0176] A common reward function construction method is to multiply the matching parameter by a positive coefficient, then subtract the absolute value of the error correction term multiplied by another positive coefficient. This calculation method produces the final reward signal. The reward signal is fed back to the reinforcement learning agent, which adjusts its strategy based on the reward signal in the hope of obtaining higher rewards in subsequent processing.
[0177] Step S150: Iteratively optimize the policy parameters of the reinforcement learning policy network according to the reward signal and the state vector until the value function estimate output by the policy evaluation module converges to a preset stable range.
[0178] After receiving the reward signal, the policy parameters of the reinforcement learning policy network need to be iteratively optimized based on the reward signal and the state vector. Iterative optimization is a process of gradually adjusting the policy parameters to continuously improve the policy. The goal is to enable the reinforcement learning agent to make better decisions and improve the effectiveness of geological disaster monitoring data processing.
[0179] Step S151: The state vector, processing action set, reward signal and next state vector are combined into an experience sample and stored in the experience replay buffer pool. The next state vector is a new state representation of the geological disaster monitoring environment after executing the processing action set.
[0180] After executing the set of processing actions, the geological hazard monitoring environment enters a new state, which is represented by the next state vector. The current state vector, the selected set of processing actions, the reward signal obtained, and the next state vector are combined to form an experience sample.
[0181] An experience sample contains complete information about the agent's interaction with the environment at a given moment. It records the agent's state, the action taken, the reward received, and the new state of the environment after the action was executed. This experience sample is stored in the experience replay buffer pool, which acts like a database for storing a large number of experience samples. As the agent continues to interact with the environment, the number of experience samples in the experience replay buffer pool continues to increase.
[0182] Step S152: randomly sampling a preset number of experience samples from the experience replay buffer pool to form a batch training sample set, which is used to reduce the variance of the policy parameter update.
[0183] For effective training, a preset number of experience samples must be randomly sampled from the experience replay buffer. This number is determined based on training requirements and computing resources. The purpose of random sampling is to break the correlation between experience samples and prevent the agent from falling into local optimal solutions during training.
[0184] A preset number of sampled experience samples are combined to form a batch training set. A batch training set contains experience samples from multiple different moments in time, covering a wide range of interactions between the agent and the environment. Using a batch training set for training can reduce the variance of policy parameter updates. If only a single experience sample is used for parameter updates at a time, the update direction may be affected by the randomness of that sample, leading to unstable updates. Using a batch training set, however, comprehensively considers information from multiple samples, making parameter updates more stable and reliable.
[0185] Step S153: Input the state vector in the batch training sample set into the policy evaluation module, calculate the estimated value of the value function of the current policy, input the next state vector in the batch training sample set into the target policy network, and calculate the estimated value of the target value function. The target policy network is a parameter lagged copy of the policy evaluation module.
[0186] After obtaining a batch of training samples, the state vectors in the sample set are input into the policy evaluation module. Based on the current policy and network parameters, the policy evaluation module evaluates the value of the corresponding action for each state vector and calculates an estimated value function under the current policy. This estimate reflects the expected return from taking different actions under the current policy.
[0187] At the same time, the next-state vector from the batch training examples is fed into the target policy network. The target policy network is a lagged copy of the policy evaluation module, and its parameters are updated relatively slowly. This provides a relatively stable target to avoid excessive fluctuations during training. The target policy network processes the next-state vector based on its own parameters and calculates an estimate of the target value function. This estimate represents the expected reward from the next state under ideal conditions.
[0188] Step S154: Calculate a temporal difference error based on the reward signal, the current value function estimate, and the target value function estimate. The temporal difference error is used to measure the deviation between the value function estimate and the target value.
[0189] Temporal difference error (TDE) is an important concept in reinforcement learning. It measures the deviation between the value function estimate and the target value. TDE is calculated by combining the reward signal, the current value function estimate, and the target value function estimate.
[0190] A common calculation method is to add the reward signal to the estimated target value function, multiply it by a discount factor, and then subtract the estimated current value function from the reward signal. The result is the temporal difference error. The discount factor is a number between 0 and 1 that indicates the degree of emphasis on future rewards. The closer the discount factor is to 1, the more emphasis is placed on future rewards; the closer it is to 0, the more emphasis is placed on current rewards.
[0191] The temporal difference error reflects the accuracy of the current strategy evaluation. If the error is large, it means that there is a large deviation between the current value function estimate and the target value, and the strategy parameters need to be adjusted. If the error is small, it means that the strategy evaluation is relatively accurate and the strategy parameters are relatively stable.
[0192] Step S155: using the temporal difference error to update the network weight parameters of the strategy evaluation module through the back propagation algorithm, and then adjusting the strategy parameters of the strategy improvement module through the gradient ascent algorithm based on the updated value function estimate.
[0193] After obtaining the temporal difference error, the backpropagation algorithm is used to update the network weight parameters of the policy evaluation module. Backpropagation is an algorithm used to calculate gradients. It uses the temporal difference error to calculate the gradient of the error with respect to each network weight parameter, layer by layer, starting from the output layer. The gradient indicates the direction and magnitude of the error as the weight parameter changes.
[0194] Based on the calculated gradient, the network weight parameters of the policy evaluation module are adjusted. The adjustment direction is to reduce the error. By continuously adjusting the weight parameters, the output of the policy evaluation module is brought closer to the target value, thereby improving the accuracy of the policy evaluation.
[0195] After updating the network weight parameters in the policy evaluation module, the policy parameters in the policy improvement module are adjusted using the gradient ascent algorithm based on the updated value function estimate. The goal of the gradient ascent algorithm is to increase the value function estimate, as a larger value function estimate indicates a better policy. By calculating the gradient of the value function estimate with respect to the policy parameters, the policy parameters are adjusted along the positive direction of the gradient, gradually optimizing the policy and enabling the agent to obtain higher rewards in subsequent interactions.
[0196] Step S156: Repeat the steps of experience sample storage, batch sampling, value function estimation, temporal difference error calculation and parameter update, calculate the variance of the value function estimation value output by the strategy evaluation module every preset number of iterations, and when the variance is less than the preset threshold, determine that the value function estimation value converges to the preset stable range, and stop the strategy parameter optimization.
[0197] The steps of experience sample storage, batch sampling, value function estimation, temporal difference error calculation, and parameter update are repeated repeatedly. Each repetition is an iteration. During the iteration process, the policy parameters are constantly adjusted and the policy is continuously improved.
[0198] After a preset number of iterations, the variance of the value function estimate output by the policy evaluation module is calculated. The variance reflects the fluctuation of the value function estimate. A large variance indicates that the value function estimate is unstable and the policy is constantly changing; a small variance indicates that the value function estimate is relatively stable.
[0199] The preset threshold is determined based on the specific application scenario and requirements. When the calculated variance is less than the preset threshold, the value function estimate is considered to have converged to the preset stable range. This means that the output of the policy evaluation module has stabilized and the policy parameters have reached a relatively optimal state. At this point, optimization of the policy parameters ceases.
[0200] After stopping optimization, the reinforcement learning policy network can be used to process actual geological disaster monitoring data. The agent can select the optimal set of processing actions based on the input state vector, processing geological disaster monitoring data efficiently and accurately, thereby improving geological disaster monitoring and early warning capabilities.
[0201] Figure 2 A schematic diagram of exemplary hardware and software components of a reinforcement learning-based geological hazard monitoring data processing system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application, is shown. For example, the processor 120 can be used in the reinforcement learning-based geological hazard monitoring data processing system 100 and is used to perform the functions of the present application.
[0202] The geological disaster monitoring data processing system 100 based on reinforcement learning can include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, for example, a disk, ROM, or RAM, or any combination thereof. Exemplarily, the geological disaster monitoring data processing system 100 based on reinforcement learning can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The geological disaster monitoring data processing system 100 based on reinforcement learning also includes an I / O interface 150 between the computer and other input and output devices.
[0203] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned geological disaster monitoring data processing method based on reinforcement learning is implemented.
[0204] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A geological disaster monitoring data processing method based on reinforcement learning, characterized in that: The method comprises: Constructing a state representation model for geological hazard monitoring data, wherein the state representation model is used to convert the continuously collected geological hazard monitoring data into a state vector recognizable by a reinforcement learning agent, wherein the state vector includes sequence correlation features and spatial correlation features of the geological hazard monitoring data; Initializing a reinforcement learning policy network, wherein the reinforcement learning policy network includes a policy evaluation module and a policy improvement module, wherein the policy evaluation module is used to calculate an estimated value function of the current policy, and the policy improvement module is used to adjust policy parameters based on the estimated value function; Executing a strategy selection operation based on the state vector and the reinforcement learning strategy network to generate a processing action set for geological disaster monitoring data, wherein the processing action set includes a data filtering operation, a feature enhancement operation, and an anomaly recognition operation; Obtaining a feedback signal from the geological hazard monitoring environment on the processing action set, and generating a reward signal based on the feedback signal and the estimated value of the value function, wherein the reward signal is used to evaluate the impact of the processing action set on the processing effect of the geological hazard monitoring data; Iteratively optimize the policy parameters of the reinforcement learning policy network according to the reward signal and the state vector until the value function estimate output by the policy evaluation module converges to a preset stable range.
2. The geological disaster monitoring data processing method based on reinforcement learning according to claim 1 is characterized in that: The construction of the state representation model of geological disaster monitoring data includes: Performing time window division processing on the continuously collected geological disaster monitoring data to obtain multiple monitoring data windows with a time sequence, each monitoring data window contains geological disaster monitoring data within a preset time length; Extracting the numerical change sequence of the geological disaster monitoring data within each monitoring data window, and calculating the correlation parameter of the numerical change sequence between adjacent monitoring data windows, wherein the correlation parameter is used to characterize the correlation degree of the monitoring data change trend within different time windows; Performing a sequence correlation analysis on multiple monitoring data windows based on the correlation parameter to generate sequence correlation features of the geological disaster monitoring data, wherein the sequence correlation features include a trend continuity index of the numerical change sequence and a distribution density of mutation points; The geological disaster monitoring data in each monitoring data window is divided into spatial dimensions to obtain monitoring data subsets of multiple spatial sub-regions, and the correlation coefficient between the monitoring data subsets of different spatial sub-regions is calculated. The correlation coefficient is used to characterize the degree of mutual influence of the geological disaster monitoring data between the spatial sub-regions; A spatial correlation feature is constructed based on the correlation coefficient, and the spatial correlation feature is subjected to feature splicing processing with the sequence correlation feature to generate a state vector including the sequence correlation feature and the spatial correlation feature, wherein the dimension of the state vector matches the state input dimension of the reinforcement learning agent.
3. The geological disaster monitoring data processing method based on reinforcement learning according to claim 2 is characterized in that: The serial correlation analysis is performed on multiple monitoring data windows based on the correlation parameter to generate serial correlation features of geological disaster monitoring data. The serial correlation features include trend continuity indicators and mutation point distribution density of numerical change sequences, including: Sorting the plurality of monitoring data windows in chronological order to obtain an ordered monitoring data window sequence, and calculating the sequence correlation strength of adjacent windows in the ordered monitoring data window sequence based on the correlation parameter, wherein the sequence correlation strength is positively correlated with the numerical value of the correlation parameter; Extracting window segments whose sequence correlation strength is continuously higher than a preset correlation threshold from the ordered monitoring data window sequence, marking the window segments as trend continuation segments, and calculating the length ratio of the trend continuation segments as the trend continuity indicator, where the length ratio is the ratio of the number of windows included in the trend continuation segments to the total number of windows in the ordered monitoring data window sequence; Perform first-order difference processing on the numerical change sequence of each monitoring data window to obtain a differential sequence, and determine the potential mutation point based on the point where the absolute value of the differential sequence exceeds the preset differential threshold; Count the total number of potential mutation points in the ordered monitoring data window sequence, and calculate the ratio of the total number of potential mutation points to the total number of windows in the ordered monitoring data window sequence as the mutation point distribution density; The trend continuity index and the mutation point distribution density are subjected to feature normalization processing, and the standardized trend continuity index and mutation point distribution density are spliced into the serial correlation feature of geological disaster monitoring data.
4. The geological disaster monitoring data processing method based on reinforcement learning according to claim 1 is characterized in that: The initialization reinforcement learning policy network, wherein the reinforcement learning policy network includes a policy evaluation module and a policy improvement module, wherein the policy evaluation module is used to calculate the estimated value function of the current policy, and the policy improvement module is used to adjust the policy parameters based on the estimated value function, includes: Determine a network architecture of a reinforcement learning policy network, wherein the network architecture comprises an input layer, a hidden layer, and an output layer, wherein the number of neurons in the input layer is the same as the dimension of the state vector, and the number of neurons in the output layer is the same as the action dimension of the processing action set; Setting initial parameters of the value function of the strategy evaluation module, wherein the initial parameters of the value function include a state feature weight matrix and a bias vector, wherein the state feature weight matrix is used to perform weighted processing on different feature dimensions of the state vector; Configuring a policy parameter update rule for the policy improvement module, wherein the policy parameter update rule determines the parameter adjustment direction and adjustment range based on the difference between the value function estimate and the target value function; Constructing an experience replay cache pool, which is used to store experience samples consisting of state vectors, processing action sets, reward signals, and next state vectors generated during the interaction between the reinforcement learning agent and the geological disaster monitoring environment; Initialize the network weight parameters of the strategy evaluation module and the strategy improvement module, wherein the network weight parameters are initialized using a random normal distribution, and the initialized network weight parameters meet the preset value range constraints.
5. The geological disaster monitoring data processing method based on reinforcement learning according to claim 4 is characterized in that: The policy parameter update rules of the configuration policy improvement module include: defining a target value function, wherein the target value function is an expected cumulative reward function constructed based on a reward signal and an estimated value function of a next state; The difference between the estimated value of the value function and the target value function is calculated as the parameter adjustment error. When the parameter adjustment error is positive, it is determined that the policy parameters need to be adjusted in the direction of increasing the estimated value of the value function. When the parameter adjustment error is negative, it is determined that the policy parameters need to be adjusted in the direction of decreasing the estimated value of the value function. Set the learning rate coefficient for policy parameter adjustment. The learning rate coefficient is used to control the amplitude of each parameter update. The value range of the learning rate coefficient is dynamically adjusted according to the training stage of the policy network. The adjustment step size of the policy parameters is calculated based on the absolute value of the parameter adjustment error and the learning rate coefficient. The larger the absolute value of the parameter adjustment error, the larger the adjustment step size; the larger the learning rate coefficient, the larger the adjustment step size; The parameter adjustment direction and the adjustment step size are combined into a policy parameter update rule, which is used to guide the policy improvement module to correct the policy parameters in each iteration process.
6. The geological disaster monitoring data processing method based on reinforcement learning according to claim 1 is characterized in that: The executing strategy selection operation based on the state vector and the reinforcement learning strategy network to generate a set of processing actions for geological disaster monitoring data includes: Inputting the state vector into the input layer of the reinforcement learning strategy network, performing nonlinear feature transformation processing on the state vector through the hidden layer to generate a high-dimensional strategy feature vector, wherein the dimension of the high-dimensional strategy feature vector is higher than the dimension of the state vector; Calling the strategy evaluation module to perform value function estimation processing on the high-dimensional strategy feature vector and calculate the expected cumulative reward value corresponding to different processing actions in the current state. The expected cumulative reward value is used to measure the long-term utility of the processing action; Execute ε-greedy strategy selection based on the expected cumulative reward value and a preset exploration rate parameter, randomly selecting a processing action within a preset probability range or selecting the processing action with the maximum expected cumulative reward value, wherein the exploration rate parameter is used to balance exploration and exploitation of strategy selection; Determining a corresponding action parameter configuration according to the selected processing action, wherein the action parameter configuration includes a filter window size of a data filtering operation, an enhancement strength coefficient of a feature enhancement operation, and a recognition threshold of an anomaly recognition operation; The processing action is combined with the corresponding action parameter configuration to generate a processing action set including a data filtering operation, a feature enhancement operation and an anomaly recognition operation, and each operation in the processing action set is executed in a preset order.
7. The method for processing geological disaster monitoring data based on reinforcement learning according to claim 6, characterized in that: The ε-greedy strategy selection is performed based on the expected cumulative reward value and a preset exploration rate parameter, randomly selecting a processing action within a preset probability range or selecting a processing action with the maximum expected cumulative reward value, and the exploration rate parameter is used to balance the exploration and utilization of the strategy selection, including: generating a random number uniformly distributed within a preset interval, and comparing the random number with a preset exploration rate parameter; When the random number is less than the exploration rate parameter, an exploration operation is performed to randomly select a processing action from the processing action set, where the random selection is based on a uniform probability distribution; When the random number is greater than or equal to the exploration rate parameter, the utilization operation is executed, and the processing action with the maximum expected cumulative reward value is selected from the processing action set. If there are multiple processing actions with the same maximum expected cumulative reward value, one of the processing actions is randomly selected; After completing a preset number of strategy selection operations, the value of the exploration rate parameter is reduced according to a preset decay rate. The decay rate is a positive number less than 1, so that the exploration rate parameter gradually decreases as the number of training iterations increases; After the exploration rate parameter decreases to the preset minimum threshold, the decay is stopped and the exploration rate parameter is kept at the preset minimum threshold.
8. The geological disaster monitoring data processing method based on reinforcement learning according to claim 1 is characterized in that: The step of obtaining a feedback signal from the geological hazard monitoring environment to the processing action set and generating a reward signal based on the feedback signal and the value function estimate includes: Applying the processing action set to the geological disaster monitoring data to obtain a processed monitoring data result, wherein the processed monitoring data result includes a filtered data sequence, an enhanced feature set, and an abnormality recognition result mark; Collecting actual monitoring data of the geological disaster monitoring environment, wherein the actual monitoring data includes geological structure change information and environmental impact factor data; Comparing and analyzing the processed monitoring data results with the actual monitoring situation data, and calculating a matching parameter between the data processing results and the actual situation, wherein the matching parameter is used to characterize the accuracy of the processing action set; Calling the policy evaluation module to obtain the estimated value of the value function corresponding to the current processing action set, and calculating the difference between the matching parameter and the estimated value of the value function as an error correction term; A reward function is constructed based on the matching parameter and the error correction term, and a reward signal is calculated through the reward function. The numerical value of the reward signal is positively correlated with the matching parameter and negatively correlated with the absolute value of the error correction term.
9. The geological disaster monitoring data processing method based on reinforcement learning according to claim 1 is characterized in that: The iteratively optimizing the policy parameters of the reinforcement learning policy network according to the reward signal and the state vector until the value function estimate output by the policy evaluation module converges to a preset stable range includes: The state vector, the processing action set, the reward signal and the next state vector are combined into an experience sample and stored in an experience replay buffer pool, wherein the next state vector is a new state representation of the geological disaster monitoring environment after executing the processing action set; Randomly sampling a preset number of experience samples from the experience replay buffer pool to form a batch training sample set, which is used to reduce the variance of the policy parameter update; Input the state vector in the batch training sample set into the policy evaluation module to calculate the value function estimate of the current policy, input the next state vector in the batch training sample set into the target policy network to calculate the target value function estimate, and the target policy network is a parameter lagged copy of the policy evaluation module; Calculating a temporal difference error based on the reward signal, the current value function estimate, and the target value function estimate, wherein the temporal difference error is used to measure the deviation between the value function estimate and the target value; Using the temporal difference error to update the network weight parameters of the strategy evaluation module through the back propagation algorithm, and then adjusting the strategy parameters of the strategy improvement module through the gradient ascent algorithm based on the updated value function estimate; Repeat the steps of experience sample storage, batch sampling, value function estimation, temporal difference error calculation and parameter update, calculate the variance of the value function estimate output by the strategy evaluation module every preset number of iterations, and when the variance is less than a preset threshold, determine that the value function estimate converges to a preset stable range, and stop strategy parameter optimization.
10. A geological disaster monitoring data processing system based on reinforcement learning, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the geological disaster monitoring data processing method based on reinforcement learning as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Marine ranch disaster decision-making method based on reinforcement learning
CN115587713A
Invertible-reasoning policy and reverse dynamics for causal reinforcement learning
WO2023167576A2
Cited By
Cross-platform user behavior analysis method and system based on transfer learning
CN120811792A
Automatic control system and method for hydrogeological test
CN121832319A