A method and system for autonomous maintenance decision of a concrete dam based on reinforcement learning

CN122840927APending Publication Date: 2026-09-29NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611073085.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]但现有维护方案对于混凝土坝维护决策的判定不够准确,使得混凝土坝的维护不及时,导致混凝土坝无法保持长期稳定工作

Benefits of technology

[0020]与现有技术相比,本发明具有以下有益效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840927A_ABST
    Figure CN122840927A_ABST
Patent Text Reader

Abstract

The application provides a kind of concrete dam autonomous maintenance decision method and system based on reinforcement learning, belongs to concrete dam maintenance management technical field.The method constructs maintenance decision as reinforcement learning task, directly with crack width, leakage flow, upstream water level and daily rainfall and so on continuous measured value as state input, with the minimum maintenance cost of full execution maintenance strategy as target, independently trains maintenance strategy in the virtual simulation environment constructed by historical data.After training, maintenance strategy decision network can dynamically output decision instruction according to the real-time state of target concrete dam, has autonomous learning, autonomous discrimination ability to crack micro-increase but leakage sharp drop and other contradictory signals, avoids waste by repairing early, repair late accident.The application changes traditional fixed cycle passive repair into on-demand active decision based on continuous monitoring data, reduces the life cycle maintenance cost under the premise of safety guarantee.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of concrete dam maintenance and management technology, and in particular to a method and system for autonomous maintenance decision-making of concrete dams based on reinforcement learning. Background Technology

[0002] Concrete dams are structures made primarily of concrete, cast or compacted, and are widely used in water conservancy and hydropower projects. During long-term operation, concrete dams are subject to the combined effects of water pressure, temperature changes, and chemical erosion, which can cause cracks and affect their stable operation (such as flood control). Therefore, timely maintenance of concrete dams is necessary to ensure their long-term stability.

[0003] However, the existing maintenance plan is not accurate enough in making decisions on the maintenance of concrete dams, which leads to untimely maintenance of concrete dams and makes it impossible for concrete dams to maintain long-term stable operation. Summary of the Invention

[0004] This invention provides a method and system for autonomous maintenance decision-making of concrete dams based on reinforcement learning. The method uses a reinforcement learning-based maintenance strategy decision-making network to analyze the state vector of a target concrete dam to determine its maintenance strategy. Since this network is trained using training vectors capable of identifying self-healing and pseudo-deterioration signals of the concrete dam, its decision analysis of the state vectors can effectively identify these signals, improving the accuracy of the maintenance strategy and enabling more timely maintenance. This, in turn, helps the concrete dam maintain long-term stable operation.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a reinforcement learning-based autonomous maintenance decision-making method for concrete dams, comprising: firstly, acquiring the state vector of a target concrete dam, the state vector indicating the cracks and leakage conditions of the concrete dam; then, using the state vector of the target concrete dam and a reinforcement learning-based maintenance strategy decision network, determining the maintenance strategy for the target concrete dam; wherein, the maintenance strategy decision network is trained using multiple training vectors and corresponding maintenance strategies; the training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating the maintenance strategy; the self-healing and pseudo-deterioration signals include crack self-healing signals and rainfall leakage signals; the maintenance strategy is to maintain the natural state or to repair it.

[0006] In one implementation of the first aspect, the training vectors are obtained through simulation using a state prediction model. The state prediction model includes a natural state prediction sub-model and an attenuation coefficient. The natural state prediction sub-model is used to identify self-healing and pseudo-deterioration signals of the concrete dam and simulate the state vector of the concrete dam in the current time period to obtain the state vector of the concrete dam in the next time period when the maintenance strategy is to maintain its natural state. The attenuation coefficient represents the difference between the state vector of the repaired concrete dam and the state vector of the concrete dam in its natural state; the product of the attenuation coefficient and the state vector of the concrete dam in its natural state in the previous time period is the state vector of the concrete dam in the next time period when the maintenance strategy is repair.

[0007] In one implementation of the first aspect, the natural state prediction sub-model is trained with physical constraint terms constructed from self-healing and pseudo-deterioration signals.

[0008] In one implementation of the first aspect, the physical constraint terms satisfy: in, Represents physical constraint terms. This indicates that the physical constraints are determined by the self-healing confidence level. .

[0009] This indicates the confidence level of the self-healing term for cracks. Used to measure the degree to which crack growth slows down satisfy: This indicates the confidence level of the rainfall infiltration term. Used to measure the degree of continuous decrease in leakage. satisfy: Indicates the slowdown threshold. Indicates the steepness coefficient. Indicates time t Increment of crack width over time; This indicates the number of consecutive days the leakage rate has decreased. Indicates the threshold for the number of days of decline. This indicates the cumulative net decrease in leakage within the preset window period. Indicates the cumulative decline threshold; Indicates time t Incremental leakage flow rate This indicates the upper limit of allowed fluctuations.

[0010] In one implementation of the first aspect, the natural state prediction sub-model is a supervised learning regression model.

[0011] In one implementation of the first aspect, the maintenance strategy of the target concrete dam is determined by using the state vector of the target concrete dam and a maintenance strategy decision network based on reinforcement learning. This includes: inputting the state vector of the target concrete dam into the maintenance strategy decision network, analyzing the state vector, and outputting the maintenance strategy of the target concrete dam.

[0012] In one implementation of the first aspect, the training process of the maintenance strategy decision network includes: constructing a candidate network, which comprises a first policy network and a first value network, both connected to the feature extraction layer at their inputs; the first policy network outputs a maintenance policy; and the first value network outputs the maintenance cost corresponding to the maintenance policy. An experience set is constructed using multiple training vectors and corresponding maintenance policies; the experience set includes multiple experiences, each of which includes a training vector, maintenance policy, and maintenance cost for the concrete dam within an execution cycle; the execution cycle is the time required to execute the maintenance policy. With the objective of minimizing the maintenance cost output by the first value network, the candidate network is trained and updated using gradient descent and the experience set, resulting in a trained candidate network; the trained candidate network includes a second policy network and a second value network. The second policy network is then determined as the maintenance strategy decision network.

[0013] In one implementation of the first aspect, the feature extraction layer includes a fully connected layer. The first policy network includes a first hidden layer and a softmax output layer connected in sequence. The first value network includes a second hidden layer and a linear output layer connected in sequence.

[0014] In one implementation of the first aspect, maintenance is classified as minor or major repair, and minor and major repairs are determined by clustering algorithms.

[0015] Secondly, this invention provides an autonomous maintenance decision-making system for concrete dams based on reinforcement learning, including a state vector acquisition module and a maintenance strategy determination module. The state vector acquisition module acquires the state vector of the target concrete dam; the state vector indicates the cracking and leakage conditions of the concrete dam. The maintenance strategy determination module uses the state vector of the target concrete dam and a reinforcement learning-based maintenance strategy decision network to determine the maintenance strategy for the target concrete dam; wherein, the maintenance strategy decision network is trained using multiple training vectors and corresponding maintenance strategies; the training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating the maintenance strategy; the self-healing and pseudo-deterioration signals include crack self-healing signals and rainfall leakage signals; the maintenance strategy is to maintain the natural state or to repair it.

[0016] Thirdly, the present invention provides an electronic device including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the method described in the first aspect above or any implementation thereof.

[0017] Fourthly, the present invention provides a computer-readable storage medium including computer program instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.

[0018] Fifthly, the present invention provides a computer program product, including computer program instructions, which, when executed on a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.

[0019] The technical effects corresponding to the second to fifth aspects and their possible implementations can be referred to the above description of the technical effects of the first aspect and its possible implementations, and will not be repeated here.

[0020] Compared with the prior art, the present invention has the following beneficial effects.

[0021] The reinforcement learning-based autonomous maintenance decision-making method for concrete dams provided by this invention first acquires the state vector of the target concrete dam indicating cracks and leakage. Then, a reinforcement learning-based maintenance strategy decision-making network analyzes the state vector to derive the maintenance strategy for the target concrete dam. In this process, the maintenance strategy decision-making network is trained using multiple training vectors and corresponding maintenance strategies. Since these training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating maintenance strategies, the network's decision analysis of the state vectors can identify these signals, improving the accuracy of the maintenance strategy and enabling more timely maintenance, thus contributing to the long-term stable operation of the concrete dam. Attached Figure Description

[0022] Figure 1 This is one of the schematic diagrams of an autonomous maintenance decision-making method for concrete dams based on reinforcement learning provided in the embodiments of this application; Figure 2 This is a schematic diagram of the training process of the maintenance strategy decision network provided in the embodiments of this application; Figure 3 This is a schematic diagram of the network structure of the candidate network provided in the embodiments of this application; Figure 4This is a schematic diagram of the training process of the natural state prediction sub-model provided in the embodiments of this application; Figure 5 This is a schematic diagram of the training process of the candidate network provided in the embodiments of this application; Figure 6 This is the second schematic diagram of a reinforcement learning-based autonomous maintenance decision-making method for concrete dams provided in the embodiments of this application; Figure 7 This is a flowchart illustrating a reinforcement learning-based autonomous maintenance decision-making method for concrete dams provided in an embodiment of this application. Figure 8 This application provides a schematic diagram of the structure of an autonomous maintenance decision-making system for concrete dams based on reinforcement learning. Detailed Implementation

[0023] In the specification and claims of this invention, the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order of objects.

[0024] In the embodiments of this application, "and / or" indicates a relationship between objects. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist simultaneously.

[0025] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0026] In the description of this invention, unless otherwise stated, "multiple" means two or more. For example, multiple training vectors means two or more training vectors.

[0027] The method and system provided in this application relate to the maintenance and management of concrete dams, and can use a maintenance strategy decision network based on reinforcement learning to determine the maintenance strategy of a target concrete dam.

[0028] Understandably, during long-term operation, concrete dams are subject to the combined effects of water pressure, temperature changes, and chemical erosion, leading to continuous degradation of structural indicators such as cracks and leakage. Scientifically determining the timing of maintenance is a core issue in balancing the safety and economy of concrete dams.

[0029] Existing maintenance decision-making methods for concrete dams mainly fall into two categories. The first is the fixed-cycle maintenance method, which involves periodic maintenance according to a preset timeframe. However, this method is disconnected from the actual degradation state of the concrete dam. If timely repairs are not possible, premature repairs lead to wasted funds, while delayed repairs may allow small cracks to develop into major structural damage. The second category is alarm methods based on single physical quantity thresholds. These methods set limits for crack width or leakage flow, triggering an alarm when these limits are exceeded. However, cracks and leakage in concrete dams are not isolated changes; there is a physical coupling between them. For example, slow crack growth coupled with decreasing leakage is a normal manifestation of material self-healing after the cracks are filled with calcium carbonate deposits. The single-threshold method, focusing only on crack growth, may misinterpret this contradictory signal as deterioration, leading to ineffective maintenance. Similarly, a sudden increase in leakage due to heavy rain without any change in crack size is a temporary, pseudo-deterioration induced by the external environment, and the single-threshold method will also generate false alarms. The root cause of these shortcomings is that these methods lack the ability to identify the coupled degradation patterns of multiple physical quantities and cannot dynamically balance maintenance costs with failure risks throughout the entire lifecycle.

[0030] To address the problem that existing maintenance schemes in the background art are not accurate enough in determining maintenance decisions for concrete dams, resulting in untimely maintenance and the inability of concrete dams to maintain long-term stable operation, this application provides a reinforcement learning-based autonomous maintenance decision-making method and system for concrete dams. The method uses a reinforcement learning-based maintenance strategy decision-making network to perform decision analysis on the state vector of the target concrete dam to determine the maintenance strategy. Since the maintenance strategy decision-making network is trained using training vectors capable of identifying self-healing and pseudo-deterioration signals of the concrete dam, its decision analysis of the state vector can identify these signals, improving the accuracy of the maintenance strategy and ensuring more timely maintenance of the concrete dam, thereby contributing to the long-term stable operation of the concrete dam.

[0031] For example, the reinforcement learning-based autonomous maintenance decision-making method for concrete dams provided in this embodiment of the invention can be executed by an electronic device with processing capabilities, such as a computer or server. Taking a computer as an example, the hardware components of the computer may include: a processor, memory, a network interface, a user interface, a communication bus, etc.

[0032] The processor is used to control electronic devices to perform related processing and computation tasks, such as obtaining the state vector of the target concrete dam and determining the maintenance strategy of the target concrete dam. The processor may include a central processing unit (CPU), an AI processing unit (such as a GPU, TPU, NPU, LPU, etc.), or other processors. The processor can be single-core or multi-core; for example, the processor may include multiple CPUs or multiple GPUs.

[0033] Memory is used to store computer instructions and related data, such as the state vector of a target concrete dam, the maintenance strategy decision network, and the maintenance strategy for the target concrete dam. Memory can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, disk storage media, or other magnetic storage devices, or any other medium capable of storing program code or data accessible by a computer. Optionally, memory can be integrated into the processor, or it can be independent of the processor.

[0034] A network interface is used for communication between a computer and other devices or communication networks. A network interface can be a transceiver with transmit and receive capabilities. Optionally, a network interface may include standard wired interfaces or wireless interfaces (such as Wi-Fi interfaces, Bluetooth interfaces, and 5G interfaces).

[0035] The communication bus is used to enable communication between different components. For example, the processor, memory, network interface and user interface mentioned above can be interconnected through the communication bus.

[0036] The user interface may include a display screen and an input unit (such as a keyboard). Optionally, the user interface may also include a standard wired interface or a wireless interface.

[0037] Those skilled in the art will understand that the computer described above may include more or fewer components, or combine certain components, or have different component arrangements; the embodiments of this application do not limit this.

[0038] like Figure 1 As shown in the figure, the autonomous maintenance decision-making method for concrete dams based on reinforcement learning provided in this application includes S101-S102.

[0039] S101. Obtain the state vector of the target concrete dam.

[0040] In this embodiment, the aforementioned state vector indicates the cracking and leakage conditions of the concrete dam. Optionally, the state vector may include crack width, leakage flow rate, upstream water level, and rainfall, etc. The state vector can be acquired on a daily, five-day, ten-day, or monthly basis; when acquired on a daily basis, the state vector may include crack width, daily leakage flow rate, upstream water level, and daily rainfall, etc. This embodiment does not further limit the parameters included in the state vector or the acquisition frequency.

[0041] In one application scenario, the process of obtaining the state vector of the target concrete dam is as follows: monitoring devices such as crack gauges, piezometers, and displacement sensors are installed around the target concrete dam, and the crack width of the target concrete dam is obtained through these monitoring devices. L Leakage flow Q Upstream water level H and rainfall P The state vector that constitutes the target concrete dam S =[ L , Q , H , P ].

[0042] In one embodiment of the above application scenario, the Three Gorges Dam is used as the target concrete dam. It is understood that the Three Gorges Dam is a concrete gravity dam with a maximum height of 181 meters and a total installed capacity of 22.5 million kilowatts. The first generating units began operation in 2003, and the entire project was completed in 2009. It has been in operation for over 20 years, with a remaining design service life of approximately 70 years. The dam safety monitoring system collects daily data on crack width, seepage flow, upstream water level, and daily rainfall. Historical monitoring records are continuous and complete since the construction period. Historical maintenance archives record all maintenance events during the operational period. Therefore, the daily crack width data from 2003 to 2024 can be extracted from the Three Gorges Dam's dam safety monitoring system. Leakage flow Upstream water level and daily rainfall Data, constructing the daily state vector of the Three Gorges Dam: .

[0043] S102. Using the state vector of the target concrete dam and a maintenance strategy decision network based on reinforcement learning, determine the maintenance strategy of the target concrete dam.

[0044] Specifically, the state vector of the target concrete dam S After inputting the maintenance strategy decision network, the maintenance strategy decision network processes the state vector. SThe analysis outputs a maintenance strategy for the target concrete dam, which can be either maintaining its natural state or undergoing repairs. Continuing with an example application scenario from S101, a reinforcement learning-based maintenance strategy decision-making network is deployed in the Three Gorges Dam's control server or edge computing device. This network connects to the monitoring equipment via a local area network. The monitoring equipment then encapsulates the monitored data into a state vector. S Input the above maintenance strategy decision network, the above maintenance strategy decision network on the state vector S After analysis, the maintenance strategy for the target concrete dam is output.

[0045] Optionally, the aforementioned maintenance can be either minor or major. Minor maintenance corresponds to a maintenance strategy with shorter maintenance time and lower cost, while major maintenance corresponds to a maintenance strategy with longer maintenance time and higher cost. This distinction between the two types of maintenance strategies can be made using clustering algorithms (such as K-Means clustering). In some cases, when the number of historical maintenance events is insufficient to support clustering, the median cost can be used as a boundary: events above the median are classified as major maintenance, and those below the median as minor maintenance.

[0046] For example, the direct cost and actual duration of each maintenance event in multiple maintenance events can be used as clustering features. A clustering algorithm can be used to divide the aforementioned multiple maintenance events into major repair groups and minor repair groups, and the event vector of minor repairs can be calculated for each group. Event vectors for overhaul .

[0047] The above event vector It can satisfy the following formula (1).

[0048] in, To maintain the strategy type, the value is set to minor repair. Major overhaul ; To determine the maintenance duration, take the median of the number of days for similar maintenance cycles; Repair rate for crack width The mean, Repair rate for crack width Standard deviation; For direct expenses; To account for operational losses during maintenance, the median of operational losses for similar maintenance periods was used. The above crack width repair rate... It satisfies the following formula (2).

[0049] in, The width of the cracks in the concrete dam at the end of maintenance. This refers to the width of the cracks in the concrete dam before repairs.

[0050] Furthermore, when the maintenance strategy output by the maintenance strategy decision network is minor repair, the event vector of minor repair can be... As a reference for maintenance strategies; when the maintenance strategy output by the maintenance strategy decision network is a major overhaul, the event vector of the major overhaul can be used. As a reference for maintenance strategies.

[0051] Continuing with the above embodiments, clustering results show that the first group had an average cost of approximately 800,000 yuan and an average time of approximately 15 days, corresponding to minor repairs; the second group had an average cost of approximately 12 million yuan and an average time of approximately 65 days, corresponding to major repairs. For each repair within the minor repair group, the crack width repair rate was calculated. The average repair rate for minor repairs was calculated and statistically analyzed using the above formula (2). Standard deviation Average repair rate of major overhauls Standard deviation Since the improvement effect of leakage flow is reflected in the decay curve, no separate repair rate index is set. Therefore, the median time for minor repair is 15 days, the median direct cost is 800,000 yuan, and the operating loss is 0 (minor repair is carried out in the corridor and does not affect power generation). The median time for major repair is 65 days, the median direct cost is 12 million yuan, and the operating loss is calculated based on the average daily power generation loss during the shutdown period. The Three Gorges Dam generates an average of about 240 million kWh per day, with an on-grid electricity price of about 0.25 yuan / kWh and an average daily power generation revenue of about 60 million yuan. If the major repair is completely shut down, the operating loss for 65 days is about 3.9 billion yuan; in actual major repairs, only some units are shut down, and the median operating loss in historical major repair records is about 1.2 million yuan. The event vector of minor repair is obtained by corresponding to the above formula (1). Event vectors for overhaul , , .

[0052] From the perspective of network training, the above-mentioned maintenance strategy decision network is trained by multiple training vectors and corresponding maintenance strategies. The training vectors are the state vectors of the concrete dam after identifying the self-healing and pseudo-deterioration signals of the concrete dam and simulating the maintenance strategy. The self-healing and pseudo-deterioration signals include crack self-healing signals and rainfall leakage signals.

[0053] In one implementation, such as Figure 2 As shown, the training process of the maintenance strategy decision network can include the following steps 1-4.

[0054] Step 1: Construct candidate networks.

[0055] The aforementioned candidate network includes a feature extraction layer, a first policy network, and a first value network. The input of the feature extraction layer is coupled to the input of the candidate network, and the output of the feature extraction layer is connected to the inputs of both the first policy network and the first value network. The outputs of the first policy network and the first value network are coupled to the output of the candidate network. In the candidate network, the feature extraction layer may include a fully connected layer; the first policy network may include a first hidden layer and a softmax output layer connected in sequence; and the first value network may include a second hidden layer and a linear output layer connected in sequence.

[0056] Furthermore, the aforementioned feature extraction layer receives and extracts the state vector. S Based on the vector features in the data, the first policy network performs decision analysis on the vector features to determine the maintenance strategy among the maintenance strategies. It should be understood that the first policy network outputs the probability of maintaining the natural state and the probability of repair, and determines the maintenance strategy with the higher probability as the maintenance strategy among the maintenance strategies. The output of the first value network performs value analysis on the vector features and outputs the maintenance cost among the maintenance strategies.

[0057] Continuing with an example of an application scenario in S101, such as... Figure 3 As shown, the feature extraction layer is a fully connected layer with 128 neurons and the activation function is tanh. The first hidden layer in the first policy network contains 64 neurons, and the softmax output layer contains 3 neurons, outputting the probability of maintaining the natural state. And the probability of repair, the above probability of repair includes the probability of minor repair. and the probability of major repairs The second hidden layer of the first value network described above contains 64 neurons, and the linear output layer contains 1 neuron, outputting the state value. The training parameters were set as follows: experience buffer capacity 4096 records, mini-batch size 128 records, number of training epochs 10, and pruning factor. GAE smoothing parameters Learning rate .

[0058] Step 2: Build and train the state prediction model.

[0059] In some embodiments, the state prediction model may include a natural state prediction sub-model and a decay coefficient. The natural state prediction sub-model can provide a training vector when the maintenance strategy is to maintain the natural state. The natural state prediction sub-model combined with the decay coefficient can provide a training vector when the maintenance strategy is to repair.

[0060] The aforementioned natural state prediction sub-model is used to identify self-healing and pseudo-deterioration signals of concrete dams and simulate the state vector of the concrete dam in the current time period to obtain the state vector of the concrete dam in the next time period when the maintenance strategy is to maintain the natural state. For example, the aforementioned natural state prediction sub-model can be a supervised learning regression model (such as XGBoost) or other network models. This application embodiment does not limit the type of the aforementioned natural state prediction sub-model.

[0061] Taking the above natural state prediction sub-model as an example of a supervised learning regression model, such as Figure 4 As shown, the training process of the above natural state prediction sub-model may include the following steps 2.1-2.2.

[0062] Step 2.1: Construct the state vector dataset.

[0063] Obtain the crack length of multiple concrete dams within the first time period. Leakage flow Upstream water level and daily rainfall The daily time-series data were obtained, and missing values ​​were filled in using linear interpolation to obtain multi-day state vectors for multiple concrete dams. ,and The first time period mentioned above refers to the non-maintenance period of the concrete dam and the period during which the concrete dam structure is not affected by maintenance. For example, for some maintenance strategies with short maintenance time and low cost (corresponding to small changes in the concrete dam structure), the period during which the concrete dam structure is affected by maintenance can be determined as 1 year; for some maintenance strategies with long maintenance time and high cost (corresponding to large changes in the concrete dam structure), the period during which the concrete dam structure is affected by maintenance can be determined as 3 years.

[0064] The daily state vectors are selected from the multi-day state vectors contained in the aforementioned state vector dataset. The change vector for the next day was calculated. and the above daily state vectors and the above-mentioned next-day change vector After normalization, the normalized daily state vector is obtained. and normalized next-day change vector Integrating the above and Obtain the state vector dataset.

[0065] Step 2.2: Construct and train a natural state prediction sub-model using the state vector dataset.

[0066] A supervised learning regression model was selected, with the daily state vector as input and the actual change on the next day as the learning objective. By minimizing the loss function (the weighted sum of data fitting loss and physical constraint penalty term), a natural state prediction sub-model that grasps the state evolution law of concrete dams under no-intervention conditions was obtained.

[0067] During the training of the aforementioned natural state prediction sub-model, the model receives the denormalized state vector of the concrete dam for that day. Output the change vector for the next day. The predicted value for the next day's state is obtained by adding the current day's state to the predicted change. After training, the model's prediction accuracy is evaluated using test set data. Input vectors from the test set are fed into the model one by one, and the root mean square error between the predicted and actual changes is calculated. The model is considered trained successfully when the error meets preset accuracy requirements (e.g., prediction error for crack width change not exceeding 0.01 mm, prediction error for leakage flow rate change not exceeding 0.02 L / min, etc.).

[0068] Optionally, the loss function described above can be constructed using physical constraint terms derived from self-healing and pseudo-deterioration signals. Then the loss function described above... It satisfies the following formula (3).

[0069] in, This represents the data fitting loss term, which is the mean squared error between the predicted state vector increment and the actual state vector increment. express ; This represents a physical constraint term. The higher the self-healing confidence, the stronger the physical constraint penalty, and vice versa.

[0070] The above It satisfies the following formula (4).

[0071] in, This indicates that the physical constraints are determined by the self-healing confidence level. .

[0072] The above This indicates the confidence level of the self-healing term for cracks. It is used to measure the degree of slowdown in crack growth, and is obtained by comparing the daily crack increment with the historical average crack increment. It satisfies the following formula (5).

[0073] The above This indicates the confidence level of the rainfall infiltration term. It is used to measure the degree of continuous decrease in leakage, and is obtained by multiplying the ratio of the number of consecutive days of leakage decrease to the saturation threshold and the ratio of the cumulative net decrease within the window period to the saturation decrease threshold. It satisfies the following formula (6).

[0074] The above Indicates the slowdown threshold. Indicates the steepness coefficient. Indicates time t Increment of crack width over time; This indicates the number of consecutive days the leakage rate has decreased. Indicates the threshold for the number of days of decline. This indicates the cumulative net reduction in leakage within a preset window period. The length is taken as 2 to 3 times the average duration of a continuous decline in leakage during the non-maintenance period in historical timeframes, to ensure the preset window period. It can completely cover a typical leakage descent process; Indicates the cumulative decline threshold; Indicates time t Incremental leakage flow rate This represents the upper limit of permissible fluctuations, calculated as the upper quartile of the non-maintenance period leakage increase in historical data.

[0075] Continuing with an example from the application scenario in S101, we will filter daily data outside of the non-maintenance period and the preset maintenance impact period. Construct an input vector from the filtered data. and output vector The first 80% of the samples, ordered chronologically, were used as the training set, and the last 20% as the test set. All inputs were Z-score normalized. XGBoost was chosen as the supervised learning regression model. During training, the loss function... The values ​​are selected from {0.01, 0.05, 0.1, 0.5, 1.0} via a grid search. In this embodiment, the values ​​are... .

[0076] Regarding loss calculation, the embodiment uses the historical average daily increase in cracks during non-maintenance periods. mm / day, standard deviation mm / day, slowing down threshold mm / day, steepness coefficient The number of consecutive days of declining leakage during historical non-maintenance periods is above the upper quartile. The cumulative decline is in the upper quartile. Upper quartile of leakage increment L / min. Daily crack increment during a certain prediction. mm / day, calculated Leakage has been decreasing continuously. Daily, cumulative net decline Calculated Self-healing confidence If the model predicts the increase in leakage at this point... L / min, exceeding ,but Apply appropriate punishment. If the crack grows faster, ,but , The penalty will automatically become invalid.

[0077] After training, the results were evaluated using a test set. The root mean square error (RMSE) for predicting crack width variation was 0.007 mm, and the RMSE for predicting leakage flow variation was 0.013 L / min, both meeting the preset accuracy requirements.

[0078] The aforementioned attenuation coefficient represents the difference between the state vector of the repaired concrete dam and the state vector of the concrete dam in its natural state. The product of the aforementioned attenuation coefficient and the state vector of the concrete dam in its natural state over a previous time period can be the state vector of the concrete dam in the next time period when the maintenance strategy is repair.

[0079] The following is a calculation process for the above attenuation coefficient, including steps 2.3-2.4.

[0080] Step 2.3: Obtain the state sequence of the concrete dam over several consecutive days after the repair.

[0081] Step 2.4: Using the natural state prediction sub-model trained in Step 2.2 above, predict the state sequence of the concrete dam for several consecutive days after the above maintenance to obtain the predicted change vector for the next day. The attenuation coefficient is calculated by using the increment of the predicted state vector and the actual change vector for the next day.

[0082] Specifically, take the date after the repair. State vector of concrete dam Input the natural state prediction sub-model to predict the next day's change vector under the natural state. Simultaneously, the actual next-day change vector is obtained from the state sequence in step 2.3. The attenuation coefficient for the day is calculated based on the crack width, and the calculation formula is shown in formula (7) below.

[0083] Continue to perform maintenance on the continuous n The daily state vector increment is calculated to obtain the maintenance attenuation coefficient sequence. For all historical maintenance of the same maintenance type, the attenuation coefficient sequences of each maintenance are aligned day by day and then averaged to obtain the average attenuation coefficient sequence for that type of maintenance. .

[0084] The above average decay coefficient sequence was analyzed using the exponential recovery function. The attenuation coefficient is obtained by performing least squares fitting. The above attenuation coefficient It satisfies the following formula (8).

[0085] in, Indicates time t The attenuation coefficient of the concrete dam after maintenance compared to the concrete dam under natural conditions. Indicates the initial attenuation coefficient. This represents the recovery time constant.

[0086] After obtaining the natural state prediction sub-model and attenuation coefficient through steps 2.1 to 2.4 above, when the state prediction model receives the state vector of the concrete dam on that day... And the number of days since the last maintenance was completed. First, the natural state prediction sub-model is called to predict the increment of the state vector for the next day. Then Multiply by the attenuation coefficient The output maintenance strategy is the change vector of the day following the maintenance. The state vector of the concrete dam on that day Change vector of the next day The sum, as the maintenance strategy, serves as the state vector of the concrete dam on the day following maintenance. The above. It can satisfy the coefficient formula.

[0087] in, The attenuation coefficient representing the crack width, The attenuation coefficient representing the leakage flow rate of the crack is mentioned above. and The calculation method can be referred to the above formula (8), and will not be repeated in the embodiments of this application.

[0088] Continuing with an example from the application scenario in S101, during the calculation of the attenuation coefficient, for each minor repair, a continuous 365-day state sequence is taken after the repair is completed. The attenuation coefficient is calculated daily. This yields the attenuation coefficient sequence for this maintenance. The attenuation sequences for all maintenance in the minor repair group are averaged daily, and the exponential recovery function is applied. The initial attenuation coefficient of the minor repair is obtained by performing least squares fitting. Recovery time constant Heaven. The same principle applies to major repairs; thus, one obtains... , sky.

[0089] Step 3: Construct an experience set using multiple training vectors obtained from the state prediction model simulation and the corresponding maintenance strategies.

[0090] In this embodiment, the aforementioned experience set may include multiple experiences, each of which includes a training vector, maintenance strategy, and maintenance cost for the concrete dam within an execution cycle. The training vector refers to the state vector simulated by the aforementioned state prediction model. The execution cycle refers to the time required to execute the maintenance strategy. Optionally, the aforementioned experience set can be constructed by encapsulating the aforementioned state prediction model into an engine. The engine encapsulation process is as follows.

[0091] With the set execution cycle T The number of days included represents the number of executions. The process involves iteratively inputting the state vector of the concrete dam for each day into the aforementioned state prediction model, which then outputs the change for the following day. Finally, the execution cycle is returned. T New state after internal maintenance strategy ends This award R and safety red line signs done Three variables.

[0092] It should be noted that the above execution cycle T This refers to the sum of the time required to perform maintenance actions and the time during which the maintenance actions affect the concrete dam structure. The aforementioned new state... This refers to the execution cycle of the predictive model simulation. T The state vector of the concrete dam on the final day. The aforementioned safety red line marker. done This is used to indicate whether a concrete dam has collapsed and is no longer functional. The collapse of a concrete dam is primarily determined by whether either the crack width or the leakage flow exceeds a failure threshold. The time indicates that the concrete dam has collapsed; otherwise, the concrete dam has not collapsed.

[0093] The above-mentioned awards R For negative maintenance costs, the above R It can be calculated using the following formulas (9) and (10).

[0094] The above Maintenance costs are direct expenses. and operational losses The sum of the above. For risk costs, The following formulas (11) and (12) are satisfied.

[0095] The above , The failure threshold is set for a predetermined crack width and leakage flow rate; Basic cost coefficient; Risk sensitivity coefficient; This is a risk index. (The above...) and The risk cost is determined by simultaneously solving a problem using two engineering anchor points. First, the expected risk cost is set under two typical conditions. Anchor point one represents the safe state (e.g., both cracks and leaks are at 30% of the failure threshold, with a risk index of...). =0.18), at this point the daily risk cost should be negligible, set as C1 = 0.001 million yuan. Anchor point two is a dangerous state (e.g., cracks and leaks are both at 90% of the failure threshold, risk index). = 1.62), at this point the daily risk cost should force the agent to act immediately, let's say C2 = 100,000 yuan. Substituting the two sets of values ​​into the above formula (11) yields two equations, which can be solved to obtain and .

[0096] After encapsulating the aforementioned state prediction model into an engine, a state vector is selected from the state vectors of the concrete dam over a historical time period as the initial state of the concrete dam. This initial state is then input into the first policy network, which outputs the probabilities of various maintenance policies and randomly selects a maintenance policy according to the probability distribution of the output. The initial state and the maintenance policy are then input into the encapsulated state prediction model to execute the cycle of that maintenance policy. T And return the result after the execution cycle ends. , R and done Record as an experience Repeating the above process yields multiple experiences, which are then integrated to form an experience set. Understandably, this experience set includes multiple experiences, each of which contains the state vector of the concrete dam, maintenance strategy, and maintenance cost for a given execution cycle.

[0097] Optionally, this experience It can satisfy: From a storage perspective, the cache contains a sequence of experiences within a round, arranged chronologically and numbered t=0,1,2,...,T, corresponding to the 1st, 2nd,...,T+1th decisions, respectively. The state vector at the t-th decision is denoted as... After a decision is made, a maintenance strategy is implemented. The new state that the concrete dam enters after the execution cycle ends is the state for the next decision. .

[0098] Continuing with an example from the application scenario in S101, we set a crack width failure threshold. mm, leakage flow failure threshold L / min. Solve for the basic cost coefficient. and risk sensitivity coefficient Under safe conditions (both cracks and leaks are at 30% of the failure threshold, risk index) The daily risk cost target is 0.001 million yuan; under hazardous conditions (both cracks and leaks are at 90% of the failure threshold, risk index...), the risk is... The daily risk cost target is 800,000 yuan. (Solution) , Take the remaining service life of the Three Gorges Dam. Average time span of a single decision per year Years, count Set the values ​​for each action parameter to: No maintenance. Execution time Day, direct cost Operational losses Minor repairs Execution time Day, direct cost RMB 10,000, operating loss Overhaul Execution time Day, direct cost RMB 10,000, operating loss Ten thousand yuan. After encapsulation, the engine returns to a new state after the action cycle has ended. This award R and safety red line signs done Three variables.

[0099] Continuing with the description of the above embodiments, an initial state is randomly selected from historical safe states, such as measured data on a certain day. The first policy network outputs the probability distribution of the three actions. Actions are randomly sampled according to this probability. (Not maintained). and Input the virtual engine, which takes the current day as the starting point, extracts 30 consecutive days of water level and rainfall sequences from the same period in other historical years of the Three Gorges Dam, calls the state prediction model under natural conditions to update the state and accumulates risk costs daily, and returns to the new state after 30 days. ,award and The first strategy network in Next choice log probability Experience Store in the cache. To continue sampling for the new current state, each decision generates experience points that are stored in a cache until the cache is full (4096 entries). If after a certain step... The round ends.

[0100] Step 4: With the goal of minimizing the maintenance cost of the first value network output, use gradient descent and the above experience set to perform reinforcement learning on the candidate network.

[0101] In one implementation, such as Figure 5 As shown, step 4 above may include steps 4.1-4.3 below.

[0102] Step 4.1, Learning from experience.

[0103] Following the sorted order of the experience set, the process is performed in reverse order, starting from the last experience (t=T). For example, for the t-th decision, the current state is taken. and the next state The first value network was used to value the results separately. and .like If the state is terminated (i.e., done=1), then the first value network valuation is used to obtain the result. .

[0104] After obtaining the valuation of the first value network, the advantage value of experience and the target return are calculated. Specifically, the prediction deviation at step t is calculated first. The calculation formula is: ;in As a discount factor, satisfying ,in The average time span of a single decision. This is a time constant, representing the effective decision-making horizon; a larger value is taken when the remaining years are long or the risk appetite is low, and a smaller value is taken when the remaining years are short or the cost control pressure is high.

[0105] Then calculate the advantage value at step t. The calculation formula is: ;in For smoothing parameters, when When the value approaches 0, the dominance estimate is mainly determined by the prediction bias of the current step. Decision. When When the value approaches 1, the advantage estimation incorporates more real reward information for subsequent steps. The preset value is determined based on the estimation accuracy of the first value network during training. When the estimation accuracy is high, a smaller value (e.g., 0.9) is used to reduce variance. When the estimation accuracy is low, a larger value (e.g., 0.95 or 0.99) is used to introduce more real rewards to compensate for estimation bias. This represents the advantage value calculated at step t+1. For the last piece of experience (t=T)... It is considered as 0.

[0106] Finally, calculate the target return at step t. The calculation formula is: .

[0107] Then, starting from t=T, repeat the above process until t=0 is reached, and for each experience... Assign the corresponding and Then, the order of experience is randomly shuffled.

[0108] Continuing with an example from the application scenario in S101, we process 4096 experience records in the cache, arranged by time, in reverse order. We take the second-to-last record in this example... ) and the last one ( Experience demonstration of calculation process: Let the last experience be... of ,action ,award , , .Will and Inputting the first value network into each network yields the following results: , Calculate the predicted deviation. .because No follow-up steps, advantage value Target return Forward processing ,set up , , , Using the calculated Recursion: , Process sequentially until... Each piece of experience is assigned a corresponding [characteristic / characteristic]. and Then, the order of experience is randomly shuffled.

[0109] Step 4.2: Update the first value network to obtain the second value network.

[0110] Learn from small batches of experience in stages, and take the state of experience for each batch. and the corresponding target return ,Will Input the first value network and obtain the current estimated value through forward propagation. Calculate the mean squared error loss. Mean squared error loss The calculation formula is as follows.

[0111] Again Find the gradient, which satisfies the following formula.

[0112] The gradient above represents the deviation between the projected return and the target return. A positive deviation indicates that the projected return is too high, the gradient is positive, and the parameters are adjusted to decrease the projected return. A negative deviation indicates that the projected return is too low, the gradient is negative, and the parameters are adjusted to increase the projected return.

[0113] After obtaining the gradients described above, a chain rule is used to propagate these gradients backward from the output layer of the first value network to the input layer. For example, suppose the first value network has a total of... Layer, number The weight matrix of the layer is The bias vector is Activation value Then the gradient of each layer parameter satisfies the following formula.

[0114] The process of updating the parameters of each layer using gradient descent satisfies the following formula.

[0115] In the formula, The learning rate controls the step size for each parameter update. The larger the value, the greater the single-step adjustment range, and the faster the convergence may be, but it is prone to oscillation. The smaller the value, the smoother the adjustment and the more stable the training, but the slower the convergence. After multiple iterations, the first value network converges its prediction value for each state, i.e.: Thus, a second value network is obtained.

[0116] Continuing with an example of an application scenario in S101, let's take... ,correspond .Will Input the first value network to get the current estimated value. Calculate the mean square error loss .right The gradient is The gradient is backpropagated from the output layer to the weight matrix of the shared feature extraction layer and the value head using the chain rule. and bias vector learning rate Update parameters; after the update, the first value network (i.e., the second value network) affects... The forecast will be directed towards The directions are converging.

[0117] Step 4.3: Update the first policy network to obtain the second policy network.

[0118] Learn from small batches of experience in stages, and extract the state from each batch of experience. ,action Advantage value And the log probability of the action recorded during sampling. As The state Input the current first policy network to obtain the probability distribution of each action under the current policy, and then select the action. Corresponding log probability Calculate the probability ratio between the old and new strategies. The above probability ratio The calculation formula is as follows.

[0119] In the formula, This indicates that the current strategy is more inclined to choose this action than the old strategy. This indicates a weakening of the tendency. The PPO trimming loss is then calculated. The aforementioned PPO trimming loss The calculation formula is as follows.

[0120] In the formula, This is the cutting factor. Will Limited to Within the range, The cutting factor is 0.1 to 0.3.

[0121] right Find the gradient when When the value exceeds the clipping range and is truncated, the gradient is zero, and this rule is not updated; when... When within the cutting range, the gradient is calculated using the following formula.

[0122] The gradient is backpropagated from the output to the parameters of each layer of the first policy network using the chain rule. Let the first policy network have a total of... Layer, number The weight matrix of the layer is The bias vector is Then the gradient of each layer parameter satisfies the following formula.

[0123] After obtaining the gradients of the parameters in each layer, the gradient descent method is used to update the parameters of each layer. The gradient descent method described above satisfies the following formula.

[0124] After multiple iterations, the action probability distribution of the first policy network for each state gradually converges to the optimal policy, i.e. Thus, the second policy network is obtained. In the formula, In the state The optimal action probability distribution that maximizes long-term cumulative reward is given below.

[0125] After training, the parameters of the second policy network are saved as the maintenance policy decision network deployed to the real concrete dam. The first value network mentioned above only assists in policy updates during the training phase and is no longer used after deployment.

[0126] Continuing with an example from the application scenario in S101, let's take a piece of experience: actions under the old strategy log probability (Corresponding probability approximately 0.657), the logarithmic probability of the current first policy network outputting the same action. (Corresponding probability approximately 0.684). Calculate the probability ratio. This experience advantage value (Positive advantage, the action was better than expected). Crop limit lower limit . Within the cutting range, PPO cutting loss .right The gradient is The gradient is backpropagated to the weight matrix and bias vector of the feature extraction layer and the policy head using the chain rule, with a learning rate... Update parameters to increase the action in the state. The selection probability is determined by the following steps. After completing 10 rounds of learning with 4096 experience points, the cache is cleared, and the network re-enters the sampling phase with the updated first policy, iterating in a loop. The total training steps are 2 million. During training, the average total cost over the most recent 100 rounds is monitored, and the process terminates early when this value no longer decreases for 15 consecutive training cycles. This embodiment converges at approximately 1.5 million steps, with an average total cost of approximately 5.8 million per round.

[0127] During training, the aforementioned maintenance strategy decision-making network autonomously learned to recognize the unique seasonal patterns of the Three Gorges Reservoir area: during the flood season, rising water levels accelerate crack expansion, but when leakage increases simultaneously, it needs to determine whether the increase is driven by water pressure or structural damage; during the dry season, as water levels recede and cracks partially close, vigilance is needed if leakage continues to increase; during periods of rapid water level drop, the dam body experiences drastic stress changes, and cracks and leakage may exhibit short-term abnormal fluctuations, during which the agent learns to remain alert and not rush into repairs. After training, the second strategy network was deployed as the maintenance strategy decision-making network on an edge computing server in the Three Gorges Dam Management Center's computer room, connected to the dam safety monitoring automation system via a local area network, and receiving real-time data streams from crack gauges, piezometers, water level gauges, and rain gauges. The system triggers a decision every 30 days.

[0128] During the operation of the aforementioned maintenance strategy decision-making network, taking a certain decision cycle as an example: the current date is August 15, 2024, and the system extracts the monitoring data for that day from the real-time data stream. S =[1.68,0.52,162.5,18.2]. Z-score standardization is performed based on the mean and standard deviation of each dimension saved during the training phase. The standardized state vector is then obtained. .Will The system inputs a loaded maintenance strategy decision network. After one forward propagation, it outputs three probabilities for actions: no maintenance (0.71), minor repair (0.24), and major repair (0.05). The decision with the highest probability, "no maintenance," is selected as the current decision instruction. The system pushes the decision result to the monitoring terminal screen, displaying "Recommendation: No maintenance, continue observation (71% confidence level)," and simultaneously writes to the maintenance decision log, recording the time, original state vector, normalized parameters, probabilities of each action, and the final decision. Maintenance personnel continue daily monitoring based on the recommendation, without scheduling any maintenance work. After 30 days, the system automatically collects data for the next cycle and enters the next round of decision-making. When a decision outputs a minor repair, maintenance personnel execute the repair according to the minor repair engineering plan, which is expected to take 15 days. If it is a major repair, the major repair process is initiated, which is expected to take 65 days. During the repair period, the system does not make new decisions. After the repair is completed, the new state after the repair is completed serves as the starting point, and a new 30-day decision cycle begins.

[0129] In addition, combined Figure 1 ,like Figure 6 As shown, the above method also includes S103.

[0130] S103. Regularly update and maintain the strategy decision-making network.

[0131] Specifically, such as Figure 7 As shown, after the above-mentioned maintenance strategy decision network has run for a preset period, the newly accumulated monitoring and maintenance data will be merged with the original historical dataset. The natural degradation prediction model and maintenance effect model will be retrained using the expanded dataset. The parameters of the currently deployed strategy network will be used as the initial values, and incremental training will be performed using a learning rate lower than that used during the initial training. After completion, the online maintenance strategy decision network will be replaced by offline verification.

[0132] Continuing with an example from the S101 application scenario, the system continuously records complete information for each decision during operation. After one year of operation, approximately 12 new decision records, along with corresponding monitoring data and actual maintenance results, are accumulated. This new data is merged with the existing historical dataset, and the state prediction model is retrained using the expanded dataset. Using the currently deployed policy network parameters as initial values, the learning rate is reduced to... The cache capacity was maintained at 4096 entries, and 50,000 steps were fine-tuned. After fine-tuning, a new version of the policy network parameter file was generated, and offline verification was performed on a mixed test set of old and new versions. Before fine-tuning, the average total cost per round was approximately 5.8 million yuan, which was reduced to approximately 5.52 million yuan after fine-tuning, a decrease of approximately 4.8%. After verifying that the performance was not inferior to the current online model, a hot replacement was performed, and the new version of the policy network was put into operation as the maintenance policy decision network.

[0133] In summary, the reinforcement learning-based autonomous maintenance decision-making method for concrete dams provided in this application first obtains the state vector of the target concrete dam indicating cracks and leakage. Then, a reinforcement learning-based maintenance strategy decision-making network is used to analyze the state vector to derive the maintenance strategy for the target concrete dam. In this process, the maintenance strategy decision-making network is trained using multiple training vectors and corresponding maintenance strategies. Since the training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating maintenance strategies, the network's decision analysis of the state vectors can identify these signals, improving the accuracy of the maintenance strategy and enabling more timely maintenance, thus contributing to the long-term stable operation of the concrete dam.

[0134] Accordingly, embodiments of this application provide an autonomous maintenance decision-making system for concrete dams based on reinforcement learning, such as... Figure 8 As shown, it includes a state vector acquisition module 501 and a maintenance strategy determination module 502.

[0135] The state vector acquisition module 501 is used to acquire the state vector of the target concrete dam; the state vector indicates the cracks and leakage conditions of the concrete dam. For example, the state vector acquisition module 501 is used to implement S101 of the above method.

[0136] The maintenance strategy determination module 502 is used to determine the maintenance strategy of the target concrete dam by using the state vector of the target concrete dam and a maintenance strategy decision network based on reinforcement learning. The maintenance strategy decision network is trained using multiple training vectors and corresponding maintenance strategies. The training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating the maintenance strategy. The self-healing and pseudo-deterioration signals include crack self-healing signals and rainfall leakage signals. The maintenance strategy is to maintain the dam in its natural state or to perform repairs. For example, the maintenance strategy determination module 502 is used to implement step S102 of the above method.

[0137] The modules of the reinforcement learning-based autonomous maintenance decision-making system for concrete dams described above can also be used to execute other steps in the above method embodiments. All relevant content involved in the above method embodiments can be referred to in the functional description of the corresponding functional module, and will not be repeated here.

[0138] This application also provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the methods in the above embodiments. The processor can implement the state vector acquisition module 501 and the maintenance strategy determination module 502; the memory can also be used to store the state vector of the target concrete dam, the maintenance strategy decision network, and the maintenance strategy of the target concrete dam, etc.

[0139] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, performs the methods described in the above embodiments.

[0140] This application also provides a computer program product, which includes computer program instructions that, when run on a computer, execute the methods described above.

[0141] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for autonomous maintenance decision-making of concrete dams based on reinforcement learning, characterized in that, include: Obtain the state vector of the target concrete dam; The state vector indicates the cracks and leakage conditions of the concrete dam; The maintenance strategy for the target concrete dam is determined using the state vector of the target concrete dam and a maintenance strategy decision network based on reinforcement learning. The maintenance strategy decision network is trained using multiple training vectors and corresponding maintenance strategies. The training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating the maintenance strategy. The self-healing and pseudo-deterioration signals include crack self-healing signals and rainfall leakage signals. The maintenance strategy is to maintain the natural state or to repair it.

2. The method as described in claim 1, characterized in that, The training vectors are obtained through simulation using a state prediction model; the state prediction model includes a natural state prediction sub-model and a decay coefficient. The natural state prediction sub-model is used to identify self-healing and pseudo-deterioration signals of concrete dams and to simulate the state vector of concrete dams in the current time period, so as to obtain the state vector of concrete dams in the next time period when the maintenance strategy is to maintain the natural state. The attenuation coefficient represents the difference between the state vector of the repaired concrete dam and the state vector of the concrete dam in its natural state; the product of the attenuation coefficient and the state vector of the concrete dam in its natural state in the previous time period is the state vector of the concrete dam in the next time period when the maintenance strategy is repair.

3. The method as described in claim 2, characterized in that, The natural state prediction sub-model is trained using physical constraint terms constructed from the self-healing and pseudo-deterioration signals.

4. The method as described in claim 3, characterized in that, The physical constraint terms satisfy: in, Represents physical constraint terms. This indicates that the physical constraints are determined by the self-healing confidence level. , This indicates the confidence level of the crack self-healing term. Used to measure the degree to which crack growth slows down satisfy: This indicates the confidence level of the rainfall infiltration term. Used to measure the degree of continuous decrease in leakage. satisfy: Indicates the slowdown threshold. Indicates the steepness coefficient. Indicates time t Increment of crack width over time; This indicates the number of consecutive days the leakage rate has decreased. Indicates the threshold for the number of days of decline. This indicates the cumulative net decrease in leakage within the preset window period. Indicates the cumulative decline threshold; Indicates time t Incremental leakage flow rate This indicates the upper limit of allowed fluctuations.

5. The method as described in claim 3, characterized in that, The natural state prediction sub-model is a supervised learning regression model.

6. The method as described in claim 1, characterized in that, The step of determining the maintenance strategy for the target concrete dam using the state vector of the target concrete dam and a reinforcement learning-based maintenance strategy decision network includes: The state vector of the target concrete dam is input into the maintenance strategy decision network, which analyzes the state vector and outputs the maintenance strategy for the target concrete dam.

7. The method as described in claim 1 or 6, characterized in that, The training process of the maintenance strategy decision network includes: A candidate network is constructed, comprising a first policy network and a first value network, both of which are connected to the feature extraction layer at their inputs; the first policy network outputs the maintenance policy; and the first value network outputs the maintenance cost corresponding to the maintenance policy. An experience set is constructed using multiple training vectors and corresponding maintenance strategies. The experience set includes multiple experiences, each of which includes the training vector, maintenance strategy, and maintenance cost of the concrete dam within an execution cycle. The execution cycle is the time required to execute the maintenance strategy. With the goal of minimizing the maintenance cost output by the first value network, the candidate network is trained and updated using gradient descent and the experience set to obtain the trained candidate network; the trained candidate network includes a second policy network and a second value network. The second strategy network is determined as the maintenance strategy decision network.

8. The method as described in claim 7, characterized in that, The feature extraction layer includes a fully connected layer; The first policy network includes a first hidden layer and a Softmax output layer connected in sequence; The first value network includes a second hidden layer and a linear output layer connected in sequence.

9. The method as described in claim 1 or 7, characterized in that, The repair is either a minor repair or a major repair, and the major repair and the minor repair are classified by a clustering algorithm.

10. A reinforcement learning-based autonomous maintenance decision-making system for concrete dams, characterized in that, It includes a state vector acquisition module and a maintenance strategy determination module; The state vector acquisition module is used to acquire the state vector of the target concrete dam; the state vector indicates the cracks and leakage of the concrete dam. The maintenance strategy determination module is used to determine the maintenance strategy of the target concrete dam by using the state vector of the target concrete dam and a maintenance strategy decision network based on reinforcement learning; wherein, the maintenance strategy decision network is obtained by training multiple training vectors with corresponding maintenance strategies; the training vectors are the state vectors of the concrete dam after identifying self-healing and pseudo-deterioration signals and simulating maintenance strategies; the self-healing and pseudo-deterioration signals include crack self-healing signals and rainfall leakage signals; the maintenance strategy is to maintain the natural state or to repair it.