Bridge group multi-target maintenance decision-making method fusing evolutionary algorithm and artificial intelligence

By building a collaborative iteration mechanism between weight combination optimization and maintenance strategy learning, combined with multi-objective evolution algorithm and A2C reinforcement learning, the problem of dynamic trade-offs and low computing efficiency in multi-objective maintenance decisions of bridge group is solved, and adaptive adjustment and efficient optimization of bridge group maintenance strategies are achieved.

CN120525522APending Publication Date: 2025-08-22UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511020141.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing deep reinforcement learning methods and multi-objective evolution algorithms have problems such as inability to dynamically weigh the maintenance costs and structural failure risks, large calculation overhead, and low search efficiency in the multi-objective maintenance decision of bridge groups, which is difficult to meet the needs of multi-objective optimization during the entire life cycle of bridge groups.

Method used

Build an iterative mechanism that coordinates weight combination optimization and maintenance strategy learning, generates initial populations through multi-objective evolution algorithms, combines A2C reinforcement learning agents to learn the optimal maintenance strategy in the Markov decision-making environment, and optimizes weight combination through a closed-loop feedback mechanism to achieve collaborative optimization of multi-objective maintenance decisions.

Benefits of technology

The adaptive adjustment of bridge group maintenance strategies has been achieved, the intelligence level and practical value of the decision-making system have been improved, and multiple goals can be weighed dynamically under different management preferences, which has improved the scientificity and adaptability of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525522A_ABST
    Figure CN120525522A_ABST
Patent Text Reader

Abstract

The invention provides a bridge group multi-target maintenance decision-making method fusing an evolutionary algorithm and artificial intelligence, and relates to the technical field of civil engineering and artificial intelligence crossing. The method comprises the steps of defining bridge group maintenance cost and structure failure risks, representing preferences of decision makers for different decision targets by weight combinations, and constructing a multi-target maintenance decision optimization model; encoding the weight combination into an individual of a multi-objective evolutionary algorithm, and randomly generating an initial population; aiming at each generation of weight combination, constructing a bridge group Markov decision-making environment; learning an optimal maintenance strategy by adopting an A2C training reinforcement learning agent; the optimal maintenance strategy is evaluated, and an evaluation result is used as individual fitness to be fed back to the multi-objective evolutionary algorithm; using a multi-objective evolutionary algorithm to perform evolutionary search on the multi-objective weight combination; and through a closed-loop feedback mechanism, outputting a Pareto optimal solution set containing an optimal maintenance strategy under various weight combinations, thereby realizing collaborative optimization of weight optimization and strategy learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the interdisciplinary field of civil engineering and artificial intelligence, and in particular to a multi-objective maintenance decision-making method for bridge groups that integrates evolutionary algorithms and artificial intelligence. Background Art

[0002] As an important component of the urban road network, the safety of bridge groups is largely related to the operational efficiency and public safety of urban transportation systems. With the increase in service life and traffic loads, the safety challenges faced by bridges are becoming increasingly severe. Faced with the reality of increasing bridge aging, frequent structural defects, and tight maintenance resources, the traditional "passive repair" model can no longer meet the comprehensive requirements of modern cities for the economy and safety of transportation infrastructure. Therefore, the maintenance model of bridge groups is gradually shifting to a preventive maintenance model based on full life cycle management requirements and collaborative optimization of multiple objectives. This model aims to scientifically control maintenance costs and improve resource allocation efficiency while ensuring structural safety, thereby achieving the collaborative optimization of "reducing full life cycle maintenance costs" and "reducing the risk of structural failure."

[0003] In early research, maintenance strategies for individual bridges were typically developed based on regular inspection data and expert experience, or through traditional optimization methods such as linear programming and dynamic programming. While these approaches have a solid theoretical foundation and clear decision-making logic, they lack specificity and timeliness when faced with the maintenance decision-making challenges of large-scale urban bridge clusters with significant differences in operating conditions and uncertain future evolution, and are unable to produce preventive maintenance strategies with practical application value.

[0004] With the rapid development of artificial intelligence technology, deep reinforcement learning (DRL) has gradually become an important tool in the field of urban bridge maintenance decision-making due to its advantages of active learning and data-driven. This type of method models the bridge group maintenance decision-making problem as a Markov decision process (MDP). Through the continuous interaction between the reinforcement learning agent and the Markov decision environment, it learns the optimal maintenance strategy that can maximize long-term returns in a high-dimensional state space (Lei X, Xia Y, Deng L, et al. A deep reinforcement learning framework for life-cycle maintenance planning of regional deteriorating bridges using inspection data[J]. Structural and Multidisciplinary Optimization, 2022, 65(5): 149.). In this method framework, the state space is usually composed of multi-dimensional feature vectors such as the structural status indicators and service years of all bridges, while the action space covers typical maintenance plans such as "monitoring, minor repairs, major repairs, and replacements." Existing studies have applied deep reinforcement learning methods to the maintenance decision-making problems of single bridges and bridge networks, showing good generalization ability and engineering adaptability. It is particularly suitable for the full life cycle maintenance decision-making problem of bridge groups.

[0005] At the same time, maintenance decisions for bridge groups need to take into account multiple conflicting objectives, such as the maintenance cost of the entire life cycle and the risk of structural failure. Multi-objective evolutionary algorithms (MOEAs), represented by the non-dominated sorting genetic algorithm II (NSGA-II), have the ability to search for Pareto optimal solutions to multi-objective optimization problems in parallel in the objective space. In typical applications, this type of algorithm encodes annual maintenance plans, budget allocations, and maintenance measures into individual genes, and uses crossover and mutation operations to iteratively evolve to generate diverse maintenance strategies (Shen Z, Liu Y, Liu J, et al. A Decision-Making Method for Bridge Network Maintenance Based on Disease Transmission and NSGA-II[J]. Sustainability, 2023, 15(6): 5007.). Existing studies have applied multi-objective evolutionary algorithms to maintenance decisions for single bridges and bridge networks, demonstrating good generalization and engineering adaptability, and are particularly suitable for multi-objective maintenance decision-making problems for bridge groups.

[0006] Although deep reinforcement learning methods have demonstrated significant advantages in learning maintenance strategies, and multi-objective evolutionary algorithms have also demonstrated excellent performance in parallel search for multi-objective solutions, the two have yet to achieve deep integration and synergy. In existing research, deep reinforcement learning methods, when dealing with multi-objective optimization problems, typically weight multiple objectives such as "reducing lifecycle maintenance costs" and "reducing the risk of structural failure" into a single scalar reward function based on fixed weights, implicitly assuming the existence of a global optimal strategy. This approach ignores the differentiated demands for objective weights under different management preferences, lacks the ability to adapt to multiple management preferences and weight combinations, and is unable to effectively support the strategy in achieving dynamic trade-offs between objectives such as "reducing lifecycle maintenance costs" and "reducing the risk of structural failure," thus limiting its applicability in multi-objective maintenance strategy formulation. Meanwhile, while multi-objective evolutionary algorithms, such as NSGA-II, can generate Pareto-optimal solutions covering a variety of objective preferences, they are essentially offline static optimization methods that do not explicitly model the dynamic evolution of the service state of a bridge group. They lack a mechanism for learning strategies through interaction with the environment, making them difficult to address the uncertainty and dynamic changes inherent in maintenance decision-making. Furthermore, when modeling bridge group maintenance decision-making in a high-dimensional state space, the length of individual chromosome codes increases linearly with the number of bridges and the planned lifespan. This leads to a dramatic expansion of the search space dimension and a significant increase in computational overhead. This, in turn, leads to problems such as low search efficiency and convergence difficulties, severely limiting their practicality and responsiveness in lifecycle maintenance management.

[0007] Based on the above background, there is an urgent need to propose a hybrid framework for multi-objective maintenance decision-making of bridge groups that can integrate multi-objective search capabilities and long-term strategy learning capabilities, so as to achieve the joint optimization of maintenance strategies and objective weights under complex management preferences, so as to break through the bottlenecks of existing methods in multi-objective adaptability, strategy generalization ability and actual engineering usability, and promote the intelligent development of maintenance decision-making for urban bridge groups.

[0008] In summary, existing deep reinforcement learning methods and multi-objective evolutionary algorithms have the following main problems:

[0009] First, in the multi-objective maintenance decision-making problem for bridge groups, existing deep reinforcement learning-based methods typically employ a fixed-weight approach, simplifying multiple decision objectives into a single scalar reward function, thereby transforming the approach into a single-objective optimization paradigm and learning a globally optimal maintenance policy through interaction with the environment. This approach suffers from the following key limitations: it implicitly assumes the existence of a single globally optimal maintenance policy, lacks adaptability to different management preferences and weight combinations, and struggles to achieve a dynamic trade-off between maintenance costs and structural failure risks. Consequently, the decision-making system lacks the ability to adaptively adjust decision objectives and is unable to respond to the differentiated cost-risk regulation requirements of different maintenance phases.

[0010] Secondly, multi-objective evolutionary algorithms, as typical offline static optimization methods, lack a mechanism for learning strategies through interaction with the environment and are unable to account for the uncertainty inherent in bridge state evolution. Furthermore, when modeling bridge group maintenance decision-making in a high-dimensional state space, the length of individual chromosome codes increases linearly with the number of bridges and the planned lifespan. This leads to a dramatic expansion of the search space dimension and a significant increase in computational overhead. This, in turn, leads to low search efficiency and convergence difficulties, severely limiting their practicality and responsiveness in lifecycle maintenance management.

[0011] Multi-objective maintenance decision-making for bridge groups requires a comprehensive trade-off between multiple conflicting objectives, such as "reducing maintenance costs throughout the life cycle" and "reducing the risk of structural failure." Existing deep reinforcement learning methods typically only design reward functions for a single objective, making it difficult to fully characterize the dynamic balance between multiple objectives, such as maintenance costs and structural failure risks. Furthermore, such methods are based on the assumption that there is a single global optimal maintenance strategy, which does not hold true in the context of multi-objective optimization and cannot meet the actual needs of multi-objective maintenance decision-making for bridge groups. While existing multi-objective evolutionary algorithms can generate maintenance strategies under different objective preferences, they lack the ability to learn maintenance strategies for bridge groups throughout their life cycle in high-dimensional state spaces and are unable to generate detailed maintenance plans for specific weights.

[0012] Therefore, the current deep reinforcement learning and multi-objective evolutionary algorithm methods have not yet formed an effective fusion mechanism, and cannot simultaneously achieve "maintenance strategy learning ability based on environmental feedback" and "multi-objective strategy diversity expression ability" in the same framework. It is difficult to meet the comprehensive decision-making needs of dynamic balance and collaborative optimization of multiple objectives (such as maintenance costs and structural failure risks) of bridge groups throughout their life cycle. Summary of the Invention

[0013] To address the technical issues in existing technologies that make it difficult to meet the comprehensive decision-making requirements for dynamic trade-offs and collaborative optimization of multiple objectives (such as maintenance costs and structural failure risks) for bridge groups throughout their lifecycles, the present invention provides a multi-objective maintenance decision-making method for bridge groups that integrates evolutionary algorithms and artificial intelligence. By constructing an iterative mechanism that coordinates weighted combination optimization and maintenance strategy learning, it dynamically optimizes maintenance strategies and weight coefficients, solves the problem of collaborative optimization of multi-objective maintenance decisions for bridge groups, and enhances the adaptability of decision-making solutions to different objective preferences. The technical solution is as follows:

[0014] On the one hand, a multi-objective maintenance decision-making method for a bridge group is provided that integrates an evolutionary algorithm and artificial intelligence. The method is implemented by a multi-objective maintenance decision-making device for a bridge group, and the method includes:

[0015] S1, define the maintenance cost and structural failure risk function, use weight combination to represent the decision maker's preference for different decision objectives, and build a multi-objective maintenance decision optimization model;

[0016] S2, encoding the weight combination into individuals of a multi-objective evolutionary algorithm and randomly generating an initial population;

[0017] S3, for each generation of weight combinations, a Markov decision environment for the bridge group is constructed based on each weight combination; the reward function in the Markov decision environment is the weighted sum of maintenance cost and structural failure risk;

[0018] S4, in the Markov decision environment, using A2C to train a reinforcement learning agent to learn an optimal maintenance policy for a group of bridges; wherein A2C represents the dominant actor-critic algorithm;

[0019] S5, evaluate the obtained optimal maintenance strategy, and feed the evaluation result as individual fitness back to the multi-objective evolutionary algorithm in step S6 to guide the next round of weight combination evolution;

[0020] S6, using a multi-objective evolutionary algorithm to perform evolutionary search on the generated population to obtain a new generation of weight combinations;

[0021] S7, repeat steps S3 to S6 until the specified number of training steps is reached, and output a Pareto optimal solution set containing the optimal maintenance strategies under multiple weight combinations as a supporting basis for multi-objective maintenance decision-making of bridge groups.

[0022] Furthermore, the maintenance cost and structural failure risk functions are defined, and weight combinations are used to represent the decision maker's preferences for different decision objectives. The multi-objective maintenance decision optimization model is constructed, including:

[0023] Model the cost of repairing a group of bridges, where the repair cost is expressed as:

[0024]

[0025] in, represents the maintenance cost, represents the total construction cost associated with bridge type c, represents the cost coefficient associated with the bridge state s, represents the cost coefficient associated with maintenance action a;

[0026] The structural failure risk is measured by the total cost of bridge damage and the risk coefficient; the structural failure risk is expressed as:

[0027]

[0028] in, represents the structural failure risk, represents the safety risk factor associated with the bridge status s;

[0029] Maintenance cost and structural failure risk are treated as two independent objective functions, and dynamic weights are introduced to express the optimization tendency under different management preferences. A multi-objective maintenance decision optimization model is constructed. The multi-objective maintenance decision optimization model is expressed as:

[0030]

[0031]

[0032] in, represents the maintenance cost weight; represents the structural failure risk weight; and are dynamic weights, which are decision variables for multi-objective optimization and are used to reflect the decision maker’s preference for different objectives. The constraints are: .

[0033] Furthermore, encoding the weight combination into individuals of a multi-objective evolutionary algorithm and randomly generating an initial population includes:

[0034] According to the “Maintenance Cost Weight – Structural failure risk weight " in the form of a two-dimensional vector, and The chromosomes of individuals encoded as multi-objective evolutionary algorithms, where + = 1, and 、 All values ​​are in the closed interval from 0 to 1;

[0035] In satisfaction + = 1, random sampling is used to generate N weight vectors as the initial population, where N is the preset population size.

[0036] Furthermore, for each generation of weight combinations, constructing a Markov decision environment for a bridge group according to each weight combination includes:

[0037] A Markov decision environment for a group of bridges was constructed. The state transition probability matrix in the Markov decision environment was set based on annual bridge inspection data, satisfying the assumption that bridges can degrade at most one level in a year without maintenance, and that different maintenance actions correspond to different repair outcomes.

[0038] Calculate the maintenance cost based on the state of the bridge group at time t and the set of maintenance actions selected by the reinforcement learning agent for the bridge group and structural failure risk ;

[0039] Using the current weight combination ( , ), according to the formula Calculating the reward function , embed the reward function into the constructed Markov decision environment for the next step of reinforcement learning agent training.

[0040] Furthermore, the obtained optimal maintenance strategy is evaluated, and the evaluation result is fed back to the multi-objective evolutionary algorithm in step S6 as the individual fitness to guide the next round of weighted combination evolution, including:

[0041] After completing the reinforcement learning training corresponding to each weight combination, the optimal maintenance strategy obtained by executing the reinforcement learning agent in the Markov decision environment is used to accumulate the maintenance cost of each time step within a preset simulation cycle to obtain a cumulative maintenance cost;

[0042] Synchronously calculate the structural failure risk corresponding to each time step within the same preset simulation cycle to obtain the cumulative structural failure risk;

[0043] The accumulated maintenance cost and the accumulated structural failure risk are combined into a fitness vector, and the fitness vector is fed back to the multi-objective evolutionary algorithm in step S6 to guide a new round of weight update.

[0044] Furthermore, the multi-objective evolutionary algorithm is NSGA-II; wherein NSGA-II is a non-dominated sorting genetic algorithm-II;

[0045] The multi-objective evolutionary algorithm is used to perform evolutionary search on the generated population to obtain a new generation of weight combinations including:

[0046] NSGA-II is applied to the generated population, and fast non-dominated sorting, crowding distance calculation, elite retention, crossover and mutation operations are performed in sequence to obtain a new generation of weight combinations.

[0047] Furthermore, the generated population is subjected to NSGA-II, which sequentially performs fast non-dominated sorting, crowding distance calculation, elite retention, crossover, and mutation operations to obtain a new generation of weight combinations including:

[0048] Perform non-dominated sorting on the generated initial population and divide individuals into different non-dominated levels according to the Pareto dominance relationship;

[0049] The crowding distance is calculated within each non-dominated layer to measure the individual distribution density and maintain population diversity;

[0050] A binary tournament selection mechanism based on non-dominated rank and crowding degree is adopted, which gives priority to individuals with better non-dominated rank and selects individuals with larger crowding distance when the ranks are the same to form the parent population;

[0051] Perform crossover operation on parent individuals to generate offspring individuals;

[0052] Performing a mutation operation on the offspring individuals to expand the search space;

[0053] The parent population is merged with the child population, and fast non-dominated sorting and crowding distance calculation are performed. Based on the non-dominated level and crowding distance priority of the individuals, the top N individuals are retained to form a new generation of weighted combinations; where N is the preset population size.

[0054] On the other hand, a multi-objective maintenance decision-making device for a bridge group is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the multi-objective maintenance decision-making methods for a bridge group that integrates the evolutionary algorithm and artificial intelligence is implemented.

[0055] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned multi-objective maintenance decision-making methods for bridge groups that integrate evolutionary algorithms and artificial intelligence.

[0056] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0057] In the above technical solution, a multi-objective maintenance decision optimization model based on the maintenance cost and structural failure risk of bridge groups is constructed to accurately express the conflicts and synergies between multiple objectives in the maintenance management of bridge groups. Then, based on the multi-objective evolutionary algorithm, an evolutionary search is performed on the multi-objective weight combination to obtain the optimization direction covering different preference trade-offs. Furthermore, a reward function-based Markov decision environment for bridge groups is constructed based on each set of weight combinations. In the Markov decision environment, A2C training is used to train the reinforcement learning agent to learn the optimal maintenance strategy. Finally, through a closed-loop feedback mechanism, the evaluation result obtained based on the optimal maintenance strategy is fed back to the multi-objective evolutionary algorithm as the fitness to guide a new round of weight updates, thereby realizing the coordinated iteration of weight optimization and strategy learning, so that the bridge group maintenance strategy has the ability to adaptively adjust, thereby improving the intelligence level and practical value of the decision-making system. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0059] Figure 1 This is a flow chart of a multi-objective maintenance decision-making method for bridge groups that integrates evolutionary algorithms and artificial intelligence, provided by an embodiment of the present invention;

[0060] Figure 2 1 is a flow chart of a multi-objective evolutionary algorithm (taking NSGA-II as an example) provided in an embodiment of the present invention;

[0061] Figure 3 This is a flow chart of solving weight combinations based on a multi-objective evolutionary algorithm (taking NSGA-II as an example) provided by an embodiment of the present invention;

[0062] Figure 4 It is a detailed flowchart of a multi-objective maintenance decision-making method for a bridge group that integrates a multi-objective evolutionary algorithm and artificial intelligence, provided by an embodiment of the present invention;

[0063] Figure 5 Schematic diagram of a multi-objective hybrid decision-making framework provided by an embodiment of the present invention;

[0064] Figure 6 It is a structural diagram of a multi-objective maintenance decision-making device for a bridge group provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0066] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0067] In the embodiments of the present invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same. The terms "of," "corresponding," and "corresponding" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same.

[0068] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0069] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0070] The embodiment of the present invention provides a multi-objective maintenance decision-making method for a bridge group that integrates evolutionary algorithms and artificial intelligence. The method can be implemented by a multi-objective maintenance decision-making device for a bridge group, which can be a terminal or a server. Figure 1 The flowchart of the multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence is shown. The processing flow of the method may include the following steps:

[0071] S1, define the maintenance cost and structural failure risk function, use weight combination to represent the decision maker's preference for different decision objectives, and build a multi-objective maintenance decision optimization model; specifically, the following steps may be included:

[0072] S11: Model the maintenance costs of bridge groups. The cost calculation takes into account the construction cost baseline of the bridge, the influence coefficient of the bridge condition, and the cost differences caused by different maintenance actions. Based on the bridge type, bridge condition, and possible maintenance measures, the maintenance cost is defined as:

[0073]

[0074] in, represents the maintenance cost, represents the total construction cost associated with bridge type c, represents the cost coefficient associated with the bridge state s, represents the cost coefficient associated with maintenance action a;

[0075] S12 measures the risk of structural failure by using the total cost of bridge damage and the risk factor, converting the risk of structural failure into quantifiable economic losses. This process uses the bridge condition level as the core variable and combines it with the risk factor for evaluation. Ultimately, the structural failure risk is expressed in the form of economic cost and is defined as follows:

[0076]

[0077] in, represents the risk of structural failure, represents the safety risk factor associated with the bridge status s;

[0078] In S13, maintenance cost and structural failure risk are treated as two independent objective functions, and dynamic weights are introduced to express the optimization tendency under different management preferences. A multi-objective maintenance decision optimization model is constructed to provide a quantitative basis for subsequent weight combination search and strategy learning. The multi-objective maintenance decision optimization model is expressed as:

[0079]

[0080]

[0081] in, represents the maintenance cost weight; represents the structural failure risk weight; and are dynamic weights, which are decision variables for multi-objective optimization and are used to reflect the decision maker’s preference for different objectives. The constraints are: .

[0082] S2, encoding the weight combination into individuals of a multi-objective evolutionary algorithm, and randomly generating an initial population; specifically, the following steps may be included:

[0083] S21, according to the “Maintenance cost weight – Structural failure risk weight " in the form of a two-dimensional vector, and The chromosomes of individuals encoded as multi-objective evolutionary algorithms, where + = 1, and 、 All values ​​are in the closed interval from 0 to 1;

[0084] S22, in meeting + = 1, random sampling is used to generate N weight vectors as the initial population, where N is the preset population size.

[0085] In this embodiment, through steps S21-S22, the weight combination is formally encoded as an individual of the multi-objective evolutionary algorithm, and a certain number of weight combinations are randomly initialized to form the initial evolution population, such as Figure 2 and Figure 3 shown.

[0086] S3, such as Figure 4 and Figure 5 As shown in FIG, for each generation of weight combinations, a Markov decision environment for a bridge group is constructed based on each weight combination; wherein the reward function in the Markov decision environment is the weighted sum of the maintenance cost and the structural failure risk; specifically, the following steps may be included:

[0087] S31, constructing a Markov decision environment for the bridge group; wherein the state transition probability matrix in the Markov decision environment is set based on the annual inspection data of the bridges, satisfying the assumption that the bridges will degrade at most one level within a year without maintenance, and that different maintenance actions correspond to different repair effects;

[0088] S32, calculate the maintenance cost based on the state of the bridge group at time t and the set of maintenance actions selected by the reinforcement learning agent for the bridge group and structural failure risk ;

[0089] S33, using the current weight combination ( , ), according to the formula Calculating the reward function , embed the reward function into the constructed Markov decision environment for the next step of training the reinforcement learning agent.

[0090] In this embodiment, in the strategy learning phase, each weight combination is used as the current optimization preference to construct a reward function in a Markov decision environment.

[0091] S4, such as Figure 4 and Figure 5 As shown, in the Markov decision environment, A2C is used to train a reinforcement learning agent (specifically, a reinforcement learning maintenance strategy agent) to learn the optimal maintenance strategy for a group of bridges. A2C represents the dominant actor-critic algorithm. Specifically, the following steps may be included:

[0092] S41, when using A2C to train a reinforcement learning agent, constructs a dual-branch structure consisting of a policy network (Actor) and a value network (Critic). These networks share two fully connected hidden layers, each containing 64 neurons and using the Tanh function as the activation function. After the shared layer, each layer is connected to its own output layer. The policy network outputs the probability distribution of each maintenance action under different states, while the value network estimates the state-value function.

[0093] S42, using the Adam optimizer to update the parameters of the policy network and the value network, setting the learning rate to 0.0007, the discount factor to 0.99, the number of multi-step reward steps to 100, the gradient norm clipping threshold to 0.5, and the value function loss weight to 0.5;

[0094] S43, network initialization: After determining the weight combination and constructing the reward function in step S33, initialize the policy network and value network including the shared two FC hidden layers (described in S41);

[0095] S44, set the time steps included in a simulation cycle to 100, representing the 100-year service life of the bridge group, and the time step is 1 year. At each time step, the maintenance action is selected according to the action probability distribution output by the current strategy network and interacts with the Markov decision environment, recording the state ,action ,award , next state ;

[0096] S45, Introducing Advantage Function in A2C Critic feedback can be used to measure the relative quality of a specific maintenance action. The calculation formula is:

[0097]

[0098] in, is the action value function, which means that in state Take action and the expectation of the cumulative reward that can be obtained by continuing to act according to the strategy; Is the state value function, which means that in state The algorithm starts with the expected cumulative reward obtained by following the policy actions. At each time step, the policy gradient is calculated based on the advantage function. The policy network and value network are jointly updated using the Adam optimizer. After completing the preset 1,000 simulation cycles (a total of 100,000 time steps), the optimal repair policy corresponding to the current weight combination is output.

[0099] In this example, based on A2C, a reinforcement learning agent is trained through multiple rounds of interaction with a Markov decision environment to learn the optimal maintenance strategy within it. The reinforcement learning state space covers all bridge status levels, and the action space encompasses executable maintenance options, such as no repair, minor repair, major repair, and replacement.

[0100] S5, evaluates the obtained optimal maintenance strategy and feeds the evaluation result as individual fitness back to the multi-objective evolutionary algorithm in step S6 to guide the next round of weighted combination evolution; specifically, the following steps are included:

[0101] S51, after completing the reinforcement learning training corresponding to each weight combination, using the reinforcement learning agent to execute the optimal maintenance strategy in the Markov decision environment, the maintenance cost of each time step within a preset simulation period is accumulated over time to obtain the cumulative maintenance cost within the service period;

[0102] S52, synchronously counting the structural failure risk corresponding to each time step in the same preset simulation cycle to obtain the cumulative structural failure risk;

[0103] In step S53, the accumulated maintenance cost and the accumulated structural failure risk are combined into a fitness vector, and the fitness vector is fed back to the multi-objective evolutionary algorithm in step S6, so that the multi-objective evolutionary algorithm re-executes operations such as fast non-dominated sorting, crowding distance calculation, elite retention, crossover and mutation based on the fitness vector, updates the weighted combination population, and enters the next round of evolution.

[0104] In this embodiment, after the reinforcement learning agent is trained, the maintenance strategy it outputs is evaluated, and the cumulative maintenance cost and cumulative structural failure risk generated under the current weight combination are calculated. This is used as the fitness indicator of the individual and fed back to the multi-objective evolutionary algorithm to guide the next round of evolution.

[0105] S6, using a multi-objective evolutionary algorithm to perform evolutionary search on the generated population to obtain a new generation of weight combinations;

[0106] In this embodiment, the multi-objective evolutionary algorithm is NSGA-II, wherein NSGA-II is a non-dominated sorting genetic algorithm-II, and NSGA-II includes: fast non-dominated sorting, crowding distance calculation, elite retention (i.e., selection), crossover and mutation operations.

[0107] In this embodiment, Figure 2 and Figure 3 As shown, NSGA-II is applied to the generated population, and fast non-dominated sorting, crowding distance calculation, elite retention, crossover and mutation operations are performed in sequence to obtain a new generation of weight combinations; specifically, the following steps may be included:

[0108] S61, the generated initial population W N Perform non-dominated sorting and classify individuals into different non-dominated levels according to the Pareto dominance relationship;

[0109] S62, the crowding distance is calculated within each non-dominated layer to measure the individual distribution density and maintain population diversity;

[0110] S63, adopts a binary tournament selection mechanism based on non-dominated rank and crowding degree, giving priority to individuals with better non-dominated rank, and selecting individuals with larger crowding distance when the ranks are the same to form the parent population;

[0111] S64, performing a crossover operation on the parent individuals to generate offspring individuals;

[0112] In this embodiment, a simulated binary crossover operator is applied to the parent individuals at a crossover rate of 0.9 to generate offspring individuals.

[0113] S65, performing a mutation operation on the offspring individuals to expand the search space;

[0114] In this embodiment, a polynomial mutation operator is applied to the offspring individuals at a mutation rate of 0.1 to expand the search space.

[0115] S66: Merge the parent population with the child population, perform fast non-dominated sorting and crowding distance calculation, and retain the top N individuals to form a new generation weighted combination based on the non-dominated level and crowding distance priority of the individuals; where N is the preset population size.

[0116] In this embodiment, in the evolutionary stage, NSGA-II is used as the outer optimization mechanism to classify individuals in the population into good and bad ones through non-dominated sorting, and the crowding distance indicator is introduced to control the distribution uniformity of the solutions. On this basis, genetic operations such as elite retention, crossover, and mutation are applied to iteratively update the population to achieve continuous optimization of the multi-objective weight combination and gradually approach the Pareto optimal solution set composed of non-dominated solutions.

[0117] In this embodiment, in order to better understand the non-dominated solution, a brief description is given:

[0118] Assume that for all objectives, solution x is better than solution y, then x is said to dominate y. If solution x is not dominated by other solutions in the solution set, then x is called a non-dominated solution.

[0119] S7, repeat steps S3 to S6 until the specified number of training steps is reached, and output a Pareto optimal solution set containing the optimal maintenance strategies under multiple weight combinations as a supporting basis for multi-objective maintenance decision-making of bridge groups.

[0120] In this embodiment, Figure 4 As shown in the figure, a multi-objective maintenance decision optimization model is constructed through S1; a multi-objective hybrid decision framework with collaborative nesting of weight optimization and strategy learning is constructed through S2-S6 to realize dynamic optimization of multi-objective maintenance strategies. Among them, S2, S5 and S6 use NSGA-II to realize weight combination evolution; S3 and S4 use A2C to learn maintenance strategies. In other words, the multi-objective hybrid decision framework is a multi-objective hybrid decision framework that integrates NSGA-II and A2C; S7 realizes model training.

[0121] In this embodiment, the two processes of weight evolution and strategy learning are repeatedly executed to form a closed-loop collaborative optimization mechanism. As the population continues to evolve and the reinforcement learning agent continues to optimize, the resulting maintenance strategy set gradually converges to a set of non-dominated solutions that can adapt to management needs under different preferences and have high performance, high stability, and high interpretability.

[0122] In summary, in the above technical solution, by constructing a multi-objective maintenance decision optimization model based on the maintenance cost and structural failure risk of bridge groups, the conflict and coordination relationship between multiple objectives in the maintenance management of bridge groups is accurately expressed; then, based on the multi-objective evolutionary algorithm, an evolutionary search is performed on the multi-objective weight combination to obtain the optimization direction covering different preference trade-offs; further, based on each set of weight combinations, a reward function-based Markov decision environment for bridge groups is constructed. In the Markov decision environment, A2C training is used to train the reinforcement learning agent to learn the optimal maintenance strategy; finally, through a closed-loop feedback mechanism, the evaluation result obtained based on the optimal maintenance strategy is fed back to the multi-objective evolutionary algorithm as the fitness to guide a new round of weight update, thereby realizing the coordinated iteration of weight optimization and strategy learning, so that the bridge group maintenance strategy has the ability of adaptive adjustment, and improves the intelligence level and practical value of the decision-making system.

[0123] The multi-objective bridge maintenance decision-making method proposed in this invention, which integrates evolutionary algorithms and artificial intelligence, aims to achieve a dynamic trade-off between multiple objectives such as economy and safety in bridge maintenance decisions, thereby improving the scientific nature, adaptability, and feasibility of the decision-making strategy. Compared with existing technologies, the embodiments of this invention have at least the following significant technical effects:

[0124] 1. Improve the comprehensive optimization capabilities of maintenance decisions

[0125] By integrating the advantages of multi-objective evolutionary algorithms and A2C, the present invention enables bridge maintenance decisions to take into account multiple optimization objectives simultaneously. By dynamically adjusting the weight combination during the continuous evolution process, it effectively approaches the Pareto frontier and achieves collaborative optimization among different objectives.

[0126] 2. Dynamically adapt to trade-offs between different objectives

[0127] The present invention establishes a closed-loop optimization mechanism in which weight optimization and strategy learning are nested with each other. Weight optimization and strategy learning operate alternately, enabling the system to continuously correct weight settings and dynamically adjust maintenance strategies during the evolution process, adapting to the trade-off requirements between different objectives. It has good flexibility and adaptability, and can continuously output optimized maintenance strategies based on the actual operating status and management preferences of the bridge group, effectively controlling maintenance costs while ensuring structural safety, and significantly improving the level of intelligent maintenance decision-making.

[0128] 3. Improve decision-making accuracy and adaptability

[0129] This invention effectively improves the accuracy of decision-making strategies in the context of multi-objective trade-offs. By using a reinforcement learning agent to deeply train maintenance plans under each weighted combination, it ensures that each strategy achieves optimal balance within the multi-objective framework, adapting to the strategy requirements of different bridges, different status levels, and different budget conditions in a complex bridge network.

[0130] 4. Improve computational efficiency and solution quality

[0131] The present invention introduces congestion calculation and elite selection strategy during the evolution process, which enables NSGA-II to maintain the diversity and representativeness of solutions when dealing with multi-objective optimization problems, while accelerating convergence, improving optimization efficiency and strategy quality, and meeting the engineering application requirements in large-scale infrastructure scenarios.

[0132] 5. Strong scalability, suitable for multi-field decision optimization

[0133] The hybrid optimization framework proposed in this paper has a clear structure, independent modules, and strong versatility and scalability. In addition to bridge maintenance decision-making, it can also be applied to other complex decision-making problems with multi-objective optimization characteristics, such as energy scheduling, traffic management, and asset operation and maintenance. It provides theoretical support and technical means for intelligent decision-making in various fields.

[0134] Figure 6 is a structural diagram of a multi-objective maintenance decision-making device for a bridge group provided by an embodiment of the present invention. Optionally, the multi-objective maintenance decision-making device 610 for a bridge group may include a first processor 2001 .

[0135] Optionally, the bridge group multi-objective maintenance decision-making device 610 may further include a memory 2002 and a transceiver 2003 .

[0136] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0137] The following combination Figure 6 The components of the bridge group multi-objective maintenance decision-making device 610 are described in detail:

[0138] The first processor 2001 is the control center of the bridge group multi-objective maintenance decision-making device 610 and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), or application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).

[0139] Optionally, the first processor 2001 may execute various functions of the bridge group multi-objective maintenance decision-making device 610 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002 .

[0140] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 are shown in FIG.

[0141] In a specific implementation, as an embodiment, the bridge group multi-objective maintenance decision-making device 610 may also include multiple processors, such as Figure 6 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0142] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0143] Alternatively, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and accessed through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0144] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0145] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0146] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and communicate with the bridge group multi-objective maintenance decision-making device 610 through an interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0147] It should be noted that Figure 6 The structure of the bridge group multi-objective maintenance decision-making device 610 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0148] In addition, the technical effects of the bridge group multi-objective maintenance decision-making device 610 can refer to the technical effects of the bridge group multi-objective maintenance decision-making method that integrates evolutionary algorithm and artificial intelligence described in the above method embodiment, and will not be repeated here.

[0149] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0150] It should also be understood that the memory in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0151] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0152] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0153] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0154] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0155] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0156] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0157] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0158] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0159] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0160] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.

[0161] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multi-objective maintenance decision-making method for bridge groups that integrates evolutionary algorithms and artificial intelligence, characterized by: The method comprises: S1, define the maintenance cost and structural failure risk function, use weight combination to represent the decision maker's preference for different decision objectives, and build a multi-objective maintenance decision optimization model; S2, encoding the weight combination into individuals of a multi-objective evolutionary algorithm and randomly generating an initial population; S3, for each generation of weight combinations, a Markov decision environment for the bridge group is constructed based on each weight combination; the reward function in the Markov decision environment is the weighted sum of maintenance cost and structural failure risk; S4, in the Markov decision environment, using A2C to train a reinforcement learning agent to learn an optimal maintenance policy for a group of bridges; wherein A2C represents the dominant actor-critic algorithm; S5, evaluate the obtained optimal maintenance strategy, and feed the evaluation result as individual fitness back to the multi-objective evolutionary algorithm in step S6 to guide the next round of weight combination evolution; S6, using a multi-objective evolutionary algorithm to perform evolutionary search on the generated population to obtain a new generation of weight combinations; S7, repeat steps S3 to S6 until the specified number of training steps is reached, and output a Pareto optimal solution set containing the optimal maintenance strategies under multiple weight combinations as a supporting basis for multi-objective maintenance decision-making of bridge groups.

2. The multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence according to claim 1 is characterized in that: The maintenance cost and structural failure risk functions are defined, and weight combinations are used to represent the decision maker's preferences for different decision objectives. The multi-objective maintenance decision optimization model is constructed, including: Model the cost of repairing a group of bridges, where the repair cost is expressed as: ; in, represents the maintenance cost, represents the total construction cost associated with bridge type c, represents the cost coefficient associated with the bridge state s, represents the cost coefficient associated with maintenance action a; The structural failure risk is measured by the total cost of bridge damage and the risk coefficient; the structural failure risk is expressed as: ; in, represents the structural failure risk, represents the safety risk factor associated with the bridge status s; Maintenance cost and structural failure risk are treated as two independent objective functions, and dynamic weights are introduced to express the optimization tendency under different management preferences. A multi-objective maintenance decision optimization model is constructed. The multi-objective maintenance decision optimization model is expressed as: ; ; in, represents the maintenance cost weight; represents the structural failure risk weight; and are dynamic weights, which are decision variables for multi-objective optimization and are used to reflect the decision maker’s preference for different objectives. The constraints are: .

3. The multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence according to claim 1 is characterized in that: The step of encoding the weight combination into individuals of a multi-objective evolutionary algorithm and randomly generating an initial population comprises: According to the maintenance cost weight – Structural failure risk weight " in the form of a two-dimensional vector, and The chromosomes of individuals encoded as multi-objective evolutionary algorithms, where + = 1, and 、 All values ​​are in the closed interval from 0 to 1; In satisfaction + = 1, random sampling is used to generate N weight vectors as the initial population, where N is the preset population size.

4. The multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence according to claim 1 is characterized in that: For each generation of weight combinations, constructing a bridge group Markov decision environment based on each weight combination includes: A Markov decision environment for a group of bridges was constructed. The state transition probability matrix in the Markov decision environment was set based on annual bridge inspection data, satisfying the assumption that bridges can degrade at most one level in a year without maintenance, and that different maintenance actions correspond to different repair outcomes. Calculate the maintenance cost based on the state of the bridge group at time t and the set of maintenance actions selected by the reinforcement learning agent for the bridge group and structural failure risk ; Using the current weight combination ( , ), according to the formula Calculating the reward function , embed the reward function into the constructed Markov decision environment for the next step of reinforcement learning agent training.

5. The multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence according to claim 1 is characterized in that: The obtained optimal maintenance strategy is evaluated, and the evaluation result is fed back to the multi-objective evolutionary algorithm in step S6 as the individual fitness to guide the next round of weighted combination evolution. After completing the reinforcement learning training corresponding to each weight combination, the optimal maintenance strategy obtained by executing the reinforcement learning agent in the Markov decision environment is used to accumulate the maintenance cost of each time step within a preset simulation cycle to obtain a cumulative maintenance cost; Synchronously calculate the structural failure risk corresponding to each time step within the same preset simulation cycle to obtain the cumulative structural failure risk; The accumulated maintenance cost and the accumulated structural failure risk are combined into a fitness vector, and the fitness vector is fed back to the multi-objective evolutionary algorithm in step S6 to guide a new round of weight update.

6. The multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence according to claim 1 is characterized in that: The multi-objective evolutionary algorithm is NSGA-II; wherein NSGA-II is a non-dominated sorting genetic algorithm-II; The multi-objective evolutionary algorithm is used to perform evolutionary search on the generated population to obtain a new generation of weight combinations including: NSGA-II is applied to the generated population, and fast non-dominated sorting, crowding distance calculation, elite retention, crossover and mutation operations are performed in sequence to obtain a new generation of weight combinations.

7. The multi-objective maintenance decision-making method for bridge groups integrating evolutionary algorithm and artificial intelligence according to claim 6 is characterized in that: The generated population is subjected to NSGA-II, which sequentially performs fast non-dominated sorting, crowding distance calculation, elite retention, crossover, and mutation operations to obtain a new generation of weight combinations including: Perform non-dominated sorting on the generated initial population and divide individuals into different non-dominated levels according to the Pareto dominance relationship; The crowding distance is calculated within each non-dominated layer to measure the individual distribution density and maintain population diversity; A binary tournament selection mechanism based on non-dominated rank and crowding degree is adopted, which gives priority to individuals with better non-dominated rank and selects individuals with larger crowding distance when the ranks are the same to form the parent population; Perform crossover operation on parent individuals to generate offspring individuals; Performing a mutation operation on the offspring individuals to expand the search space; The parent population is merged with the child population, and fast non-dominated sorting and crowding distance calculation are performed. Based on the non-dominated level and crowding distance priority of the individuals, the top N individuals are retained to form a new generation of weighted combinations; where N is the preset population size.

8. A multi-objective maintenance decision-making device for a bridge group, characterized by: The multi-objective maintenance decision-making device for bridge groups includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Hierarchical multi-agent bridge and tunnel group maintenance decision-making method and device based on GA-RL

    CN120806945A

  • Electric tool operation parameter optimization system based on machine learning

    CN120911310A

  • Fault positioning and maintenance management method and device based on vehicle power assembly

    CN122114893A