Multi-objective assembly type bridge engineering construction organization optimization method

By employing an algorithm combining Monte Carlo tree search and deep reinforcement learning in prefabricated bridge engineering, a multi-objective optimization objective function was constructed, solving the balance between construction period and cost, and realizing intelligent optimization of construction strategies and efficiency improvement.

CN120930877APending Publication Date: 2025-11-11HUAZHONG UNIV OF SCI & TECH +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511145768.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies are insufficient for simultaneously balancing multiple objectives of schedule and cost optimization in prefabricated bridge engineering, and construction organization optimization mainly relies on experience-driven approaches and lacks systematic digital rules.

Method used

An algorithm combining Monte Carlo tree search and deep reinforcement learning is used to construct a multi-objective optimization function by optimizing the construction sequence and resource allocation. The DQN network is then used to predict the construction value and generate a construction strategy that balances shortening the construction period and controlling costs.

Benefits of technology

It enables rapid calculation of construction strategies in prefabricated bridge engineering, improves the intelligence and efficiency of construction organization planning, and can dynamically adapt to on-site deviations to optimize the construction period and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930877A_ABST
    Figure CN120930877A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field related to intelligent construction, and particularly relates to a multi-objective assembly type bridge engineering construction organization optimization method. The method comprises the following steps that a multi-objective optimization objective function in the bridge construction process is solved through a Monte Carlo tree search method, and the optimization objective function takes the number of working days and the minimum cost in the bridge construction process as optimization objectives and takes a construction sequence as a decision variable; by introducing a hybrid algorithm of deep reinforcement learning and a Monte Carlo tree search method and a network prediction action value, an MCTS evaluates a state income by simulating a construction path, the two cooperatively generate a construction strategy considering both construction period shortening and cost control, dynamic parameters such as a component state and team and group working hours are coded into a state space, and the state space is optimized. And team scheduling and resource allocation are optimized by using a reward function guidance algorithm, and a strategy is dynamically updated according to field deviation, so that the goal of intelligently formulating an assembly type bridge engineering construction organization plan with relatively good construction period and cost is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent construction technology, and more specifically, relates to a multi-objective optimization method for the construction organization of prefabricated bridge engineering. Background Technology

[0002] Prefabricated bridge engineering is characterized by large components, small spaces, tight schedules, and diverse constraints. Ensuring that the construction organization plan is adapted to the new dynamic construction scenario of prefabricated construction and considering the balance between multiple objectives is a key challenge.

[0003] Traditional methods for construction organization planning, such as manually arranged double-symbol network diagrams and Gantt charts, suffer from low efficiency, high labor input, a singular focus on schedule, and difficulty in cost control. On the one hand, traditional single-objective optimization methods struggle to balance the interplay between multiple objectives, failing to deeply integrate with modern construction organization. On the other hand, current construction organization optimization relies heavily on experience, making it difficult to quantify the cost logic of new prefabricated engineering scenarios and lacking systematic digital rules. Therefore, a multi-objective optimization method for prefabricated bridge engineering construction organization is urgently needed. Summary of the Invention

[0004] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a multi-objective prefabricated bridge engineering construction organization optimization method to solve the problem of the inability to balance multi-objective optimization.

[0005] To achieve the above objectives, according to one aspect of the present invention, a multi-objective method for optimizing the construction organization of prefabricated bridge engineering is provided, the method comprising the following steps: The Monte Carlo tree search method is used to solve the multi-objective optimization objective function in the bridge construction process. The optimization objective function takes minimizing the number of working days and cost in the bridge construction process as the optimization objective, and the construction sequence as the decision variable. The optimization objective function is as follows: ; in, Indicates the number of working days. X represents the total cost, and X represents the construction sequence. Total hours of all working days.

[0006] More preferably, the constraint condition of the optimization objective function is an environmental action constraint, as follows: ; ; Where x represents the action and i is the construction team number. j represents cross-index. k represents the process type. n is the total number of spans of the bridge to be constructed, n1 is the number of construction teams of type 0, n2 is the number of construction teams of type 1, n3 is the number of construction teams of type 2, and n4 is the number of construction teams of type 3.

[0007] More preferably, the Monte Carlo tree search method includes the following steps: S1 sets the loop count to cnt=0; S2 takes the current node as the parent node, randomly samples a path from the parent node to the child node from the Monte Carlo tree, updates the current node to the child node pointed to by the selected path, updates the environment state space of the current node and calculates the environment reward. S3 Determine if the current node is a leaf node. If it is not a leaf node, return to step S2. If it is a leaf node, then determine whether the current node's environment state space is in a terminated state; If it is not in a terminated state, start from the current node and expand to a new node, and use the expanded node as the current node, then return to S3; If it is in a terminated state, cnt = cnt + 1, initialize the environment parameters of the current node, calculate the cumulative reward of the current environment, and establish the relationship between the decision variable corresponding to the path to the current node and the solution in the solution set; if there is no solution in the solution set that dominates the decision variable, then store the path to the current node and the cumulative reward of the current environment in the solution set; otherwise, do not store them. S4 checks if the current loop count cnt has reached the preset iteration count. If not, it returns to S2; otherwise, it ends.

[0008] More preferably, the method for constructing the Monte Carlo tree is as follows: S11 Initialize the root node; S12 Determines whether the current node has been fully expanded; If the current node is not fully expanded and the randomly generated number p does not meet the preset conditions, then expand the new node with the current node as the starting point, and take the expanded node as the current node, and return to step S12. When the current node is not fully expanded and the randomly generated number p satisfies the preset conditions, or when the current node is fully expanded and the current node is not in a terminated state, select a node from the child nodes as the current node and return to step S12. The process ends when the current node is fully expanded and is in a terminated state. More preferably, whether the current node is fully expanded refers to whether the child nodes of the current node correspond one-to-one with the nodes that satisfy the current environment action constraints. A one-to-one correspondence indicates full expansion, while a non-one-to-one correspondence indicates incomplete expansion.

[0009] More preferably, the selection of a node as the current node from among the child nodes is performed in the following manner: (1) Determine whether the number of visits to each child node has reached the preset value. If the preset value is reached, update the PUCT value according to the following formula; otherwise, update the UCT value according to the following formula. ; ; In the formula, C is a constant. It is the number of times the node's parent node has been visited. It is the number of times the node is accessed. It is the average value of the nodes. is the prior probability of the path; s refers to the state space, and a refers to a legal action when the state space is s; (2) Select the child node corresponding to the maximum value of PUCT value and UCT value as the current node.

[0010] More preferably, the process of expanding new nodes from the current node is performed according to the following steps: Randomly select a point from the points that satisfy the current environment action constraints as an expansion node, calculate the value of the expansion node, and update the Q value (Q(s,a)) and the number of visits (N(s,a)) of all nodes on the path leading to the current node in reverse order from the current node. Q is the reward value for continuing construction until the project is completed after selecting construction action a when the actual construction state is in state space s. The update method for Q value and number of visits is as follows: Determine if the number of visits to a node on the path is 0. If it is not 0, then Q = Q + (value of the extended node - Q) / (historical number of visits to the node + 1). Otherwise, Q equals the value of the extended node. The node's access count increases by one.

[0011] More preferably, the value of the node is calculated using a DQN network, where the input of the DQN network is the node's environmental state space, and the output is the node's value.

[0012] More preferably, the method for determining whether the current node's environmental state space is in a terminated state is: the current node's process state matrix does not contain 0 and does not contain 1.

[0013] More preferably, in step S2, the random selection is a random sampling based on the probability distribution of the PUCT values ​​of each child node.

[0014] In summary, the technical solutions conceived by this invention have the following beneficial effects compared with the prior art: 1. The multi-objective prefabricated bridge construction organization optimization method proposed in this invention solves the problem that single-objective optimization methods in the prior art cannot simultaneously balance the game relationship between multiple objectives and cannot be deeply coupled with modern construction organization by taking the construction period and cost as optimization objectives to form two parallel optimization objective functions. 2. This invention introduces a hybrid algorithm combining Deep Reinforcement Learning (DRL) and Monte Carlo Tree Search (MCTS) with a Deep Q-learning Network (DQN) to predict action value. MCTS evaluates state benefits by simulating construction paths. The two work together to generate a construction strategy that balances shortening the construction period and controlling costs. Dynamic parameters such as component status and shift work hours are encoded into a state space. The reward function guides the algorithm to optimize shift scheduling and resource allocation. Furthermore, the strategy can be dynamically updated based on on-site deviations, achieving the goal of intelligently formulating a prefabricated bridge construction organization plan with optimal construction period and cost.

[0015] 3. In this invention, the DQN is used to calculate node value in the MCTS search algorithm. Its greatest advantage is that it can quickly calculate node value regardless of the problem dimension. Traditional methods use a rollout simulation method to calculate node value, which requires updating the simulation environment step by step to the final state to obtain the node value. However, updating the simulation environment step by step requires additional computing power, especially for high-dimensional problems. Not only is the time of a single rollout very long (related to the tree depth), but many rollouts are also required (related to the total number of nodes in the tree), resulting in extremely low algorithm efficiency. Although this method does not reduce the number of rollouts, it significantly reduces the time of a single rollout, thus greatly improving the overall efficiency. It is suitable for high-dimensional complex problems such as the optimization of construction organization for prefabricated bridge engineering involving multiple components and personnel.

[0016] 4. This invention emphasizes the depth-first search algorithm in the MCTS search algorithm. Due to the nature of engineering problems, rewards are sparse, leading to poor direction and low search efficiency in traditional methods during the early exploration stages. Emphasizing depth-first search allows for a rapid attainment of a non-zero reward at the termination state, solving the "zero reward trap" problem. Combined with a balancing strategy, it compensates for the breadth-first search deficiency and can escape local optima. This method also provides valuable experience for training DQN networks. When DQN is sufficiently accurate, weakening the depth-first search can yield a broader search, but for specific tree branches, a more refined search using the depth-first algorithm is still necessary. Attached Figure Description

[0017] Figure 1 This is an overall flowchart of a multi-objective prefabricated bridge construction organization optimization method constructed according to a preferred embodiment of the present invention.

[0018] Figure 2 This is a flowchart of the Monte Carlo search algorithm constructed according to a preferred embodiment of the present invention.

[0019] Figure 3 This is a flowchart of constructing a Monte Carlo tree according to a preferred embodiment of the present invention.

[0020] Figure 4 This is a flowchart of the DQN algorithm constructed according to a preferred embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0022] like Figure 1 As shown, a multi-objective optimization method for the construction organization of prefabricated bridge engineering is proposed, which includes the following steps: S1. User inputs initial parameters: construction start date, daily working hours, and bridge work section information; Initial parameters include: construction start date, daily working hours, span of each span of the bridge to be constructed, components to be constructed, and component composition types.

[0023] Specifically, the start date of construction must be a year-month-day string of length 8, such as "20250101"; the daily working hours must be a floating-point number in hours, such as 8.5h.

[0024] In one embodiment of the present invention, a prefabricated bridge with a total of 19 spans needs to be constructed, including 18 piers, 18 cap beams and 19 beam segments.

[0025] Specifically, the bridge construction section information includes: the span of each span of the bridge to be constructed, the components to be constructed, and the component composition type. The span of each span is represented sequentially by a string in the form a*b+c*d, where the left multiplication character (e.g., ac) represents the quantity, and the right multiplication character (e.g., bd) represents the span, in meters. In this embodiment, "1*24+6*32+2*24+10*32" represents the first span as 24m, the second to seventh spans as 32m, and so on. The components to be constructed and their composition types are required to be in string encoding format. D, G, and L represent the component type, referring to piers, cap beams, and beam segments respectively. The numeric characters are the span numbers, indicating which span. For example, D3 represents the pier component of the third span. All components to be constructed are input as a set; completed bridge structural components within the bridge construction section are not included.

[0026] S2. User settings for other parameters and constraints: unit price of machinery and materials for each person, composition of construction teams, construction efficiency information, number of construction teams, and information on completed work areas; Other parameters and constraints include: daily and monthly wages for all positions, rental and access costs for machinery, unit price of components and fuel costs for machinery, types and numbers of personnel in each construction team, and construction efficiency for different types of components.

[0027] Specifically, in one embodiment of the present invention, the unit price of personnel, machinery and materials includes: management personnel, safety officers, quality inspectors, surveyors, crane operators, general workers, installers, hoisting workers, signalmen, aerial work platform operators, truck cranes, truck crane operators, truck crane fuel costs, truck crane entry and exit fees, flatbed truck transportation costs, component quality, etc., in yuan, yuan / day, yuan / month, or tons, and the data type is floating point.

[0028] The construction team consists of the number of workers and machinery required for the three types of components: piers, cap beams, and beam segments. For example, the construction team for pier components includes 1 manager, 1 safety officer, 1 quality inspector, 2 surveyors, 1 crane operator, 2 general workers, 1 hoisting worker, 1 signalman, 1 truck crane operator, 1 aerial work platform operator, 1 truck crane operator, and 1 flatbed truck operator. Each construction team is independent of the others.

[0029] Construction efficiency information includes: for different types of components, the expected construction time obtained from actual construction cases is used as the construction efficiency. The type is a floating-point number and the unit is hours. For example, Da corresponds to 2.0h and Db corresponds to 2.5h. It is only related to the type of component.

[0030] The number of construction teams is represented by n1 pier construction teams, n2 cap beam construction teams, n3 beam segment construction teams, and n4 general construction teams (capable of performing construction work on all components). n1, n2, n3, and n4 are non-negative integers and are of integer data type.

[0031] The information on completed work surfaces is represented as a set of completed component codes, which is consistent with the type of information on components to be constructed in S2. This stage can be skipped if there are no completed components. During the implementation phase, when there is a deviation between the actual progress and the plan, the completed components can be input into the system, and the algorithm can be called again to generate a construction organization plan for the remaining components to adapt to changes in actual construction.

[0032] S3. Call the multi-objective optimization algorithm: output the optimal strategy for the construction organization plan; (a) Setting decision variables, objective function and constraints The decision variables for the optimization problem in this embodiment are as follows: ; The decision variable is a sequence of actions, which can be viewed as a vector. This represents an action. A decision variable corresponds to a path from the root node to a leaf node in the terminal state.

[0033] The objective function for the optimization problem in this embodiment is as follows: ; Indicates the number of working days. Represent the cost, and minimize both.

[0034] The environment model contains two functions, which are part of the environment model. As the environment is updated, it can output the values ​​of these two functions when it reaches the final state. 1. ; It indicates Total duration of all workdays, final time step.

[0035] Let the sequence of decision variables be: ; Where i, j, and k are the codes for action x.

[0036] The state evolves to: ; For the final state set, It can then be expressed as: ; 2. ; In the second formula, m represents the month. Let $\mathbf{m}$ represent the usage of construction teams in month $m$, which is a row vector. $\mathbf{price}$ represents the monthly cost of each construction team, which is a column vector. That is the total cost.

[0037] Each m under It is obtained from the environmental input decision variable X, that is, the following formula: Let the sequence of decision variables be: ; Where i, j, and k are the codes for action x.

[0038] Let the interval step number corresponding to month m be: ; It can then be expressed as: ; The 1(.) function is an indicator function; it returns 1 if the condition is true and 0 otherwise.

[0039] (ii) Setting up the environmental model A. Initial input parameters: 1) The number of spans n of the prefabricated bridge to be constructed.

[0040] 2) The number of four types of construction teams, n1, n2, n3, and n4, where n1 to n3 represent the number of construction teams that can only construct piers, cap beams, and beams, respectively, and n4 represents the number of all-round construction teams, that is, teams that can construct all three types of components.

[0041] 3) The component construction efficiency matrix duration_matrix has a size of n*3. The value of (j, k) in this matrix represents the average construction efficiency of the k-th component in the (j+1)-th span. k takes values ​​of 0, 1, and 2, corresponding to the three types of components: pier, cap beam, and beam segment, respectively.

[0042] 4) Bridge span information bridge_span_st, string type, such as 1*24+6*32+2*24, which means the first span is 24m, the second to seventh spans are 32m, and the eighth and ninth spans are 24m.

[0043] 5) Crane travel speed, transport_speed B. The node's environment state space s is defined, including the following parameters: 1) The process state matrix, with a size of n*3, has a value S. ps The value of (j, k) represents the construction status of the k-th component in the (j+1)-th span: -1 indicates that the process does not exist, 0 indicates that the process has not been constructed, 1 indicates that the process is under construction, and 2 indicates that the process has been completed.

[0044] 2) The remaining construction time matrix, with a size of n*3, is S. rt The value of (j, k) represents the remaining construction time of the k-th component in the (j+1)-th span: the value is the remaining completion time, and -1 indicates that the process is not under construction.

[0045] 3) The construction team's state matrix, with a size of 4*max(n1, n2, n3, n4), has a value S. ts The value of (a, b) represents the state of the b-th construction team of type a. The value of a ranges from 0 to 3, corresponding to the four types of construction teams: pier, cap beam, beam segment, and all-round construction team. The value of b represents the number of the corresponding construction team: a value of -1 indicates that the construction team does not exist. For example, when n1=0, (0,0) does not exist. 0 indicates that the construction team is in an idle state, and 1 indicates that the construction team is unavailable, that is, in a construction state.

[0046] 4) The construction date is a scalar, indicating the current working day. C. Action space definition, including the following parameters A path is a sequence of actions, an action is an element in the sequence, and the action space is the set of all actions.

[0047] The action is a three-dimensional matrix with a three-dimensional space size of (n1+n2+n3+n4)×n×3. The element A(i,j,k) in the three-dimensional matrix represents the construction team number, j represents the corresponding span number, and k represents the process type.

[0048] The first n1 terms of i are the numbers of the pier construction teams, n1 to n1+n2 are the numbers of the cap beam construction teams, n1+n2 to n1+n2+n3 are the numbers of the beam construction teams, and the last n4 terms are the numbers of the all-around construction teams. j represents the (j+1)th span; The values ​​of k are 0, 1, and 2, which correspond to piers, cap beams, and beam segments, respectively. A(i,j,k) represents the construction team i performing the k-th process of the (j+1)-th span.

[0049] D. Setting the reward function The reward is a three-dimensional reward. The first dimension represents the construction period, the second dimension represents the cost, and the third dimension represents the additional transportation time. Currently, a non-zero reward is only returned when the game reaches its end; otherwise, the reward is 0, i.e., [0,0,0]. 1. Calculation method of reward function R for project duration t =bk / (total_step / / 600+1), where b is the bias term, k is the coefficient, and total_step is the total number of steps. The total number of steps is updated synchronously as the process progresses, and the environment can be obtained directly.

[0050] 2. Cost reward function calculation method R c =b - k_cost * cost, where b is the bias term, k_cost is the coefficient, and cost is the total cost. The cost needs to be calculated based on the environment, with a span of 30 days, or 30 * 600 steps per span. Within the time interval, the number of construction teams that have never carried out any construction work is removed from the construction teams. The number of construction teams that have carried out construction work is multiplied by the corresponding monthly cost of the construction team to obtain the cost of the current monthly span. All monthly costs are accumulated to obtain the final cost. For example, if total_step = 20000, that is the cost for two months. The first 18000 is considered as one month, and the last 2000, which is less than one month, is also calculated using the monthly cost.

[0051] 3. Calculation method for the additional time-consuming reward function. ; Here, the code corresponding to the current action is denoted as A'(i,j',k'), and the action executed by the previous construction team i corresponding to this action is denoted as A(i,j,k). This represents the distance between the (j+1)th span and the (j'+1)th span, which can be derived from... (This was obtained from a record of all span lengths), where v represents the construction team's transport speed. , This represents the additional time consumed by construction team i during the current action.

[0052] E. Determining the legality of actions, i.e., how to obtain the set of potential child nodes to be expanded in the MCTS tree while satisfying the following five conditions: The constraints of the optimization problem in this embodiment are as follows: Each construction team (4 types: pier, cap beam, beam segment, all-rounder) can choose actions (i,j,k) at each decision point, which must satisfy the following: 1. Match construction team type with work process (except for all-around teams); the correspondence between i and k is restricted here. The values ​​of are 0, 1, 2, and 3, corresponding to the values ​​of k being 0, 1, and 2. When i ≠ 3, a match is only made when i = k are equal. This content does not need to be considered when the value is 3.

[0053] 2. The preconditions of the process represented by the action are met (for example, the cap beam process requires that the pier column process of the current span has been completed; the pier column process has no preconditions; the beam process requires that the cap beam processes of the current span j+1 and the previous span j have been completed. In particular, since there is no cap beam for the first span and no cap beam for the last span, the precondition for the former is that the cap beam of the current span j+1 has been completed, and the precondition for the latter is that the cap beam of the jth span has been completed). 3. The construction team was idle at the time of the decision-making; When selecting an action, only available construction teams can be chosen. The selected action must represent an available construction team i; otherwise, it is invalid. In the state matrix of the construction team corresponding to action i, a value of 0 indicates that it is available. Based on i, S can be obtained. ts Given (a, b), the formulas for a and b based on i are as follows: ; Here, 'c' is only used for auxiliary calculations and represents only the type of construction team i. How many teams were there before?

[0054] 4. During decision-making, the current state matrix of the process corresponds to element S. ps If (j,k) is 0, then no construction has been carried out. ;

[0055] For the column pier process (k=0): No prerequisites.

[0056] For the cap beam construction process (k=1): it is required that the pier columns of the same span have been completed; For the beam segment process (k=2): it is required that the cap beam of the same span has been completed, and if there is a previous span, its cap beam has also been completed.

[0057] 5. Action x in action sequence X satisfies the condition that it does not exceed the size of the action space. ; Action x can be represented in the form (i,j,k) and has the following calculation formula: ; ; remember ; Let i be the type of team i. Its values ​​0, 1, and 2 correspond to the process types when k is 0, 1, and 2, respectively. The value 3 corresponds to any k, i.e., k is 0, 1, or 2. This represents the state of the k-th process across j+1, with values ​​of -1 indicating non-existence, 0 indicating no construction, 1 indicating construction in progress, and 2 indicating completion. This represents the state of team i, with values ​​of -1 indicating non-existence, 0 indicating idle, and 1 indicating construction.

[0058] F. How does the environment obtain the next state based on the current state space s and the input action a, and update the next state?

[0059] First, based on the determination of legal actions, a set of legal actions 'a' in the current state space 's' is obtained. Then, the input 'a' must be one of the legal actions in the set of legal actions. Executing 'a' updates the information in 's': the process status corresponding to action 'a' is updated to "under construction," the remaining time is updated to the efficiency of action 'a', the construction team status is updated to "unavailable," and the construction date information 'day' is not updated temporarily. The efficiency of action 'a' is the element (j,k) in the construction efficiency matrix corresponding to action 'a(i,j,k)'.

[0060] At this point, an intermediate state s is obtained. Next, it is determined whether a valid action exists. If it does, the intermediate state is returned directly. If not, a jump-step strategy is needed, as follows: Let jump = the smallest non-negative value in the remaining construction time matrix, total_step += jump, and simultaneously update the process state to 2, the remaining time length to -1, and the construction team state to 0. At this point, at least one remaining time length must be 0; update it to -1, and update the corresponding process state to "completed," and change the corresponding construction team state to "idle." Then, the final updated state space s is returned. G. Method for determining whether the environment state space is a terminated state The termination state is: the current node's process state matrix contains neither 0 nor 1.

[0061] H. Node environment parameter initialization Specifically: initialize the state space s, with total_step set to 0; the set of components under construction is empty. total_step represents the virtual physical time in the environment, in minutes, and is used to represent the date information "day" in the current state space.

[0062] The state space s is initialized as follows: In the process status matrix, the value of a valid process is 0, while the values ​​of invalid processes are the last span's piers and cap beams, which are -1. In the construction team state matrix, all valid construction teams have a value of 0, while invalid construction teams have a value of -1. For example, if n1~n4 have values ​​of 1, 1, 1, and 2, then the matrix size is 4*2, and the second column of the first three rows is invalid. The remaining time state matrix takes values ​​of -1. The construction period (day) is set to 1. (iii) If Figure 2 As shown, the specific steps of the multi-objective optimization algorithm used to solve the above environment model are as follows: DM-Step 1: Read the existing MCTS tree structure, DQN network structure, and initialize the environment; DM-Step2: Set the number of loops N, and define the variable cnt=0; DM-Step3: Determine if the number of iterations has reached the target. If yes, proceed to DM-Step4; otherwise, proceed to DM-Step15. DM-Step4: Calculate the PUCT value of each child node starting from the current node; PUCT stands for Predictor Upper Confidence bound applied to Trees. It means that the upper confidence interval of the input is based on the prior probability. Nodes with high node value, high prior probability, and low access frequency will have larger PUCT values, which means they are more likely to be selected, thus guiding the direction of selection. ;

[0063] The first term of the formula refers to the average value of each node visit, which means that actions with higher average value will be selected; the second term is the exploration term. The second term, which has a higher prior probability and fewer visits, has a larger value, meaning that actions that have not been fully explored but that the network predicts have potential will be continuously tried.

[0064] Q(s,a) represents the total reward expected to be obtained after performing construction action a when the construction state is s; N(s,a) represents the total number of times the process of performing construction action a when the construction state is s has been performed in the optimization algorithm; N(s) is the sum of N(s,a), representing the total number of times the construction state s has been reached in the optimization algorithm; P(a|s)=N(s,a) / N(s) represents the proportion of the process of performing construction action a when the construction state is s in the total number of times the construction state s has been reached, based on the current exploration space of the optimization algorithm, and represents the statistical probability of performing construction action a when the construction state is s; C is the coefficient that balances exploration and utilization, balancing the values ​​of the first and second terms in the formula.

[0065] DM-Step5: Use softmax to output the vector composed of PUCT of each node as the probability distribution, randomly sample, select a child node, take the selected child node as the current node, update the state space of the environment synchronously to be consistent with the current node, and record the selected path and the environmental feedback reward rt; DM-Step6: Determine if the current node is a leaf node. If yes, proceed to DM-Step7; otherwise, proceed to DM-Step4.

[0066] DM-Step7: Determine if the current node is in a terminated state. If yes, proceed to DM-Step8; otherwise, proceed to DM-Step11. DM-Step8: cnt = cnt + 1, and reset the environment; DM-Step9: Record the path and cumulative reward ∑r and determine the dominance relationship with the non-dominated solutions in the Pareto solution set. If not dominated, proceed to DM-Step10; if dominated, proceed to DM-Step3. The dominance relationship is as follows: if the objective function value of solution A is not inferior to the objective function value of solution B, then solution B is dominated by solution A; otherwise, there is no dominance relationship between A and B.

[0067] DM-Step10: Save path and ∑r to the pareto set, and proceed to DM-Step3; DM-Step11: Randomly expand a child node to the current node, call DQN to input the current state and output the simulated prediction value; DM-Step12: Update the tree structure in reverse order, visiting all nodes on the path from back to front; DM-Step13: Determine if node N(s,a) is 0. If yes, use the weighted network output as Q(s,a) for that node; otherwise, use the weighted output as Q(s,a). , where Vsim is the weighted network output; DM-Step14: N(s,a)+=1, and set the newly expanded node as the current node, then proceed to DM-Step6; DM-Step 15: Output the Pareto set; DM-Step 16: End.

[0068] As a preferred option, the DQN network, MCTS tree structure, and environment model in DM-Step1 are obtained through different methods.

[0069] (iv) such as Figure 3 As shown, the construction of the MCTS tree structure includes the following steps: SM1: Initialize the MCTS root node; The initial state space of the root node is as follows: The root node's state space contains the process state matrix, remaining construction time matrix, construction team state matrix, and construction period as follows: 1. Process state matrix S ps The size is n*3, the corresponding elements (n-1,0) and (n-1,1) take the value -1, and the other elements take the value 0.

[0070] 2. Remaining construction time matrix S rt The size is n*3, and all elements are -1.

[0071] 3. Construction team state matrix S ts , 4*max(n1, n2, n3, n4), the first n1 elements of the first row are 0, the first n2 elements of the second row are 0, the first n3 elements of the third row are 0, the first n4 elements of the fourth row are 0, and the remaining elements are -1.

[0072] 4. The construction date is set to 1.

[0073] The parent and child node sets are both empty, the expected value of the node Q(s,a) is 0, the number of node visits N(s,a) is 0, and the action a from the previous state to the current state is empty.

[0074] SM2: Determine whether the current node has been fully expanded. If it has not been fully expanded, proceed to SM3; if it has been fully expanded, proceed to SM13. Whether the current node is fully expanded refers to whether the child nodes of the current node correspond one-to-one with the nodes that satisfy the action constraints of the current environment. If they do not correspond one-to-one, it is considered fully expanded; if they correspond one-to-one, it is considered partially expanded. SM3: Randomly generate a number , ,like Then proceed to SM4; otherwise proceed to SM7. SM4: Calculate the UCT or PUCT values ​​of the child nodes of the current node respectively; SM5: For each node, determine if N(s,a)>=N p If yes, calculate PUCT; otherwise, calculate UCT. ; ; ; Where C is the exploration coefficient, which is a constant; P(a|s) is the probability distribution; and N(s,a) and Q(s,a) are the number of node visits and the expected value of each node stored in each node. SM6: Select the node with the largest UCT or PUCT value as the current node and return SM2; SM7: Randomly expand a new node from the current node that is an expansion node. Expand the new node based on the interaction with the environment. The current node state st, the randomly selected legal action a is input into the environment, the environment returns the next state st+1, the timely reward rt, the termination information done, the action mask in the st state and the mask in the st+1 state, the next state represents the expanded new node, and the new node is used as the current node. SM8: Simulation. The current node is input into the DQN network, which outputs multiple multi-dimensional vectors. Each multi-dimensional vector represents the predicted value of the current node's state transitioning to other states. The network sets the output values ​​of illegal transitions (i.e., illegal states) to zero vectors, and weights the non-zero vectors of legal transitions by w. The maximum scalar obtained is used as the simulation return value Vsim. SM9: Reverse update the tree structure, visiting all nodes on the path from back to front; SM10: Determine if node N(s,a) is 0. If yes, use the weighted network output as Q(s,a) for that node; otherwise, use the weighted output for that node. , where Vsim is the weighted network output; SM11: N(s,a)+=1, and set the newly expanded node as the current node; SM12: Stores the previous node state st obtained from SM7, the randomly selected illegal action a, the current node state st+1 returned by the environment, the timely reward rt, and other information info into the experience pool to train the DQN network.

[0075] Other information includes: the legal action mask corresponding to st and st+1 (reflecting information about the legal action set), done flag information (True or False, reflecting whether st+1 is in a terminated state), and total_step (total number of steps, representing time in minutes).

[0076] SM13: Determine whether the fully expanded node is in an aborted state. If yes, proceed to SM14; otherwise, proceed to SM4. SM14: End.

[0077] (v) such as Figure 4 As shown, training a DQN network includes the following steps: T1: Initialize DQN network parameters. Two identical networks, one called the main network and the other called the target network. Set the number of learning iterations M, initialize m=0, batch size, learning rate lr, γ, and update the target network period N. T2: Extract batch-sized data from the experience pool. Each data entry includes the current node state st, the randomly selected legal action at, the next state st+1, the timely reward rt, the action mask in state st, and the action mask in state st+1. T3: Input st into the main network, output the corresponding Q(ai) in all action spaces, and find the Q(at) corresponding to at; T4: Input st+1 into the target network and output Q(ai) corresponding to all action spaces. Based on the action mask of the st+1 state, filter the Q(a') corresponding to the legal actions and record the maximum value as max Q(a'). T5: Calculate the tag value according to the Bellman formula. ; T6: Calculate MSE loss ; T7: Backpropagation gradient update T8: Check if ++m%N==0. If yes, copy the main network parameters to the target network parameters and proceed to T9; otherwise, proceed directly to T9. T9: Determine if m <= M; if yes, proceed to T2; otherwise, proceed to T10. T10: End.

[0078] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-objective optimization method for the construction organization of prefabricated bridge engineering, characterized in that, The method includes the following steps: The Monte Carlo tree search method is used to solve the multi-objective optimization objective function in the bridge construction process. The optimization objective function takes minimizing the number of working days and cost in the bridge construction process as the optimization objective, and the construction sequence as the decision variable. The optimization objective function is as follows: ; in, Indicates the number of working days. X represents the total cost, and X represents the construction sequence. Total hours of all working days.

2. The multi-objective prefabricated bridge construction organization optimization method as described in claim 1, characterized in that, The constraints of the optimization objective function are environmental action constraints, as follows: ; Where x represents the action and i is the construction team number. j represents cross-index. k represents the process type. n is the total number of spans of the bridge to be constructed, n1 is the number of construction teams of type 0, n2 is the number of construction teams of type 1, n3 is the number of construction teams of type 2, and n4 is the number of construction teams of type 3.

3. A multi-objective prefabricated bridge construction organization optimization method as described in claim 1 or 2, characterized in that, The Monte Carlo tree search method includes the following steps: S1 sets the loop count to cnt=0; S2 takes the current node as the parent node, randomly samples a path from the parent node to the child node from the Monte Carlo tree, updates the current node to the child node pointed to by the selected path, updates the environment state space of the current node and calculates the environment reward. S3 Determine if the current node is a leaf node. If it is not a leaf node, return to step S2. If it is a leaf node, then determine whether the current node's environment state space is in a terminated state; If it is not in a terminated state, start from the current node and expand to a new node, and use the expanded node as the current node, then return to S3; If it is in a terminated state, cnt = cnt + 1, initialize the environmental parameters of the current node, calculate the cumulative reward of the current environment, and establish the relationship between the decision variables corresponding to the path to the current node and the solutions in the solution set. If there is no solution in the solution set that dominates the decision variable, then the path to the current node and the cumulative reward of the current environment are stored in the solution set. Otherwise, do not store; S4 checks if the current loop count cnt has reached the preset iteration count. If not, it returns to S2; otherwise, it ends.

4. The multi-objective prefabricated bridge construction organization optimization method as described in claim 3, characterized in that, The Monte Carlo tree is constructed as follows: S11 Initialize the root node; S12 Determines whether the current node has been fully expanded; If the current node is not fully expanded and the randomly generated number p does not meet the preset conditions, then expand the new node with the current node as the starting point, and take the expanded node as the current node, and return to step S12. When the current node is not fully expanded and the randomly generated number p satisfies the preset conditions, or when the current node is fully expanded and the current node is not in a terminated state, select a node from the child nodes as the current node and return to step S12. The process ends when the current node is not fully expanded and is in a terminated state.

5. The multi-objective prefabricated bridge construction organization optimization method as described in claim 4, characterized in that, Whether the current node is fully expanded refers to whether the child nodes of the current node correspond one-to-one with the nodes that satisfy the current environment's action constraints. A one-to-one correspondence indicates full expansion, while a non-one-to-one correspondence indicates incomplete expansion.

6. The multi-objective prefabricated bridge construction organization optimization method as described in claim 5, characterized in that, The selection of a node from the child nodes as the current node is performed in the following manner: (1) Determine whether the number of visits to each child node has reached the preset value. If the preset value is reached, update the PUCT value according to the following formula; otherwise, update the UCT value according to the following formula. ; In the formula, C is a constant. It is the number of times the node's parent node has been visited. It is the number of times the node is accessed. It is the average value of the nodes. is the prior probability of the path; s refers to the state space, and a refers to a legal action when the state space is s; (2) Select the child node corresponding to the maximum value between PUCT and UCT as the current node.

7. The multi-objective prefabricated bridge construction organization optimization method as described in claim 4, characterized in that, The process of expanding new nodes from the current node is performed according to the following steps: Randomly select a point from the points that satisfy the current environmental action constraints as an expansion node, calculate the value of the expansion node, and update the Q(s,a) value and visit count of all nodes on the path leading to the current node in reverse order from the current node. Q(s,a) is the reward value for continuing construction until the project is completed after selecting construction action a when the actual construction state is state space s. The update method for Q(s,a) value and visit count is as follows: Determine if the number of visits to a node on the path is 0. If it is not 0, then Q = Q + (value of the extended node - Q) / (historical number of visits to the node + 1). Otherwise, Q equals the value of the extended node. The node's access count increases by one.

8. The multi-objective prefabricated bridge construction organization optimization method as described in claim 7, characterized in that, The value of a node is calculated using a DQN network, where the input to the DQN network is the node's environmental state space, and the output is the node's value.

9. The multi-objective prefabricated bridge construction organization optimization method as described in claim 4, characterized in that, The method to determine whether the current node's environment state space is in a terminated state is: the current node's process state matrix does not contain 0 and does not contain 1.

10. The multi-objective prefabricated bridge construction organization optimization method as described in claim 4, characterized in that, In step S2, the random sampling selection is based on the probability distribution of the PUCT values ​​of each child node.

Citation Information

Patent Citations

  • DAG task scheduling method based on Monte Carlo tree search

    CN109857532A

  • Heterogeneous modular robot self-reconfiguration planning method based on enhanced learning algorithm

    CN110297490A

  • DQN and MCTS-based inter-box multi-field bridge dynamic scheduling method

    CN112836974A

  • Large-span bridge construction resource scheduling optimization method

    CN118569609A