A construction site monitoring layout method based on deep reinforcement learning

By applying a multi-agent reinforcement learning model based on deep reinforcement learning in the construction site, the monitoring equipment is modeled as an agent, which solves the problems of low efficiency and poor flexibility of traditional monitoring layout methods, and achieves efficient and accurate monitoring layout and equipment optimization.

CN119849888BActive Publication Date: 2025-05-30CHONGQING UNIV IND TECH RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510329424.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-05-30
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Traditional construction site monitoring and layout methods are inefficient and have poor flexibility, and it is difficult to quickly provide layout solutions based on intelligent algorithms.

Method used

The multi-agent reinforcement learning model based on deep reinforcement learning is adopted to model the monitoring equipment as an agent, and the coordinated and adaptive arrangement of monitoring equipment is achieved through deep reinforcement learning and guidance of common goals.

Benefits of technology

The intelligent layout of construction site monitoring is realized, the layout efficiency and accuracy are significantly improved, the number and layout of monitoring equipment are balanced, the monitoring coverage is maximized, and the equipment purchase, installation and construction costs are minimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849888B_ABST
    Figure CN119849888B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for monitoring layout of construction sites based on deep reinforcement learning, which relates to the technical field of model application. In the present invention, the monitoring layout is modeled as the design of a multi-agent reinforcement learning model. Through the deep reinforcement learning of multiple agents, under the guidance of a common goal, each agent is prompted to act in a coordinated and adaptive manner, and finally the intelligent layout of construction site monitoring is realized. Moreover, the method for intelligent layout of construction site monitoring based on deep reinforcement learning proposed by the present invention can efficiently and automatically design the monitoring layout. Compared with the traditional manual experience and fixed rule methods, it can significantly improve the layout efficiency and accuracy. By treating the monitoring devices as agents for joint optimization, global coordination and adaptive layout can be achieved, avoiding the inflexibility and inefficiency problems brought by manual methods, and improving the monitoring layout efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model applications, and particularly to a construction site monitoring layout method based on deep reinforcement learning. Background Art

[0002] With the continuous expansion of the scale of construction sites, the demand for site monitoring layout has gradually increased;

[0003] Traditional monitoring layout methods often rely on manual experience and fixed rules, suffering from low efficiency and poor flexibility;

[0004] The construction site monitoring layout methods based on intelligent algorithms often require long-term iterative calculations and are difficult to quickly provide layout solutions.

[0005] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide a construction site monitoring layout method based on deep reinforcement learning to solve the technical problems raised in the background art.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A construction site monitoring layout method based on deep reinforcement learning, which at least includes the following steps:

[0008] S1: Model the monitoring layout problem as a multi-agent reinforcement learning model, with each monitoring being regarded as an agent, namely a monitoring-agent. Each monitoring-agent moves freely in the construction site, and the monitoring-agents cooperate with each other. They cooperate with each other to achieve comprehensive monitoring coverage and cost minimization, and finally determine their own positions;

[0009] S2: Build a monitoring-agent based on deep reinforcement learning. Through the guidance of deep reinforcement learning and a common goal, each agent is prompted to act in a coordinated and adaptive manner, and finally realize the intelligent layout of construction site monitoring;

[0010] S3: Build an agent network architecture based on the attention mechanism.

[0011] Further, the calculation process of modeling the monitoring layout at least includes the following steps:

[0012] Convert the monitoring layout problem from a planar design problem into a path planning problem that changes over time and with the environment;

[0013] At this time, the mathematical model conforms to the mathematical modeling principles of multi-agent reinforcement learning;

[0014] As each monitoring-agent i takes an action according to the current state and the current time step t , the environment feeds back a new state and a reward based on the current state and action ;

[0015] For each agent, its goal is to maximize the expected cumulative reward, as shown in Equation (1):

[0016]

[0017] where is the discount factor, taking 0.99, which is used to represent the current value of future rewards; N represents the number of monitoring-agents.

[0018] Furthermore, the deep reinforcement learning of the monitoring-agent at least includes the action of the monitoring-agent, the observation content of the monitoring-agent, and the reward of the monitoring-agent.

[0019] Furthermore, the action format output by each agent in the action of the monitoring-agent is the same, that is ;

[0020] where represents the type of monitor i, represents the position of monitor i in the plane, represents the height of monitor i, represents the orientation of monitor i. It should be noted that is a discrete action, and its action list is ;

[0021] where represents the type of the mth monitoring device, and 0 indicates that this agent fails, that is, this monitor is not arranged in the site;

[0022] By taking 0 as an alternative, the algorithm can not only decide the position of the agent but also the number of agents, making the design results more flexible and diverse;

[0023] In addition, in order to improve the generalization ability of the agent neural network and accelerate the convergence speed, the action spaces of these four parameters are designed as continuous actions and scaled to .

[0024] Furthermore, the observation content of the monitoring-agent at least includes the following content:

[0025] The monitoring layout cannot exceed the boundary of the construction area. Therefore, each agent needs to observe the boundary of the construction area, that is , where represents the k-th boundary coordinate point of the construction area;

[0026] Similarly, each monitoring-agent needs to observe the area of the construction structure, that is

[0027]

[0028] where represents the x coordinate of the q-th boundary point of the m-th engineering structure;

[0029] Each monitoring-agent i needs to observe its own model, the observation range, and its planar position, height, and orientation angle, that is .

[0030] represents the angle that monitoring i can observe, which is an inherent attribute of monitoring i; represents the position of monitoring i in the plane; represents the height of monitoring i; represents the orientation of monitoring i;

[0031] Each monitoring-agent i needs to observe the positions and parameters of other agents, that is .

[0032] Furthermore, the reward of the monitoring-agent is the feedback of the environment to the monitoring layout;

[0033] The reward obtained by the monitoring-agent guides the agent to move to a better position and at the same time prompts the agent to select a more suitable device type for itself. In addition, since the goal of the monitoring-agent is a common goal and there is no competitive relationship, the global reward is used as the reward for each monitoring-agent;

[0034] Of course, the reward setting of the agent needs to consider multiple aspects, as shown in Equation (2), where represents the weights of different factors, which are determined by the designer or adjusted through experiments:

[0035]

[0036] where is the coverage reward of the monitoring, used to guide the agent to cover more areas, as shown in Equation (3):

[0037]

[0038] is the area covered by the monitoring, is the area of the site;

[0039] is the reward for monitoring equipment purchase, guiding the agent to use less equipment purchase cost, as shown in Equation (4):

[0040]

[0041] represents the purchase cost of the i-th device, and N represents the total number of all devices;

[0042] is the reward for monitoring installation and construction, guiding the agent to use less installation and construction cost, as shown in Equation (5):

[0043]

[0044] where N represents the total number of all devices, that is, the fewer the number of devices, the less the installation and construction cost;

[0045] represents the number of device types, that is, the fewer the device types, the lower the installation complexity, that is, the less the cost.

[0046] Furthermore, each monitoring-agent in the intelligent agent network architecture based on the attention mechanism has its own set of action (Actor) network and evaluation (Critic) network. In addition, from the observation content of the monitoring-agent, it can be seen that the observation content of the monitoring-agent is relatively large. Therefore, it is necessary for the monitoring-agent to pay more attention to important content. The network based on the attention mechanism is used as the action (Actor) network of the agent, and at the same time, it is also used as the evaluation (Critic) network of the agent. In the proposed network, Transformer is the final output network.

[0047] Furthermore, the query Q (Query), key K (Key), and value V (Value) in the Transformer network are the three core components in the attention mechanism, and the outputs of the LSTM network are respectively used as the inputs of the three core components;

[0048] The inputs of the three LSTMs are the observation content A, the observation content A, and the observation content B respectively;

[0049] The entire network architecture can be expressed as Equation (6):

[0050]

[0051] where: LSTMQ Represents the LSTM network that forms Q the value, d k is LSTM Q the vector dimension of the output of Obv A Then it represents the observation content A, with softmax as the activation function, indicating the ratio of the exponent of this element to the sum of the exponents of all elements. Similarly, LSTM K Represents the LSTM network that forms K the value, LSTM V Represents the LSTM network that forms V the value, Obv B Then it represents the observation content B;

[0052] In the action network, take its own parameters, namely as Obv B , and other parameters as Obv A ;

[0053] In the evaluation network, take the action taken a as Obv B , and other parameters as Obv A ;

[0054] The two networks respectively obtain actions and rewards through the attention mechanism represented by Equation (6).

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] 1. In the present invention, the monitoring layout is modeled as the design of a multi-agent reinforcement learning model. Through the deep reinforcement learning of multi-agents, guided by a common goal, each agent is prompted to act in a collaborative and adaptive manner, and finally the intelligent layout of the construction site monitoring is realized;

[0057] 2. The intelligent layout method for construction site monitoring based on deep reinforcement learning proposed by the present invention can efficiently and automatically design the monitoring layout. Compared with the traditional manual experience and fixed rule methods, it can significantly improve the layout efficiency and accuracy. By regarding the monitoring devices as agents for joint optimization, it can achieve global coordination and adaptive layout, avoiding the inflexibility and inefficiency problems brought by manual methods, and improving the monitoring layout efficiency and accuracy;

[0058] 3. The method proposed by the present invention can balance the number and layout of monitoring devices, thereby maximizing the monitoring coverage range while minimizing the equipment purchase, installation, and construction costs. Through deep reinforcement learning, the agent can automatically adjust the device location and quantity according to the dynamic conditions of the site, avoiding resource waste caused by over-deployment while ensuring that the monitoring requirements of the site are met.

[0059] 4. By introducing the combination of a discrete action space (including device type and location decision) and a continuous action space, this method enables each agent to not only select its best location but also dynamically decide whether to deploy monitoring devices, increasing the flexibility and diversity of the algorithm. This method solves the limitations brought by a fixed number of agents and endows the deployment plan with more adaptability.

[0060] 5. In the present invention, the Transformer network architecture is applied to the agent's action and evaluation (Critic) networks, which can process a large amount of observation data and accurately focus on important information through the attention mechanism. This design enhances the agent's information processing ability in complex environments, improves the decision-making quality and the generalization ability of the network, and ensures that the deep reinforcement learning algorithm can quickly adapt to the changing and complex construction environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0062] Figure 1 It is a schematic diagram of monitoring - multi-agent of the present invention;

[0063] Figure 2 It is a schematic diagram of the parameters of monitoring i of the present invention;

[0064] Figure 3 It is a schematic diagram of the network architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them.

[0066] In order to achieve comprehensive monitoring of the construction site, the monitors need to cooperate with each other to cover as comprehensively as possible. However, too many monitors will lead to an increase in costs. Therefore, it is necessary to arrange the monitors reasonably and efficiently. Each monitor is regarded as an agent, and each agent moves freely in the construction site, such asFigure 1 As shown. The optimal strategy of an agent depends on the strategies of other agents, which leads to significant strategy interaction and coordination problems in the learning process.

[0067] Furthermore, a construction site monitoring layout method based on deep reinforcement learning is proposed, which at least includes the following steps:

[0068] S1: Model the monitoring layout problem as a multi-agent reinforcement learning model. Each monitoring is regarded as an agent, that is, called a monitoring-agent. Each monitoring-agent moves freely in the construction site. The monitoring-agents cooperate with each other and coordinate with each other to achieve comprehensive monitoring coverage and cost minimization, and finally determine its own position;

[0069] S2: Build a monitoring-agent based on deep reinforcement learning. Through the guidance of deep reinforcement learning and common goals, promote each agent to act in a coordinated and adaptive manner, and finally realize the intelligent layout of construction site monitoring;

[0070] S3: Build an agent network architecture based on the attention mechanism.

[0071] The calculation process of monitoring layout modeling at least includes the following steps:

[0072] Convert the monitoring layout problem from a planar design problem into a path planning problem that changes with time and environment changes;

[0073] At this time, the mathematical model conforms to the mathematical modeling principle of multi-agent reinforcement learning;

[0074] As each monitoring-agent i takes an action according to the current state and the current time step t , the environment feedbacks a new state and a reward according to the current state and the action ;

[0075] For each agent, its goal is to maximize the expected cumulative reward, as shown in Equation (1):

[0076]

[0077] Among them, is the discount factor, taking 0.99, which is used to represent the current value of future rewards; N represents the number of monitoring-agents.

[0078] The deep reinforcement learning of the monitoring-agent at least includes the action of the monitoring-agent, the observation content of the monitoring-agent, and the reward of the monitoring-agent.

[0079] For the actions of the monitoring-agent, the action formats output by each agent are the same, that is , such as Figure 2 shown;

[0080] Among them, represents the type of the i-th monitoring, represents the position of the i-th monitoring in the plane, represents the height of the i-th monitoring, represents the orientation of the i-th monitoring. It should be noted that is a discrete action, and its action list is ;

[0081] Among them, represents the type of the m-th monitoring device, and 0 indicates that this agent fails, that is, this monitoring is not arranged in the site;

[0082] By taking 0 as an alternative solution, the problem that the number of agents is fixed when designing multi-agent deep reinforcement learning is solved, so that the algorithm can not only decide the positions of the agents but also the number of agents, making the design results more flexible and diverse;

[0083] In addition, in order to improve the generalization ability of the agent neural network and accelerate the convergence speed for monitoring, the action spaces of these four parameters are designed as continuous actions and scaled to .

[0084] The observation content of the monitoring-agent includes at least the following:

[0085] The monitoring arrangement cannot exceed the boundary of the construction area. Therefore, each agent needs to observe the boundary of the construction area, that is , among which, represents the k-th boundary coordinate point of the construction area;

[0086] Similarly, each monitoring-agent needs to observe the area of the construction structure, that is

[0087]

[0088] Among them, represents the x coordinate of the q-th boundary point of the m-th engineering structure;

[0089] Each monitoring-agent i needs to observe its own model, the observation range, as well as its planar position, height and orientation angle, that is ,

[0090] represents the angle that the i-th monitoring can observe and is an inherent attribute of the i-th monitoring; Indicates the position of monitor i in the plane; Indicates the height of monitor i; Indicates the orientation of monitor i;

[0091] Each monitoring-agent i needs to observe the positions and parameters of other agents, that is .

[0092] The reward of the monitoring-agent is the feedback of the environment to the monitoring layout;

[0093] The rewards obtained by the monitoring-agents guide the agents to move to a better position, and at the same time prompt the agents to select a more suitable device type for themselves. In addition, since the goals of the monitoring-agents are common goals and there is no competitive relationship, the global reward is used as the reward for each monitoring-agent;

[0094] Of course, the reward setting of the agents needs to consider multiple aspects. As shown in Equation (2), where Indicates the weights of different factors, which are determined by the designer or adjusted through experiments:

[0095]

[0096] Among them, Is the coverage reward of the monitor, used to guide the agent to cover more areas, as shown in Equation (3):

[0097] ,

[0098] Is the area covered by the monitor, Is the area of the site;

[0099] Is the equipment purchase reward of the monitor, guiding the agent to use less equipment purchase cost, as shown in Equation (4):

[0100] ,

[0101] Represents the purchase cost of the i-th device, and N represents the total number of all devices (i.e., the number of agents);

[0102] Is the installation and construction reward of the monitor, guiding the agent to use less installation and construction cost, as shown in Equation (5):

[0103]

[0104] Among them, N represents the total number of all devices (i.e., the number of monitoring-agents), that is, the fewer the number of devices, the less the installation and construction cost;

[0105] Indicates the number of types of devices, that is, the fewer the types of devices, the lower the installation complexity, that is, the lower the cost.

[0106] In the agent network architecture based on the attention mechanism, each monitoring-agent has its own set of action (Actor) network and evaluation (Critic) network. In addition, from the observation content of the monitoring-agent, it can be seen that the observation content of the monitoring-agent is relatively large. Therefore, it is necessary for the monitoring-agent to pay more attention to important content. The network based on the attention mechanism is used as the action (Actor) network of the agent, and at the same time, it is also used as the evaluation (Critic) network of the agent. In the proposed network, Transformer is the final output network.

[0107] Query in the Transformer network Q (Query), key K (Key), and value V (Value) are the three core components in the attention mechanism, and the outputs of the LSTM network are respectively used as the inputs of the three core components;

[0108] The inputs of the three LSTMs are observation content A, observation content A, and observation content B respectively;

[0109] The entire network architecture can be expressed as Equation (6):

[0110]

[0111] Among them: LSTM Q Indicates forming Q The LSTM network of the value, d k Is LSTM Q The vector dimension of the output of; Obv A Then it represents observation content A, softmax is the activation function, indicating the ratio of the exponent of this element to the sum of the exponents of all elements. Similarly, LSTM K Indicates forming K The LSTM network of the value, LSTM V Indicates forming V The LSTM network of the value, Obv B Then it represents observation content B;

[0112] In the action network, take its own parameters, that is, As ObvB , and other parameters are used as Obv A ;

[0113] In the evaluation network, the actions taken a are used as Obv B , and other parameters are used as Obv A ;

[0114] The two networks respectively obtain actions and rewards through the attention mechanism represented by Equation (6);

[0115] In addition, the multi-agent reinforcement learning of the present invention adopts a training method of centralized training and distributed execution, and at the same time adopts an experience replay mechanism for learning. This algorithm framework can effectively handle learning tasks in large-scale and highly interactive environments.

[0116] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. A construction site monitoring arrangement method based on deep reinforcement learning, characterized in that: At least the following steps are included: S1: The monitoring layout problem is modeled as a multi-agent reinforcement learning model. Each monitor is regarded as an agent, that is, a monitoring-agent. Each monitoring-agent is free and automatic in the construction site. The monitoring-agents cooperate with each other to achieve comprehensive monitoring coverage and minimize costs, and finally determine their own positions; The calculation process for monitoring arrangement modeling includes at least the following steps: Transform the monitoring layout problem from a plane design problem to a path planning problem with time and environment changes; At this point, the mathematical model conforms to the mathematical modeling principles of multi-agent reinforcement learning; As each monitoring-agent i is based on the current state Perform an action with the current time step t , the environment is based on the current state and actions Feedback a new status and rewards ; For each agent, its goal is to maximize the expected cumulative reward, as shown in formula (1): , in, is the discount factor, which is 0.99 and is used to represent the current value of future rewards; N represents the number of monitoring-agents; S2: Build a monitoring-agent based on deep reinforcement learning. Through deep reinforcement learning and the guidance of common goals, each agent is prompted to act in a coordinated and adaptive manner, and finally realize the intelligent arrangement of construction site monitoring; The deep reinforcement learning of the monitoring-agent includes at least the actions of the monitoring-agent, the observations of the monitoring-agent and the rewards of the monitoring-agent; The action format output by each agent in the monitoring-agent action is the same, that is, ; in, Indicates the type of monitoring i, Indicates the position of monitoring i in the plane, Indicates the height of monitoring i, Indicates the direction of monitoring i. It should be noted that is a discrete action, and its action list is ; in, Indicates the type of the mth monitoring device, and 0 means that this agent is invalid, that is, this monitoring device is not deployed in the venue; By taking 0 as an alternative, the algorithm can not only decide the position of the agent but also the number of agents, making the design results more flexible and diverse; In addition, in order to monitor and improve the generalization ability of the agent neural network and accelerate the convergence speed, The action space of these four parameters is designed as a continuous action and is scaled to ; The monitoring-agent's observation content includes at least the following: The monitoring layout cannot exceed the boundary of the construction area, so each agent needs to observe the boundary of the construction area, that is, ,in, Indicates the kth construction area boundary coordinate point; Each monitoring-agent needs to observe the area of ​​the construction structure, i.e. , in, represents the x-coordinate of the qth boundary point of the mth engineering structure; Each monitoring-agent i needs to observe its own model, observation range, plane position, height and orientation angle, that is, ; It represents the angle that monitor i can observe, which is an inherent attribute of monitor i; Indicates the position of monitoring i in the plane; represents the height of monitoring i; represents the direction of monitoring i; Each monitoring agent i needs to observe the positions and parameters of other agents, that is, ; The reward of the monitoring-agent is the feedback of the environment to the monitoring arrangement; The rewards obtained by the monitoring-agent guide the agent to move to a better position and encourage the agent to choose a more suitable device type. In addition, since the goals of the monitoring-agents are common goals and there is no competition between them, the global reward is used as the reward for each monitoring-agent. Of course, the reward setting of the intelligent agent needs to take into account many aspects, as shown in formula (2), where: and Indicates the weights of different factors, determined by the designer or through experimental adjustment: , in, is the monitoring coverage reward, which is used to guide the agent to cover more areas, as shown in formula (3): , is the area covered by the monitoring, is the area of ​​the site; is the monitored equipment purchase reward, guiding the agent to use less equipment purchase cost, as shown in formula (4): , represents the purchase cost of the i-th device, and N represents the number of all devices; is the installation and construction reward for monitoring, guiding the agent to use less installation and construction costs, as shown in formula (5): , Where N represents the number of all equipment, that is, the fewer the number of equipment, the lower the installation and construction costs; Indicates the number of types of equipment, that is, the fewer the types of equipment, the less complicated the installation, that is, the lower the cost; S3: Build an attention-based intelligent network architecture.

2. A construction site monitoring arrangement method based on deep reinforcement learning according to claim 1, characterized in that: In the attention-based agent network architecture, each monitoring agent has its own action network and evaluation network. The attention-based network is used as the action network of the agent, and it is also used as the evaluation network of the agent. In the proposed network, Transformer is the final output network.

3. A construction site monitoring arrangement method based on deep reinforcement learning according to claim 2, characterized in that: The Transformer network queries Q ,key K Sum V These are the three core components in the attention mechanism, and the output of the LSTM network is used as the input of the three core components respectively; The inputs of the three LSTMs are observation content A, observation content A, and observation content B; The entire network architecture is expressed as formula (6): , in: LSTM Q Indicates the formation Q LSTM network of values, d k yes LSTM Q The vector dimensions of the output of Obv A It represents the observed content A, and softmax is the activation function, which represents the ratio of the index of an element to the sum of the indexes of all elements. LSTM K Indicates the formation K LSTM network of values, LSTM V Indicates the formation V LSTM network of values, Obv B It indicates the observation content B; In the action network, the parameters of the As Obv B , other parameters as Obv A ; In the evaluation network, the action to be taken a As Obv B , other parameters as Obv A ; The two networks obtain actions and rewards respectively through the attention mechanism expressed in formula (6).

Citation Information

Patent Citations

  • Unmanned cluster task collaboration method based on multi-agent reinforcement learning

    CN113589842A

  • Video abstract generation method based on multi-agent reinforcement learning

    CN115982407A