Building two-dimensional bidirectional module layout generation system based on artificial intelligence reinforcement learning

Through a two-dimensional, bidirectional module layout generation system based on artificial intelligence reinforcement learning, the problems of modular building design complexity and insufficient apartment diversity are solved, flexible horizontal and vertical arrangement of modules is achieved, design efficiency and quality are improved, and diverse and personalized apartment generation solutions are provided.

CN119848989BActive Publication Date: 2025-09-26CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411916344.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-09-26
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing modular buildings have problems with increased design complexity and insufficient diversity when generating apartment types. The existing module arrangement method is usually one-dimensional and single-directional, with obvious limitations, and it is difficult to meet building quality, cost and personalization requirements.

Method used

A two-dimensional bidirectional module layout generation system based on artificial intelligence reinforcement learning is adopted. Through the algorithm input unit, algorithm framework building unit, algorithm initialization unit, algorithm training unit and solution evaluation unit, a reinforcement learning framework is constructed, allowing modules to be flexibly arranged horizontally and vertically. Using multi-dimensional evaluation analysis and visualization functions, a variety of module arrangement combination schemes are output for designers to choose from.

Benefits of technology

It simplifies the design process, improves design efficiency and quality, realizes the diversity of modular buildings and the generation of personalized apartment types, has automation and rapid response capabilities, and enhances the practicality of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848989B_ABST
    Figure CN119848989B_ABST
Patent Text Reader

Abstract

The present invention relates to a two-dimensional bidirectional module layout generation system for buildings based on artificial intelligence reinforcement learning, which belongs to the field of building technology. The system automatically processes input information, including the area to be arranged, the type of module and the module parameters, and uses the prior knowledge of modular building arrangement rules to construct a reinforcement learning framework. Through training, the model outputs a variety of module arrangement combination results calculated by the reinforcement learning model. The system adopts a two-dimensional layout method, and judges whether the arrangement is completed by the remaining area, instead of limiting the width of the module to only rely on the length to judge whether the arrangement is completed. The length and width of the module are determined only by the input parameters of the algorithm, and the module can be flexibly arranged in the horizontal and vertical directions. It simplifies manual operation and also has the advantages of automation, rapid response and multi-scheme generation. The multi-dimensional evaluation analysis and visualization functions further enhance the practicality of the method, making it show broad application prospects in the field of architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of building technology and relates to a building two-dimensional bidirectional module layout generation system based on artificial intelligence reinforcement learning. Background Art

[0002] In recent years, with the rapid growth of the global population and the rapid advancement of urbanization, the construction industry is facing unprecedented challenges. Traditional construction methods can no longer meet the comprehensive demands of modern society for speed, quality, cost, and environmental sustainability. Against this backdrop, prefabricated modular construction, as an innovative construction method, has gradually attracted widespread attention and favor.

[0003] Modular construction shifts the majority of construction work to prefabrication in factories, significantly improving the efficiency and quality of the building process. This approach not only shortens the construction cycle and reduces on-site waste, but also offers the advantages of better quality control, disassembly, and reusability. Furthermore, modular construction can effectively reduce accidents and improve worker safety in a factory-based environment. These advantages have led to its widespread adoption in a variety of settings, including schools, hotels, dormitories, and offices.

[0004] However, modular construction also faces some challenges. For example, traditional construction methods offer greater flexibility in unit design. However, modular construction, due to the introduction of prefabricated modules, requires more consideration from the early stages of the design process, such as module size, shape, and connection methods. This increases design complexity, making errors more likely and ultimately increasing project costs.

[0005] On the other hand, modular building unit generation also needs to consider the issue of diversity. Because modular building prefabricated modules have certain standards and repeatability, how to achieve diversity and personalization of unit types while ensuring building quality is a difficult problem in the promotion of modular building.

[0006] To address these issues, researchers have begun exploring the use of intelligent algorithms and AI-generated content to automate the generation of modular building layouts. Reinforcement learning, an advanced machine learning method, continuously learns and optimizes decision-making strategies through interaction with the environment. Applying reinforcement learning to modular building layout generation can automatically explore various feasible layout design options based on different needs and constraints. Through continuous learning and optimization, it finds the optimal solution that meets multiple requirements, including functionality, structure, cost, user personalization, and environmental sustainability.

[0007] Currently, some methods for generating building module layouts based on reinforcement learning have been proposed in the existing technology. However, when setting the module width, the existing module arrangement methods usually set it to be the same as the area to be arranged. Only the relationship between the module length and the length of the area to be arranged is considered to determine whether the arrangement is complete. This method is a one-dimensional and unidirectional arrangement with obvious limitations. Summary of the Invention

[0008] In light of this, the present invention aims to provide a two-dimensional, bidirectional modular building layout generation system based on artificial intelligence (AI) reinforcement learning. This system automatically processes input information, including the area to be arranged, module type, and module parameters, and utilizes prior knowledge of modular building layout rules to construct a reinforcement learning framework. Through training, the model can output a variety of module layout combinations calculated by the reinforcement learning model for designers to choose from. This system utilizes a two-dimensional layout approach, determining whether the layout is complete based on the remaining area, rather than limiting the module width to length. The module's length and width are determined solely by the algorithm's input parameters, allowing for flexible module layout in both the horizontal and vertical directions. This innovation not only simplifies manual operation but also offers advantages such as automation, rapid response, and the generation of multiple solutions. Furthermore, multi-dimensional evaluation, analysis, and visualization capabilities further enhance the method's practicality, promising broad application prospects in the architectural field. This advanced approach enables designers to more efficiently design buildings, improve work efficiency, and optimize design quality.

[0009] In order to achieve the above object, the present invention provides the following technical solutions:

[0010] A building two-dimensional bidirectional module layout generation system based on artificial intelligence reinforcement learning, the system comprising: an algorithm input unit, an algorithm framework building unit, an algorithm initialization unit, an algorithm training unit, a scheme evaluation unit, and an algorithm output unit;

[0011] The algorithm input unit is used to input the area to be arranged, module size, and algorithm parameters; the area to be arranged includes the building outline and room boundaries; the module size includes the minimum length, width, and maximum length and width of the input module; the algorithm parameters include different generation targets, namely the total number of modules, module types, and module alignment weights α1, α2, and α3, which must satisfy α1+α2+α3=1;

[0012] The algorithm framework building unit is used to build a reinforcement learning framework;

[0013] The algorithm initialization unit is used to initialize the reinforcement learning environment and the intelligent agent;

[0014] The algorithm training unit is used to perform agent training in a cyclic and iterative manner and output a final solution;

[0015] The scheme evaluation unit is used to perform quantitative evaluation of the results and visual description of indicators, including quantitative statistical evaluation of the number of module types, total number of modules, and room alignment of the generated scheme;

[0016] The algorithm output unit is used to output module arrangement scheme comparison analysis, scheme floor plan and scheme data table.

[0017] Furthermore, the algorithm input unit receives three input conditions: 1) a JSON description file of the area to be arranged; 2) a set of optional modules; 3) algorithm parameters, including the weights of the three arrangement goals (total number of modules, module types, and module alignment) on the results, as well as the exploration decay rate parameter of reinforcement learning. Specifically, the following steps are included:

[0018] S1. Processing the JSON description file of the area to be arranged: Read the input JSON description file of the area to be arranged, obtain the coordinates of the boundary of each area, calculate the coordinates of each area through the coordinates of the boundary, and form a dot matrix of the entire area to be arranged;

[0019] S2, optional module set processing: Generate different modules by inputting the maximum and minimum length and width constraints and generation strategy, and access the specific module size through the module sequence number;

[0020] S3. Algorithm parameter processing: Obtain the weights α1, α2, and α3 of the total number of modules, module types, and module alignment, which must satisfy α1+α2+α3=1; obtain the exploration rate parameter ε decay .

[0021] Furthermore, the algorithm framework building unit builds a reinforcement learning algorithm framework, including the definition and setting of the environment, state, action, and reward mechanism; the environment is defined as a layout area, which is composed of the boundary of the layout area and the boundary of the internal functional area; a two-dimensional layout method is adopted, and the state setting cannot only set the distance to the boundary in a certain direction; the intelligent agent state is the dot matrix information of the entire layout area, in which the arranged areas are marked by the modules as their type numbers, and the unarranged areas are marked as -1; whether the area has been arranged with modules is judged by whether the coordinate value in the dot matrix is ​​-1, and whether the area to be arranged is completed is judged by whether there is still an area with a value of -1; the intelligent agent action obtains a list of optional module size types by the length and width range parameters of the module in the input parameters and the generation strategy, and then accesses and selects different actions through the type number; the reward mechanism provides feedback on the actions performed by the intelligent agent from three perspectives: the total number of modules, the type of modules, and the degree of module alignment.

[0022] Furthermore, the algorithm framework construction specifically includes the following steps:

[0023] S1. Build environment: Count the boundaries of each room and functional area, and store each boundary in a Json file in the form of starting and ending points;

[0024] S2. Constructing the state space: Construct the state space of the agent to obtain the state of the agent in the environment at any time. The modules are arranged starting from the upper left corner of the area, and the position of the next arrangement is randomly selected from the four vertices of the module in the previous arrangement. The dot matrix of the entire room is used as the state space. Whether the coordinate value of the area is -1 indicates which areas are not arranged and which areas are arranged. Whether the arrangement is complete is determined by whether all the areas to be arranged are not -1.

[0025] S3. Constructing action space: The agent action space is determined by the set of module sizes;

[0026] S4. Build a reward mechanism: Build a reward function to obtain the reward that the environment gives back to the agent after executing an action, that is, selecting a module of a certain size and placing it into the environment.

[0027] Furthermore, in step S4, the action execution is evaluated from three aspects based on three objectives: 1) the total number of modules; 2) the type of modules; and 3) the degree of module alignment. The specific evaluation methods include:

[0028] S41. After placing the module, calculate the positional relationship between the placed module and the layout area. There are three possible positional relationships: ① The placed module overlaps with the previously placed module; ② The size of the placed module exceeds the room boundary; ③ The placed module neither overlaps with the layout area nor exceeds the room boundary. By examining the relationship, a reward r describing the total number of modules and the type of modules is obtained from a qualitative and quantitative perspective. 1、 r2; Because the total number of modules is described by the number of times the action is executed, executing any action indicates an increase in the total number of modules, and judging whether the module type appears indicates an increase in the type;

[0029] S42. After the module is successfully placed, the degree of alignment between the modules in the entire room is determined, that is, whether there are many misaligned module vertices. The more misaligned module vertices there are, the lower the module alignment is considered to be, and a reward r3 for the module alignment degree is obtained;

[0030] S43, the sign of training termination: the remaining area of ​​the room cannot accommodate any module, and feedback r4 is given by the inventory of the remaining space size; r1, r2, and r3 are adjusted according to the input weights α1, α2, and α3. Through different feedback signals, after training, a batch of results that meet the input goals are obtained.

[0031] Furthermore, the algorithm initialization unit initialization step specifically includes:

[0032] S1. Initialize the environment, including: initializing each coordinate value of the dot matrix to -1; returning the initial position of the agent to the upper left corner of the room; clearing the historical action list;

[0033] S2. Initialize the neural network and use the deep reinforcement learning method. The agent makes decisions and evaluates actions through the neural network. The neural network includes the action evaluation network N eval and the target network N target , the initialization of the neural network includes: randomly initializing the action evaluation network N eval The weight of N target The weights are set to be equal to N eval same;

[0034] S3. Initialize the experience buffer: Using deep reinforcement learning, the agent trains the neural network by sampling the existing data in the experience buffer. Positive and negative samples are stored in different experience replay buffers. Initializing the experience replay buffer includes clearing both buffers.

[0035] S4. Initialize the parameters of the deep reinforcement learning agent, including:

[0036] S41, initialization exploration rate ε, exploration decay rate ε decay The agent generates a random number to determine whether to randomly select an action or obtain an action through a neural network. If the random number is less than ε, the action is randomly selected, otherwise the current state is input into the neural network to obtain the output action. The initial ε is set to a large value, and the agent mainly performs actions randomly. After each action is determined, ε is adjusted according to ε. decay Decay, subsequent agents will tend to use neural networks to get actions;

[0037] S42. Initialize the discount factor γ, which is used to calculate the time difference target value during training;

[0038] S43, initialize the training batch size batch_size;

[0039] S44. Initialize the target network parameter update frequency update_freq.

[0040] Furthermore, the algorithm training unit performs training by adopting the following steps:

[0041] S1. Agent selection action: At any time in state S, the agent decides to randomly select a module or select the action A with the largest Q value based on the current state S according to the environment exploration rate ε; at the same time, the exploration rate ε is determined by the exploration decay rate ε. decay reduce;

[0042] S2. Feedback from the environment: Based on the actions performed by the agent, i.e. the modules selected, the environment provides the agent with corresponding reward values ​​according to the reward mechanism. Specifically, the reward values ​​include:

[0043] S21. Record the current agent position, find the corresponding module size according to the selected action, and update the agent position;

[0044] S22. Give a reward R according to the reward mechanism;

[0045] S23. Determine the placement of the current module. If the length or width of the module plus the length or width of the agent exceeds the range of the room, the action is invalid, the module sequence is not updated, the current state Q value is updated, and the training round end flag T is set to true; if there is another module at the location where the current module is placed, the action is invalid, the module sequence is not updated, the current state Q value is updated, and the training round end flag T is set to true; if the location where the current module is placed does not exceed the range of the room and there is no other module, the action is valid, the module sequence is updated, the Q value is updated, and the position of the agent is updated;

[0046] S24. Obtain a new state S' according to the updated pointer;

[0047] S25. Depending on whether the action is valid, store the five-tuple (S, A, R, S', T) into the positive sample or negative sample experience buffer;

[0048] S3. Update the neural network: This step samples data from the experience buffer, trains the network, and completes the weight update. First, check whether the size of the positive and negative sample experience buffers is not less than batch_size / 2. If so, do not update the network; otherwise, update the network. The main process is as follows:

[0049] S31. Sample batch_size / 2 from positive and negative samples respectively, combine them into a batch of data, and input N eval , get the estimated output yhat;

[0050] S32, using the target network output to calculate the time difference target value y;

[0051] S33. Calculate the mean square error loss of yhat and y, and update N with gradient descent eval ;

[0052] S34. If the number of training times meets update_freq, N eval The weight value is copied to N target ;

[0053] S4. Terminate training and output results: Perform training according to the set number of training rounds. During the training process, store the successful arrangement results and record the result information. After the training is completed, output the Pareto image of all results, and select a series of corresponding Pareto frontier solutions as the result output based on multiple objectives. Output a variety of arrangement methods for designers to screen.

[0054] Furthermore, the scheme evaluation unit performs a quantitative evaluation of the scheme including the following steps:

[0055] S1. Count the total number of modules and calculate the length of the execution action sequence of each area to be arranged to obtain the total number of modules;

[0056] S2. Count the number of module types, calculate the number of modules, and store them in a dictionary form;

[0057] S3. Count the module alignment, calculate the number of isolated vertices of adjacent modules in each arrangement area, and determine whether the length and width of adjacent modules are the same;

[0058] S4. Output the evaluation results.

[0059] Furthermore, the algorithm output unit algorithm output includes: 1) a Pareto chart of all collected arrangement schemes; 2) a display of the Pareto frontier solutions among multiple targets; 3) a data table of the arrangement schemes of each module.

[0060] The beneficial effects of the present invention are:

[0061] This invention utilizes a two-dimensional layout approach, using the remaining area to determine whether the layout is complete, rather than limiting the width of the module to length alone. The module's length and width are determined solely by the algorithm's input parameters, allowing for flexible horizontal and vertical layouts. This innovation not only simplifies manual operations but also offers advantages such as automation, rapid response, and the ability to generate multiple solutions. Furthermore, multi-dimensional evaluation, analysis, and visualization further enhance the method's practicality, giving it broad application prospects in the architectural field. This advanced approach allows designers to design buildings more efficiently, improving work efficiency and optimizing design quality.

[0062] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0064] Figure 1 This is a flow chart of the overall operation of the present invention;

[0065] Figure 2 Schematic diagram of the area to be arranged;

[0066] Figure 3 It is the module size range or type and fixed size diagram;

[0067] Figure 4 This is the framework diagram for intelligent agent training;

[0068] Figure 5 Figure 2 is a diagram of the agent training process;

[0069] Figure 6 Provide quantitative evaluation diagram for the scheme;

[0070] Figure 7 This is the floor plan of the module layout. DETAILED DESCRIPTION

[0071] The technical solution of the present invention is described in detail below with reference to the accompanying drawings.

[0072] This invention provides a system for generating two-dimensional, bidirectional modular building layouts based on artificial intelligence (AI) reinforcement learning. The system automatically processes input information, including the area to be arranged, module type, and module parameters, and builds a reinforcement learning framework using prior knowledge of modular building layout rules. Through training, the model can output a variety of modular layout combinations calculated by the reinforcement learning model for designers to choose from. Figure 1 This is the overall operation flow chart of the present invention, and the system includes: an algorithm input unit, an algorithm framework building unit, an algorithm initialization unit, an algorithm training unit, a scheme evaluation unit and an algorithm output unit; the algorithm input unit is used to input the area to be arranged, module size, and algorithm parameters; the area to be arranged includes the building outline and the room boundary; the module size includes the minimum length, width and maximum length and width of the input module; the algorithm parameters include different generation targets, namely the weights of the total number of modules, module types, and module alignment degree; the algorithm framework building unit is used to construct a reinforcement learning framework; the algorithm initialization unit is used to initialize the reinforcement learning environment and the intelligent agent; the algorithm training unit is used to iteratively perform intelligent agent training and output the final scheme; the scheme evaluation unit is used to perform quantitative evaluation of the results and visual description of indicators, including quantitative statistical evaluation of the generated scheme in terms of the number of module types, total number of modules, and room alignment degree; the algorithm output unit is used to output module arrangement scheme comparative analysis, scheme floor plan and scheme data table.

[0073] like Figure 2 and Figure 3 As shown, the algorithm input unit inputs the area to be arranged, module size, and algorithm parameters; the area to be arranged includes the building outline and room boundary; the module size includes the minimum length, width, and maximum length and width of the input module, such as length 500-1000mm and width 400-800mm.

[0074] The algorithm input unit accepts the following inputs: 1) a JSON file describing the area to be arranged; 2) the module length and width parameters and generation strategy; 3) algorithm parameters, including the weights of the three arrangement objectives (total number of modules, module type, and module alignment) on the results, as well as the reinforcement learning exploration decay parameter. After processing, the resulting output includes: 1) a dot matrix of the area to be arranged; 2) a list of available module sizes; 3) the weights of each arrangement objective and the algorithm exploration decay rate, which serve as input to the next unit.

[0075] The algorithm framework construction unit is used to build the reinforcement learning algorithm framework, including the definition and configuration of the environment, state, actions, and reward mechanism. The environment is defined as a layout area, consisting of the layout area boundary and the boundaries of the functional areas within it. Unlike existing agent state definitions, our approach targets two-dimensional layouts, so the state setting cannot simply be the distance to the boundary in a specific direction. In our approach, the agent state is a dot matrix diagram of the entire layout area, where already arranged areas are marked by their type numbers, and unarranged areas are marked with -1. Whether a coordinate value in the dot matrix is ​​-1 determines whether a module has already been arranged in that area, and whether any remaining areas are -1 determines whether the area to be arranged has been arranged. The agent's actions are derived from the module length and width parameters in the input parameters and the generation strategy to obtain a list of available module size types. The type number is then used to access and select different actions. The reward mechanism provides feedback on the agent's actions based on the total number of modules, module type, and module alignment. Figure 4 Figure 2 is a diagram of the agent training framework.

[0076] Specifically, the algorithm input unit receives three input conditions: 1) a JSON description file of the area to be arranged; 2) a set of optional modules; 3) algorithm parameters, including the weights of the three arrangement goals (total number of modules, module types, and module alignment) on the results, as well as the exploration decay rate parameter of reinforcement learning. The specific processing process is as follows:

[0077] Processing of the JSON description file of the area to be arranged: Read the input JSON description file of the area to be arranged, obtain the coordinates of the boundary of each area, calculate the coordinates of each area through the coordinates of the boundary, and form a dot matrix of the entire area to be arranged.

[0078] Optional module set processing: Generate different modules by inputting maximum and minimum length and width constraints and generation strategies, and access specific module sizes through module serial numbers.

[0079] Algorithm parameter processing: Get the weights α1, α2, and α3 of the total number of modules, module types, and module alignment, which must satisfy α1+α2+α3=1; get the exploration rate parameter ε decay .

[0080] The algorithm framework is built, including the following steps:

[0081] Step 1. Build the environment

[0082] Count the boundaries of each room and functional area, and store each boundary in the form of starting point and end point in a Json file.

[0083] Step 2. Construct the state space

[0084] Construct the agent's state space to capture its state in the environment at any given moment. Modules are arranged starting from the top left corner of the area, with the next position randomly selected from the four vertices of the previously arranged module. A dot matrix of the entire room is used as the state space. Areas that are unarranged and arranged are indicated by whether their coordinate values ​​are -1. Arrangement is determined by whether all areas to be arranged are not -1.

[0085] Step 3. Construct the action space

[0086] The agent action space is determined by the set of module sizes.

[0087] Step 4. Build a reward mechanism

[0088] Construct a reward function to obtain the reward that the environment gives back to the agent after executing an action, i.e., selecting a module of a certain size and placing it in the environment. Based on three objectives, the action execution is evaluated from three aspects: (1) the total number of modules; (2) the type of modules; and (3) the degree of module alignment. The specific evaluation ideas are:

[0089] (1) After placing the module, calculate the positional relationship between the placed module and the layout area. There are three possible positional relationships: ① The placed module overlaps with the previously placed module; ② The size of the placed module exceeds the room boundary; ③ The placed module neither overlaps with the layout area nor exceeds the room boundary. By examining the relationship, a reward r describing the total number of modules and the type of modules is obtained from a qualitative and quantitative perspective. 1、 Since the total number of modules is described by the number of times actions are executed, executing any action indicates an increase in the total number of modules, and judging whether a module type appears indicates an increase in its type.

[0090] (2) After the module is successfully placed, the degree of alignment between the modules in the entire room is judged, that is, whether there are many misaligned module vertices. The more misaligned module vertices there are, the lower the module alignment degree is considered, and the module alignment degree reward r3 is obtained.

[0091] (3) The training ends when the remaining area of ​​the room cannot accommodate any module. Feedback r4 is given on the remaining space size.

[0092] According to the input weights α1, α2, α3, r1, r2, and r3 are adjusted. Through different feedback signals and training, a batch of results that meet the input targets are obtained.

[0093] The steps performed by the algorithm initialization unit include:

[0094] Step 1. Initialize the environment

[0095] The initialization of the environment includes:

[0096] (1) Initialize each coordinate value of the dot matrix to -1.

[0097] (2) The initial position of the agent returns to the upper left corner of the room.

[0098] (3) Clear the history action list.

[0099] Step 2. Initialize the neural network

[0100] Using deep reinforcement learning methods, the agent makes decisions and evaluates actions through a neural network. The neural network includes an action evaluation network N eval and the target network N target The initialization of the neural network includes:

[0101] (1) Randomly initialize the action evaluation network N eval The weight of .

[0102] (2) N target The weights are set to be equal to N eval same.

[0103] Step 3. Initialize the experience buffer

[0104] Using deep reinforcement learning, the agent trains its neural network by sampling data from its experience buffer. Positive and negative samples are stored in separate experience replay buffers. Initializing the experience replay buffers involves clearing both buffers.

[0105] Step 4. Initialize agent parameters

[0106] Initialize the deep reinforcement learning agent parameters, including:

[0107] (1) Initialize the exploration rate ε and the exploration decay rate ε decay The agent generates a random number to determine whether to randomly select an action or obtain an action through the neural network. If the random number is less than ε, the action is randomly selected, otherwise the current state is input into the neural network to obtain the output action. The initial ε is set to a large value, and the agent mainly performs actions randomly. After each action is determined, ε is calculated based on ε. decay decay, subsequent agents will tend to use neural networks to get actions.

[0108] (2) Initialize the discount factor γ. The discount factor is used to calculate the temporal difference target value during training.

[0109] (3) Initialize the training batch size batch_size.

[0110] Initialize the target network parameter update frequency update_freq.

[0111] Figure 5 This is a diagram of the agent training process. The execution steps of the algorithm training unit include:

[0112] Step 1: The agent chooses an action

[0113] At any moment in state S, the agent decides to randomly select a module or select the action A with the largest Q value based on the current state S according to the environment exploration rate ε. At the same time, the exploration rate ε is determined by the exploration decay rate ε. decay Decrease.

[0114] Step 2: Feedback from the environment

[0115] Based on the actions performed by the agent, that is, the modules selected, the environment will give the agent corresponding reward values ​​based on the reward mechanism. The specific process is as follows:

[0116] (1) Record the current agent position, find the corresponding module size according to the selected action, and update the agent's position.

[0117] (2) Give reward R according to the reward mechanism.

[0118] (3) Determine the placement of the current module. If the length or width of the module plus the length or width of the agent exceeds the range of the room, the action is invalid, the module sequence is not updated, the current state Q value is updated, and the training end flag T for this round is set to true. If there are other modules at the location where the current module is placed, the action is invalid, the module sequence is not updated, the current state Q value is updated, and the training end flag T for this round is set to true. If the location where the current module is placed does not exceed the range of the room and there are no other modules, the action is valid, the module sequence is updated, the Q value is updated, and the position of the agent is updated.

[0119] (4) Obtain the new state S' according to the updated pointer.

[0120] (5) Depending on whether the action is valid, store the five-tuple (S, A, R, S', T) into the positive sample or negative sample experience buffer.

[0121] Step 3: Update the neural network

[0122] This step samples data from the experience buffer, trains the network, and completes the weight update. First, check whether the size of the positive and negative sample experience buffers is not less than batch_size / 2. If so, do not update the network; otherwise, update the network. The main process is as follows:

[0123] (1) Sample batch_size / 2 from positive and negative samples respectively, combine them into a batch of data, and input N eval , and obtain the estimated output yhat.

[0124] (2) Use the target network output to calculate the time difference target value y.

[0125] (3) Calculate the mean square error loss between yhat and y, and update N with gradient descent eval .

[0126] (4) If the number of training times meets update_freq, N eval The weight value is copied to N target .

[0127] Step 4: Terminate training and output results

[0128] Training is performed for a set number of rounds, storing successful placement results and recording the resulting information. After training, a Pareto graph of all results is output, and a series of corresponding Pareto frontier solutions are selected based on multiple objectives as the output. This output provides a variety of placement options for designers to select.

[0129] Figure 6 The solution evaluation unit generates a quantitative evaluation diagram for each solution. The steps involved are as follows: Step 1: Counting the total number of modules: Calculate the length of the execution action sequence for each area to be arranged to determine the total number of modules. Step 2: Counting the number of module types: Calculate the number of modules and store them in a dictionary. Step 3: Counting the degree of module alignment: Calculate the number of isolated vertices between adjacent modules in each arrangement area and determine whether the length and width of adjacent modules are the same. Step 4: Output the evaluation results.

[0130] The algorithm output includes: (1) a Pareto chart of all collected layout plans; (2) a display of the Pareto frontier solutions for multiple targets; and (3) a data table of the layout plans for each module. Figure 7 This is the floor plan of the module layout.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A building two-dimensional bidirectional module layout generation system based on artificial intelligence reinforcement learning, characterized by: The system includes: an algorithm input unit, an algorithm framework building unit, an algorithm initialization unit, an algorithm training unit, a solution evaluation unit and an algorithm output unit; The algorithm input unit is used to input the area to be arranged, module size, and algorithm parameters; the area to be arranged includes the building outline and room boundaries; the module size includes the minimum length, width, and maximum length and width of the input module; the algorithm parameters include different generation goals, namely the total number of modules, module types, and module alignment weights α1, α2, and α3, which must satisfy α1 + α2 + α3 = 1; The algorithm framework building unit is used to build a reinforcement learning framework; The algorithm initialization unit is used to initialize the reinforcement learning environment and the intelligent agent; The algorithm training unit is used to perform agent training in a cyclic and iterative manner and output a final solution; The scheme evaluation unit is used to perform quantitative evaluation of the results and visual description of indicators, including quantitative statistical evaluation of the number of module types, total number of modules, and room alignment of the generated scheme; The algorithm output unit is used to output module arrangement scheme comparison analysis, scheme floor plan and scheme data table; The algorithm input unit receives three input conditions: 1) a JSON description file of the area to be arranged; 2) a set of optional modules; 3) algorithm parameters, including the weights of the three arrangement goals (total number of modules, module types, and module alignment) on the results, as well as the exploration decay rate parameter of reinforcement learning. The specific steps include: S1. Processing the JSON description file of the area to be arranged: Read the input JSON description file of the area to be arranged, obtain the coordinates of the boundary of each area, calculate the coordinates of each area through the coordinates of the boundary, and form a dot matrix of the entire area to be arranged; S2, optional module set processing: Generate different modules by inputting the maximum and minimum length and width constraints and generation strategy, and access the specific module size through the module sequence number; S3, algorithm parameter processing: obtain the weights α1, α2, α3 of the total number of modules, module types, and module alignment degree, which must satisfy α1 + α2 + α3 = 1; obtain the exploration rate parameter ε decay ; The algorithm framework building unit builds a reinforcement learning algorithm framework, including the definition and setting of the environment, state, action, and reward mechanism; the environment is defined as a layout area, which is composed of the boundary of the layout area and the boundary of the internal functional area; a two-dimensional layout method is adopted, and the state setting cannot only set the distance to the boundary in a certain direction; the intelligent agent state is the dot matrix information of the entire layout area, in which the arranged areas are marked by the modules as their type numbers, and the unarranged areas are marked as -1; whether the area has been arranged with modules is judged by whether the coordinate value in the dot matrix is ​​-1, and whether the area to be arranged is completed is judged by whether there is still an area with -1; the intelligent agent action obtains a list of optional module size types by the length and width range parameters of the module in the input parameters and the generation strategy, and then accesses and selects different actions through the type number; the reward mechanism provides feedback on the actions performed by the intelligent agent from three perspectives: the total number of modules, the type of modules, and the degree of module alignment.

2. The system for generating a two-dimensional bidirectional building module layout based on artificial intelligence reinforcement learning according to claim 1, characterized in that: The algorithm framework construction specifically includes the following steps: S1. Build the environment: Count the boundaries of each room and functional area, and store each boundary in a JSON file as a start point and end point. S2. Constructing the state space: Construct the state space of the agent to obtain the state of the agent in the environment at any time. The modules are arranged starting from the upper left corner of the area, and the position of the next arrangement is randomly selected from the four vertices of the module in the previous arrangement. The dot matrix of the entire room is used as the state space. Whether the coordinate value of the area is -1 indicates which areas are not arranged and which areas are arranged. Whether the arrangement is complete is determined by whether all the areas to be arranged are not -1. S3. Constructing action space: The agent action space is determined by the set of module sizes; S4. Build a reward mechanism: Build a reward function to obtain the reward that the environment gives back to the agent after executing an action, that is, selecting a module of a certain size and placing it into the environment.

3. The system for generating a two-dimensional bidirectional building module layout based on artificial intelligence reinforcement learning according to claim 2, characterized in that: In step S4, the action execution is evaluated from three aspects according to three objectives: 1) the total number of modules; 2) the types of modules; 3) Module alignment; specific evaluation methods include: S41. After placing the module, calculate the positional relationship between the placed module and the layout area. There are three possible positional relationships: ① The placed module overlaps with the previously placed module; ② The size of the placed module exceeds the room boundary; ③ The placed module neither overlaps with the layout area nor exceeds the room boundary. By examining the relationship, a reward r describing the total number of modules and the type of modules is obtained from a qualitative and quantitative perspective. 1、 r2; Because the total number of modules is described by the number of times the action is executed, executing any action indicates an increase in the total number of modules, and judging whether the module type appears indicates an increase in the type; S42. After the module is successfully placed, the degree of alignment between the modules in the entire room is determined, that is, whether there are many misaligned module vertices. The more misaligned module vertices there are, the lower the module alignment is considered to be, and a reward r3 for the module alignment degree is obtained; S43, the sign of training termination: the remaining area of ​​the room cannot accommodate any module, and feedback r4 is given by the inventory of the remaining space size; r1, r2, and r3 are adjusted according to the input weights α1, α2, and α3. Through different feedback signals, after training, a batch of results that meet the input goals are obtained.

4. The system for generating a two-dimensional bidirectional building module layout based on artificial intelligence reinforcement learning according to claim 3, characterized in that: The algorithm initialization unit initialization step specifically includes: S1. Initialize the environment, including: initializing each coordinate value of the dot matrix to -1; returning the initial position of the agent to the upper left corner of the room; clearing the historical action list; S2. Initialize the neural network and use the deep reinforcement learning method. The agent makes decisions and evaluates actions through the neural network. The neural network includes the action evaluation network N eval and target network N target , the initialization of the neural network includes: randomly initializing the action evaluation network N eval The weight of N target The weights are set to be equal to N eval same; S3. Initialize the experience buffer: Using deep reinforcement learning, the agent trains the neural network by sampling the existing data in the experience buffer. Positive and negative samples are stored in different experience replay buffers. Initializing the experience replay buffer includes clearing both buffers. S4. Initialize the parameters of the deep reinforcement learning agent, including: S41, initialization exploration rate ε, exploration decay rate ε decay The agent generates a random number to determine whether to randomly select an action or obtain an action through a neural network. If the random number is less than ε, the action is randomly selected, otherwise the current state is input into the neural network to obtain the output action. The initial ε is set to a large value, and the agent mainly performs actions randomly. After each action is determined, ε is adjusted according to ε. decay Decay, subsequent agents will tend to use neural networks to get actions; S42. Initialize the discount factor γ, which is used to calculate the time difference target value during training; S43, initialize the training batch size batch_size; S44. Initialize the target network parameter update frequency update_freq.

5. The system for generating a two-dimensional bidirectional building module layout based on artificial intelligence reinforcement learning according to claim 4, characterized in that: The algorithm training unit performs training by following the steps below: S1. Agent selection action: At any time in state S, the agent decides to randomly select a module or select the action A with the largest Q value based on the current state S according to the environment exploration rate ε; at the same time, the exploration rate ε is determined by the exploration decay rate ε. decay reduce; S2. Feedback from the environment: Based on the actions performed by the agent, i.e. the modules selected, the environment provides the agent with corresponding reward values ​​according to the reward mechanism. Specifically, the reward values ​​include: S21. Record the current agent position, find the corresponding module size according to the selected action, and update the agent position; S22. Give a reward R according to the reward mechanism; S23. Determine the placement position of the current module. If the length or width of the module plus the length or width of the agent exceeds the room range, the action is invalid, the module sequence is not updated, the current state Q value is updated, and the training round end flag T is set to true. If there is another module at the location where the current module is placed, the action is invalid, the module sequence is not updated, the current state Q value is updated, and the training round end flag T is set to true. If the location where the current module is placed does not exceed the room range and there is no other module, the action is valid, the module sequence is updated, the Q value is updated, and the agent's position is updated. S24. Obtain a new state S' according to the updated pointer; S25. Store the five-tuple (S, A, R, S', T) into the positive sample or negative sample experience buffer according to whether the action is valid; S3. Update the neural network: This step samples data from the experience buffer, trains the network, and completes the weight update. First, check whether the size of the positive and negative sample experience buffers is not less than batch_size / 2. If so, do not update the network; otherwise, update the network. The main process is as follows: S31. Sample batch_size / 2 from positive and negative samples respectively, combine them into a batch of data, and input N eval , get the estimated output yhat; S32, using the target network output to calculate the time difference target value y; S33. Calculate the mean square error loss of yhat and y, and update N with gradient descent eval ; S34. If the number of training times meets update_freq, N eval The weight value is copied to N target ; S4. Terminate training and output results: Perform training according to the set number of training rounds. During the training process, store the successful arrangement results and record the result information. After the training is completed, output the Pareto image of all results, and select a series of corresponding Pareto frontier solutions as the result output based on multiple objectives. Output a variety of arrangement methods for designers to screen.

6. The system for generating a two-dimensional bidirectional building module layout based on artificial intelligence reinforcement learning according to claim 5, characterized in that: The program evaluation unit performs program quantitative evaluation including the following steps: S1. Count the total number of modules and calculate the length of the execution action sequence of each area to be arranged to obtain the total number of modules; S2. Count the number of module types, calculate the number of modules, and store them in a dictionary form; S3. Count the module alignment, calculate the number of isolated vertices of adjacent modules in each arrangement area, and determine whether the length and width of adjacent modules are the same; S4. Output the evaluation results.

7. The system for generating a two-dimensional bidirectional building module layout based on artificial intelligence reinforcement learning according to claim 6, characterized in that: The algorithm output unit algorithm output includes: 1) a Pareto chart of all collected arrangement schemes; 2) a display of the Pareto frontier solutions among multiple targets; 3) a data table of the arrangement schemes of each module.

Citation Information

Patent Citations

  • Building layout generation method and device, computer equipment and storage medium

    CN113642090A

  • Module unit splitting method, system and equipment based on reinforcement learning and medium

    CN118153175A