Method for generating a factory layout plan for a factory by means of an electronic computing device, computer program product, computer-readable storage medium, and electronic computing device
The MARL-based method addresses the inefficiencies of existing plant layout generation by integrating a 3D physics engine and large language model to optimize plant layouts in complex environments, achieving efficient and human-friendly solutions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-04-09
AI Technical Summary
Existing methods for generating plant layout plans are inefficient and unsatisfactory in handling complex, three-dimensional real-world environments with multiple constraints and interdependencies, leading to suboptimal solutions due to the complexity of realistic boundary conditions and human factors.
A method utilizing Multi-Agent Reinforcement Learning (MARL) to generate a plant layout plan, incorporating a 3D physics engine and large language model to handle customized constraints and optimization goals, allowing for precise component placement and integration of human feedback to optimize the layout.
Enables the generation of realistic, approximate global optima plant layouts that satisfy both quantifiable and non-quantifiable human requirements, improving efficiency and reducing time and computational effort.
Smart Images

Figure EP2025077449_09042026_PF_FP_ABST
Abstract
Description
[0001] 202413507
[0002] 1
[0003] Description
[0004] Method for generating a plant layout plan for a plant using an electronic computing device, computer program product, computer-readable storage medium and electronic computing device
[0005] The following invention relates to a method for generating a plant layout plan for a plant using an electronic computing device according to claim 1. The invention further relates to a corresponding computer program product, a corresponding computer-readable storage medium and an electronic computing device.
[0006] Creating a comprehensive three-dimensional CAD model of a process or plant, especially a manufacturing plant, is a time-consuming, manual, and labor-intensive process consisting of several phases. Typically, in the first phase, also known as pre-basic or preliminary planning, engineers manually create 3D designs and 2D floor plans of the plant's sub-processes. Components are selected and connected in 3D plant engineering software according to diagrams of the process or material flow of the sub-processes and their variations. In later phases, also known as basic engineering, specific variations are selected, and their 3D designs and 2D floor plans are further detailed. Ultimately, a 3D model and floor plan of the final plant are developed.The components of the various trades are manually selected and placed in the 3D software and, if necessary, connected to each other via pipes or cables.
[0007] Depending on a client's requirements and the site conditions, such as brownfield or greenfield, the 3D designs and the final 3D model must consider numerous constraints and be arranged according to optimization goals. Potential constraints can include distances between components, areas in the greenfield or brownfield that must remain free of components, and optimization goals such as space saving, optimized material flow, or the shortest overall length of pipes or cables. The quality of the 3D designs and 2D floor plans depends heavily on the experience of the engineers developing them. In complex scenarios, such as a large number of components and constraints or intricate brownfield structures, a manual approach to optimal solutions is not practical. There are many possible ways to arrange the components, and it is impossible for humans to consider every combination.
[0008] 2. However, the quality of the early 3D design has a significant impact on, for example, the overall costs and effort of the plant project, and the design of the plant has an impact on future operating costs, which is why a near-optimal solution brings particular time savings.
[0009] The state of the art already includes approaches that address the problem of optimally arranging physical components and resources, particularly with regard to maximum efficiency, safety, and minimum cost, to develop a model of a facility. This problem is also often referred to as the Facility Layout Problem (FLP) and is classified as non-polynomial hard due to its algorithmic complexity. Optimal solutions can only be approximated. Most approaches use metaheuristics, such as genetic algorithms, simulated approximation, and particle swarm optimization. Furthermore, mathematical modeling methods, such as mixed-integer nonlinear programming (MINLP), are also employed. Intelligent solutions using artificial neural networks are also proposed in the literature, including reinforcement learning (RL) approaches.
[0010] Nevertheless, all known approaches are unable to provide solutions that address realistic problems, due to the complexity of the real world. The primary reason is that the real world is a continuous three-dimensional space, which is difficult for most known automated approaches to the facility layout problem to handle. Furthermore, every component to be placed is subject to several environmental constraints, such as its weight and size, because physical boundaries exist in the real world. Additionally, there are safety restrictions, for example, on brownfield sites, where certain areas for component placement must be avoided or prohibited due to the risk of explosion or other hazards.Besides the environment itself, the components have several interdependencies, such as free space, spatial constraints (like placing one component type above others), and many more. Finally, human aspects must also be considered, such as well-being in an environment or aesthetic aspects of the solution, which are not easily quantifiable and therefore difficult for computer systems to handle.
[0011] To meet these requirements, metaheuristic methods and mathematical modeling methods are not suitable, as they do not provide intelligent behavior for the components 202413507.
[0012] Three methods can be learned, and considering mutual dependencies and boundary conditions is too complex. Therefore, known approaches simplify the problem by treating only two-dimensional discrete grids, where the components are simplified as 2D objects and surrounded by only a few boundary conditions. As a result, the provided solution is extensively reviewed, modified, and extended by engineers to an exception, raising questions about the usefulness of these methods. Known reinforcement learning (RL) methods also cannot handle the various boundary conditions for individual components, leading to suboptimal solutions. Therefore, known reinforcement learning approaches only treat simplified discrete layouts with few components.Human feedback as a further optimization goal is used in some metaheuristic approaches, but only in a simple way by rating the solution as good or bad, which cannot lead to optimal or approximately optimal solutions because the feedback is too sparse.
[0013] In summary, the problem of optimizing plant layouts in a three-dimensional real environment, taking into account realistic boundary conditions and multiple simultaneous optimization goals, remains unsolved and is only addressed in an unsatisfactory way.
[0014] The object of the present invention is to provide a method, a computer program product, a computer-readable storage medium, and an electronic computing device by means of which the disadvantages of the prior art are overcome. In particular, it is an object of the present invention to provide a method, a computer program product, a computer-readable storage medium, and an electronic computing device by means of which an improved plant layout plan for a plant can be generated.
[0015] This problem is solved by a method, a computer program product, a computer-readable storage medium, and an electronic computing device according to the independent claims. Advantageous embodiments are specified in the dependent claims.
[0016] One aspect of the invention relates to a method for generating a plant layout plan using an electronic computer. At least one component list for the plant is provided by the electronic computer. The electronic computer then determines at least one dependency on at least one component in the component list. The plant layout plan is then generated.
[0017] 4. generated by means of a multi-agent reinforcement learning algorithm depending on at least one specific dependency and on the component list using the electronic computing device.
[0018] This makes it possible to implement a reliable system architecture for a plant based on Multi-Agent Reinforcement Learning (MARL), which includes, in particular, mutual dependencies.
[0019] In particular, a 3D physics engine is essentially proposed for placing components in environments such as brownfield or greenfield sites. Within these environments, the components are then positioned using the Multi-Agent Reinforcement Learning (MARL) method. This method can also be enhanced with a feedback loop for human instructions, which are understandable to the MARL approach, to create and refine solutions. This leads, in particular, to a workflow for overcoming the described problems according to the state of the art.
[0020] The MARL method can process customized constraints and optimization goals for individual components while simultaneously optimizing the overall plant layout, particularly the plant design plan. In this way, the electronic computing system can handle multiple boundary conditions and interdependencies between components, as well as constraints, for example, in brownfield or greenfield environments. To enable a realistic number of constraints and goals, human feedback can be used to define additional, previously unknown constraints and goals, providing the MARL approach with instructions for adjustment. This simultaneously optimizes the fulfillment of non-quantifiable human requirements.
[0021] Additional input can include, for example, a process or material flow diagram of the underlying manufacturing or chemical process. Furthermore, a list, specifically the Engineering Bill of Materials (EBOM), containing 3D models of the physical components to be placed to implement the underlying process, is provided. The Engineering Bill of Materials essentially corresponds to the component list. In the case of integrating a facility into a brownfield site, a 3D model of the brownfield is available, which is used to optimize the plant layout and configure the equipment. Potentially, additional information regarding the brownfield boundaries, such as areas where component placement is prohibited, can be included.
[0022] 5. Potential initial optimization targets to be optimized can also be provided.
[0023] As previously mentioned, the proposed solution utilizes multi-agent reinforcement learning (MARL), a subset of machine learning, which operates in an environment with multiple agents, corresponding to the respective components. Each component is treated as an independent unit with a limited set of resources, attempting to develop an optimal strategy through learning to achieve a predefined goal. Unlike traditional reinforcement learning, where only one agent operates within an environment, the components in MARL must collaborate and compete to reach their objectives. This allows the agents to learn from and adapt to each other, enabling them to make better decisions and improve their performance.
[0024] The use of the MARL approach in a continuous 3D environment to optimize plant layout generation is particularly novel compared to previous approaches. The MARL approach in a continuous 3D environment can be used, in particular, to support components in optimizing their overall goal. This enables a continuous 3D environment with, for example, collision grids, and thus significantly more precise component placement in a virtual environment than with prior art, considerably simplifying the real world. Furthermore, the MARL approach, using, for example, artificial neural networks, allows for the scaling of many components with complex constraints and dependencies compared to metaheuristic, mathematical modeling, or single-agent reinforcement learning approaches, since each component has its own neural network capable of learning complex relationships.
[0025] Brownfield and greenfield, which—as already mentioned—can be optimized accordingly, are terms used in urban planning, real estate, and site selection for industrial facilities. Brownfield refers to a plot of land or area that already contains certain infrastructure, such as old buildings, factory buildings, or contaminated soil. Extensive remediation and development measures are generally required for the new use of a brownfield site to remove existing structures, clean the soil, and create new infrastructure. Greenfield, on the other hand, refers to a plot of land or area that is not yet built upon or used.
[0026] The site was developed and is free of any existing infrastructure. Developing a greenfield site generally requires less extensive measures than developing a brownfield site, as there is no need to remediate contaminated land and a new infrastructure can be planned from scratch. Overall, the use of brownfield sites is therefore more complex than the development of greenfield facilities. On the other hand, the repurposing of brownfield sites also offers advantages, such as better access to infrastructure and other facilities, as well as reduced land use and thus greater sustainability.
[0027] In one advantageous embodiment, a three-dimensional plant layout plan is generated using a multi-agent reinforcement learning algorithm. This allows not only the creation of a two-dimensional floor plan, but, more importantly, a three-dimensional plant layout plan. Specifically, this three-dimensional plant layout plan can then be displayed on a screen of the electronic computer system. This allows a user of the electronic computer system to reliably study the generated three-dimensional plant layout plan and make necessary changes. In this way, the plant can be planned realistically.
[0028] It has also proven advantageous to determine the placement of at least one component within the plant layout plan. In particular, the placement within the three-dimensional plant layout plan is determined. Various dependencies, such as predefined component placements relative to each other or, for example, corresponding connections in a brownfield environment, can be used as criteria to determine the appropriate placement within the plant layout plan.
[0029] It has also proven advantageous to consider at least one environmental condition of the environment in which the plant is positioned when determining the plant layout plan. For example, the room layout, the room's dimensions (length, height, and width), and any fixed objects can be considered, along with other environmental conditions. In this way, the environmental condition can be specified as a boundary condition, and the plant layout plan can be reliably determined.
[0030] It has also proven advantageous to consider at least one existing spatial constraint and / or at least one future spatial constraint as an environmental condition. For example, in the brownfield area 202413507
[0031] Seven corresponding spatial constraints may exist, such as the size of the space. Furthermore, it can also be considered that, for example, a roof or other structure may need to be built in a greenfield site in the future. These constraints can then be used as boundary conditions in the plant layout plan. Thus, both existing and future spatial constraints can be taken into account, allowing for the reliable generation of the plant layout plan.
[0032] In another advantageous embodiment, a user input is captured via a large language model of the electronic computing device, and this input is considered as a boundary condition for the system design. A large language model (LLM) is a subfield of artificial intelligence (AI) trained on enormous amounts of text data to generate human-like speech. These models can understand and respond to natural language input by mimicking human conversational style, enabling them to be used for a wide variety of applications such as text generation, translation, summarizing, and answering questions. Large language models are typically built using deep learning techniques such as recurrent neural networks (RNNs) or transformers, which allow the models to learn complex patterns and dependencies in the training data.The larger the model and the more data it has been trained on, the better its ability to generate coherent and natural-sounding speech.
[0033] In particular, the use of the large language model allows for the integration of human feedback into the optimization process. This is because, in previous approaches, users could already specify all constraints, goals, dependencies, and unconscious requirements, such as aesthetics, in advance—something that is not directly feasible for a human. Through intelligent human feedback via the integrated LLM interface to the MARL method, plant designs can thus be provided in complex environments that correspond to real-world use cases.
[0034] It has also proven advantageous if at least one placement prohibition and / or placement requirement and / or placement request and / or component dependency is captured via the large language model. This allows the user to input the relevant conditions via the large language model, which are then transformed in a way that is readable by the algorithm. Thus, a 202413507
[0035] 8. Human feedback is provided, allowing for the integration of various boundary conditions by humans that were not previously foreseeable, for example, via the component list. This enables even complex user-defined requirements to be taken into account, ensuring the reliable generation of the plant layout plan.
[0036] Another advantageous design feature provides that the current calculation progress for generating the plant layout plan is displayed on a display device.
[0037] Generating the plant layout plan is a particularly time-consuming step, even in simulation. Therefore, the current calculation progress can be displayed during the simulation or calculation itself. This allows users to intervene early if calculation steps do not meet their requirements. Boundary conditions can thus be incorporated during the calculation, reducing the time and computational effort required to generate the plant layout plan.
[0038] It has proven advantageous to allow user input during the calculation process for generating the plant layout plan and to incorporate this input into the calculation. For example, the user can input information via a large language model. This allows the user to see undesired placements during the calculation process and then, via the user interface, prohibit such placements. This enables adaptation of the plant layout plan through human feedback during the calculation process. This significantly saves time in determining the plant layout plan, as adjustments can be made to the calculation or generation of the plan during the process itself.
[0039] In another advantageous embodiment, the multi-agent reinforcement learning algorithm is provided with a central compute node that evaluates each individual calculation performed by an agent. Each agent takes into account the central compute node's evaluation during its individual calculation. Thus, a higher-level compute node can be used as a so-called critic, which in turn monitors and improves the individual agents, or components, thereby enabling the agents themselves to learn.
[0040] 9. They can optimize their actions to achieve placements that optimize the overall problem without violating their individual and agent-dependent constraints. Furthermore, the language model interface can also be used to influence corresponding agent constraints and goals to globally satisfy non-quantifiable human goals, while the MARL algorithm simultaneously optimizes the main goal of the plant layout plan, thus leading to realistic approximate global optima of the plant layout plan that also satisfy individual human expectations.
[0041] Furthermore, it can be stipulated that the plant layout plan is generated based on a specific collision mesh for plant components. The collision mesh is, in particular, a virtual 3D model of, for example, the plant, which is used to calculate collisions. It is a simplified representation of objects that contains only components relevant for collision calculation, such as corners, edges, and faces. If two objects move within the MARL algorithm, for example, collision detection algorithms can use the collision mesh of both objects to determine whether they are touching or intersecting. This prevents the components from disappearing into each other or interacting unrealistically.
[0042] The presented method is, in particular, a computer-implemented method. Therefore, a further aspect of the invention relates to a computer program product with program code means which, when the program code means are executed by the electronic computing device, cause a method according to the preceding aspect to be carried out.
[0043] Furthermore, the invention also relates to a computer-readable storage medium containing at least the computer program product according to the preceding aspect.
[0044] A further aspect of the invention relates to an electronic computing device for generating a plant layout plan, comprising at least one multi-agent reinforcement learning algorithm, wherein the electronic computing device is configured to carry out the method. In particular, the method is carried out by means of the electronic computing device. 202413507
[0045] 10
[0046] A computing unit / electronic computing device can be understood, in particular, as a data processing device containing a processing circuit. The computing unit can therefore process data to perform arithmetic operations. This may also include operations to perform indexed access to a data structure, such as a lookup table (LUT).
[0047] The computing unit may, in particular, contain one or more computers, one or more microcontrollers, and / or one or more integrated circuits, for example, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), and / or one or more systems on a chip (SoCs). The computing unit may also contain one or more processors, for example, one or more microprocessors, one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more signal processors, in particular one or more digital signal processors (DSPs). The computing unit may also include a physical or virtual array of computers or other units of the aforementioned type.
[0048] In various embodiments, the computing unit includes one or more hardware and / or software interfaces and / or one or more storage units.
[0049] A storage unit can be volatile data storage, for example as dynamic random access memory (DRAM) or static random access memory (SRAM), or as non-volatile data storage, for example as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or flash EEPROM, ferroelectric random access memory (FRAM), or magnetoresistive random access memory.It can be designed as MRAM (magnetoresistive random access memory) or as phase-change random access memory, PCRAM (phase-change random access memory). 202413507
[0050] 11
[0051] Here and in the following, an artificial neural network can be understood as software code stored on a computer-readable storage medium that represents one or more interconnected artificial neurons or can replicate their function. The software code can also contain multiple software code components, which may, for example, have different functions. In particular, an artificial neural network can implement a nonlinear model or a nonlinear algorithm that maps an input to an output, where the input is given by an input feature vector or an input sequence, and the output may, for example, include a category for a classification task, one or more predicated values, or a predicated sequence.
[0052] For use cases or application situations that may arise in a method according to the invention and that are not explicitly described herein, it may be provided that, according to the method, an error message and / or a request for user feedback is issued and / or a default setting and / or a predetermined initial state is set.
[0053] Regardless of the grammatical gender of a particular term, persons with male, female or other gender identities are included.
[0054] Further features and combinations of features of the invention will become apparent from the figures and their descriptions, as well as from the claims. In particular, further embodiments of the invention need not necessarily include all features of any one of the claims. Further embodiments of the invention may have features or combinations of features that are not mentioned in the claims.
[0055] The single figure, Fig. 1, shows a schematic block diagram of an embodiment of an electronic computing device.
[0056] In the figure, identical or functionally equivalent elements are provided with the same reference symbols.
[0057] Fig. 1 shows a schematic block diagram according to an embodiment of an electronic computing device 10 for generating a plant layout plan 12 for a 202413507
[0058] 12
[0059] The system is designed using the electronic computing device 10. At least one component list 14 for the system is provided using the electronic computing device 10. At least one dependency on at least one component from the component list 14 within the system is determined using the electronic computing device 10. The system layout plan 12 is then determined using a multi-agent reinforcement learning algorithm 16, depending on the at least one determined dependency and on the component list 14, using the electronic computing device 10.
[0060] In particular, it is provided that a three-dimensional plant layout plan 12 is generated using the multi-agent reinforcement learning algorithm 16. Furthermore, it is provided that the placement of at least one component in the plant layout plan 12 is determined. It can also be provided that at least one environmental condition 18 of an environment in which the plant is positioned is taken into account when determining the plant layout plan 12. For this purpose, it can be provided, for example, that a collision network 20 is determined based on the environmental conditions 18 and the component list 14, and that the plant layout plan 12 is then determined based on the collision network 20.
[0061] Furthermore, it may be stipulated that at least one existing spatial constraint and / or at least one future spatial constraint be taken into account as an environmental condition 18. In addition, spatial dependencies of at least two components may also be considered when generating the plant layout plan 12.
[0062] Furthermore, it may be provided that a user input, in particular via an input device 24, is recorded via a large language model 22 of the electronic computing device 10 and that the input is taken into account as a boundary condition for the plant layout plan 12. In this way, at least a placement prohibition and / or the placement requirement and / or a placement request and / or a component dependency can be recorded via the large language model 22.
[0063] Furthermore, the current calculation progress for generating the plant layout plan 12 can also be displayed on a display unit 26 of the electronic computing unit 10. Additionally, the calculation process for generating 202413507 can also be displayed.
[0064] 13 of the plant layout plan 12, an input by the user can be recorded and the recorded input is taken into account during the calculation process.
[0065] Furthermore, it may be provided that the multi-agent reinforcement learning algorithm 16 is provided with a central computing node which evaluates each individual calculation of an agent, and whereby an evaluation of the central computing node by each agent is taken into account in the individual calculation.
[0066] Furthermore, Fig. 1 also shows that corresponding process parameters 28 for process control can also be taken into account in the multi-agent reinforcement learning algorithm 16 for generating the plant layout plan 12.
[0067] In particular, the proposed approach envisages providing a 3D physics engine to create, for example, a brownfield or greenfield environment as a continuous 3D setting in which components can be placed using the Multi-Agent Reinforcement Learning Algorithm 16 (MARL). This method is also linked, for example, to the large language model 22, which transforms human feedback into instructions understandable to the MARL algorithm for creating and adapting solutions. This results in a workflow capable of overcoming prior art problems. The MARL algorithm can handle customized constraints and optimization goals for individual components while simultaneously optimizing the overall layout, especially the plant layout plan 12. This allows it to process multiple boundary conditions and interdependencies between components, as well as constraints inherent in the brownfield or greenfield environment.To enable a realistic number of boundaries and goals for the solution, the large language model 22 can capture human feedback as well as constraints or goals that were previously unknown and also give instructions to the MARL algorithm to make adjustments that allow simultaneous optimization of the non-quantifiable human requirements.
[0068] The process parameter 28 of the underlying manufacturing or chemical process can be provided as input for the method. Furthermore, the component list 14, containing the 3D models of the physical component to be placed, can be provided to implement the underlying process. If a plant is to be integrated into a brownfield site, a 3D model of the brownfield is provided, which in this case is characterized in particular by the environmental conditions 18 202413507.
[0069] Figure 14 illustrates this. Additional information regarding the brownfield boundary conditions, such as areas where the placement of components is prohibited, can be provided. Potentially, initial optimization targets are provided for optimization.
[0070] In a first step, a collision mesh 20 is created for each 3D model of the components, such as meshes, point clouds, corresponding object files, part files, drawings, or the like. This mesh best approximates the actual shape of the object. The same process is performed for the 3D model of the brownfield. In addition to each collision mesh 20, the local coordinates for inputs and outputs are added to the geometry according to the process flow diagram. Consequently, the collision meshes 20 of the components and the brownfield can be loaded into a physics engine, and the components can be placed and moved within it. In this way, collisions between components or with the brownfield, as well as distances between components and surfaces, can be measured. The collision meshes 20 of the components and the brownfield, integrated into the physics engine, serve as the environment for the MARL method.In the case of a greenfield environment, the components can be freely placed inside the physics engine.
[0071] Since the 3D models are converted into collision meshes, and the MARL learns and performs the placement using a physics engine, the method is applicable to any known software used to create 3D models of plants. After the coordinates and rotations of the components have been developed by the MARL method within the physics engine, the actual 3D models of the components can be automatically placed in the software program using a simple interface to the API of the software in use.
[0072] Within the MARL method, each component is represented by a single agent that can be moved simultaneously with the others in the physics engine. The underlying problem for the MARL method is a partially observable Markov game (POMG) in which plant layout plan 12 is optimized. For each agent, the environment is dynamic and partially observable. In addition to the agents, there is a central computational node, also called the critic, for which the environment is fully observable and which predicts how well the actions of each agent will perform toward the overall optimization objectives. Each agent and the critic contain an artificial neural network (ANN) that is trained over time through interaction with the environment to produce an optimized plant layout by optimizing the plant layout problem.
[0073] The corresponding Markov game is defined by the tuple N, S, A, R, T, O, y, Q, in which the agents can interact with the environment over a specific number of episodes with a specific number of journals T. It comprises a number of N agents, each representing a specific component. The number of agents depends on the number of components required to realize the underlying process. The state space S is a three-dimensional continuous object representing the brownfield or greenfield environment. The observation O that an agent receives at each journal T is its own x, y, z position in the environment, a distance vector increasing from its output to the target point according to the process flow, and a distance vector pointing to the nearest collision object within a given area.It may also include a point cloud input of its environment as an additional input to allow the agent to learn how to navigate the environment and avoid collisions using a visual three-dimensional input. Therefore, O is a partial view of the overall state S of the environment. A is the so-called Joint Action Space, which is a continuous space that describes the direction and rotation along the x, y, and z axes within a maximum font size. Since each agent should only move within its observable space, the area is chosen according to the input point cloud or the area in which it can observe external collision networks. Q is the critic, which predicts whether the actions chosen by the agent are good or bad in order to optimize the long-term accumulated reward goal, which represents the overall optimization objectives.The critic observes the environment completely; that is, it receives the observations of all agents along with their chosen actions as its own observation. Furthermore, it can potentially receive additional information not provided to the agents. This could, for example, be a point cloud representing the entire environment or values representing the fulfillment of dependencies between groups of agents. To achieve the optimization goal and satisfy the individual boundary conditions of the agents, both agents and critic are trained using the reward function R. The reward function R is designed so that the critic learns to increase the overall optimization objectives. For example, to minimize material handling costs, the distance between material inputs and outputs should be optimized. The agents themselves receive an individual reward in each journal.This reward is positive if the agent moves its output closer to the input of other components that match its own output. Causing a collision results in a negative reward. Additionally, the agent receives a small negative reward in each magazine to incentivize overall movement. 202413507.
[0074] 16
[0075] An important aspect is that the individual rewards and the reward function for the agents and the critic are not fixed and deterministically predetermined before the optimization process, but can be modified, for example, by the output of the large language model 22 during the process. The output of the large language model 22 is thus the human-translated feedback that guides the MARL process. For example, the agent can be given an additional reward for each magazine if it increases the distance to a wall, if the human wants the agent to move its collision net 20 further away from the wall. These rewards can be predefined, and the large language model 22 then selects the appropriate predefined reward setting, which is then executed, or the large language model 22 can provide the additional reward itself in mathematical form.
[0076] As a result, the agents receive their individual reward in each journal, which helps them learn how to improve their actions based on their individual observations, taking into account the agent's individual constraints and goals, such as those defined by the large language model 22. Conversely, in each journal, the critic receives each agent's observations and actions, along with any additional information, and evaluates whether the agents' actions improve the overall order of the facility. Thus, the facility's design plan 12 is globally approximated by the critic, who monitors and improves the agents as they learn to optimize their actions to achieve positions that optimize the overall goal without violating their individual and agent-dependent constraints.Furthermore, the interface with the large language model 22 allows a user to manipulate human constraints and agent goals to fulfill non-quantifiable human objectives, while the MARL algorithm simultaneously continues to have the main goal of plant optimization, thus leading to realistic approximate global optima of the arrangement that also meets individual human expectations.
[0077] The large language model 22 offers the further advantage of being an interface that, for example, raw text in human language can be received and translated into information that the MARL process can understand and integrate. The large language model 22 can either take the raw text and assign predefined constraints, goals, or tasks to it, thus connecting it to the agent. An example of this is a requirement that a component should be attached in a specific position. The human could then, for example, specify that a particular component, such as a 202413507, must be attached to a specific position.
[0078] 17. A blue water tank is to remain in the corner, and the large language model 22 would translate this to the MARL algorithm, indicating that the corresponding agent jumps to the required position and performs the action [0,0,0], which does not result in any movement from that point onward. Consequently, the large language model 22 only needs to translate which agent is the mentioned component and to which position it must immediately move. Therefore, no further reward is required. On the other hand, the large language model 22 can also allow manipulation of the reward function to add a goal for the agent, as already described in the example with the component and the distance to a specific wall. Consequently, the large language model 22 is superior as an interface to non-intelligent interfaces, where, for example, a human has to handle a large number of parameters.
[0079] The architecture of the neural networks for agents and critics can consist of any number of the following layers: dense layers, long-term / short-term storage layers, attention layers, convolving layers, and pooling layers. Convolving and pooling layers are particularly advantageous in cases where agents obtain point clouds in their observations. Attention and long-term / short-term storage layers are useful in highly complex environments because they perform better in scenarios where time dependency is important, as is the case with the invention. Additionally, there is an advantage to increasing the number and size of layers within the agent's artificial neural network as the complexity of the environment increases, and similarly, increasing the number and size of layers within the critic's artificial neural network as the number of agents increases.Since the training time increases with larger artificial neural networks, the size of the artificial neural networks must be chosen to achieve a sensible balance between training time and performance.
[0080] The large language model 22 used can be one of the existing pre-trained models known from the prior art, as these large advanced language models 22 already demonstrate very good solutions. However, smaller large language models 22 that can be run offline on a local computer can also be used. 202413507
[0081] 18
[0082] Reference symbol list
[0083] 10 electronic computing equipment
[0084] 12 Plant layout plan 14 Component list
[0085] 16 Multi-agent reinforcement learning algorithm
[0086] 18 Environmental conditions
[0087] 20 Collision net
[0088] 22 large language model 24 input device
[0089] 26 Display unit
[0090] 28 process parameters
Claims
202413507 19 Patent claims 1. Method for generating a plant layout plan (12) for a plant using an electronic computing device (10), comprising the steps: - Providing at least a component list (14) for the plant using the electronic computing device (10); - Determining at least one dependency on at least one component of the component list (14) using the electronic computing device (10); and - Generating the plant layout plan (12) using a multi-agent reinforcement learning algorithm (16) depending on at least one specific dependency and on the component list (14) using the electronic computing device (10).
2. Method according to claim 1, characterized in that a three-dimensional plant layout plan (12) is generated using the multi-agent reinforcement learning algorithm (16).
3. Method according to claim 1 or 2, characterized in that a placement of the at least one component in the plant layout plan (12) is determined.
4. Method according to one of the preceding claims, characterized in that at least one environmental condition (18) of an environment in which the plant is positioned is taken into account when determining the plant layout plan (12).
5. Method according to claim 4, characterized in that at least one existing spatial constraint and / or at least one future spatial constraint is taken into account as an environmental condition (18).
6. Method according to one of the preceding claims, characterized in that spatial dependencies of at least two components are taken into account when generating the plant layout plan (12).
7. Method according to one of the preceding claims, characterized in that 202413507 20 via a large language model (22) of the electronic computing device (10) an input from a user is recorded and the input is taken into account as a boundary condition for the plant layout plan (12).
8. Method according to claim 7, characterized in that at least one placement prohibition and / or one placement requirement and / or one placement request and / or one component dependency are captured via the large language model (22).
9. Method according to one of the preceding claims, characterized in that a current calculation progress for generating the plant layout plan (12) is displayed on a display device (26).
10. Method according to claim 9, characterized in that during the calculation process for generating the plant layout plan (12) an input by a user can be recorded and the recorded input is taken into account during the calculation process.
11. Method according to one of the preceding claims, characterized in that the multi-agent reinforcement learning algorithm (16) is provided with a central computing node which evaluates each individual calculation of an agent, and wherein each agent takes into account an evaluation of the central computing node during the individual calculation.
12. Method according to one of the preceding claims, characterized in that the plant layout plan (12) is generated on the basis of a specific collision network (20) for the components of the plant.
13. Computer program product comprising program code means which cause an electronic computing device (10) to perform a method according to one of claims 1 to 12 when the program code means are executed by the electronic computing device (10).
14. Computer-readable storage medium comprising at least one computer program product according to claim 13. 202413507 21 15. Electronic computing device (10) for generating a plant layout plan (12) for a plant, comprising at least one multi-agent reinforcement learning algorithm (16), wherein the electronic computing device (10) is configured to perform a method according to one of claims 1 to 12.
Citation Information
Patent Citations
Automatic building space combination method based on multi-agent deep reinforcement learning
CN118296702A