Railway precast beam yard beam storage scheduling optimization method and device based on reinforcement learning

By optimizing the beam storage and scheduling of railway precast beam yards using deep neural networks based on reinforcement learning, the problem of lack of data support for beam storage and scheduling in existing technologies has been solved, thereby improving the utilization rate and work efficiency of beam yards and reducing resource waste and downtime.

CN119249549BActive Publication Date: 2026-04-07CHINA STATE RAILWAY GRP CO LTD +3
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The existing railway precast beam yards lack data-supported quantitative analysis and intelligent optimization methods for beam storage and scheduling, resulting in frequent conflicts, insufficient site utilization, long scheduling time, and difficulty in coping with changes in uncertain factors on site.

Method used

A deep neural network based on reinforcement learning is used to optimize the beam storage and scheduling of railway precast beam yards. By obtaining the simulation model and business logic model of the target beam yard, the business module is constructed using object-oriented methods. Combined with the motion rules of the beam moving machine and the bridge erecting machine, the operation behavior parameters are optimized to improve the efficiency of beam storage and scheduling.

Benefits of technology

It has achieved optimization of beam storage planning and scheduling under the constraints of pedestal, equipment and personnel resources, improved beam yard utilization and work efficiency, and reduced resource waste and downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119249549B_ABST
    Figure CN119249549B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for optimizing beam storage scheduling in railway precast beam yards based on reinforcement learning, relating to the field of railway bridge construction engineering management technology. The method includes: acquiring a pre-established simulation model of the target beam yard; wherein the simulation model includes operational behavior parameters of the target beam yard, which represent the behavior and interaction methods of each module of the target beam yard; optimizing the operational behavior parameters using a deep neural network based on reinforcement learning based on a pre-determined objective function, and optimizing the beam storage scheduling of the target beam yard based on the optimized operational behavior parameters. The method provided by this invention establishes a corresponding simulation model, which allows for convenient adjustment of model parameters and closer interaction with the beam yard's operational status; by utilizing a deep neural network based on reinforcement learning to adjust model behavior, it overcomes the limitations of beam storage scheduling optimization technology under constraints from multiple resources such as piers, equipment, and personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway bridge construction engineering management technology, and in particular to a method and device for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning. Background Technology

[0002] Precast beam yards require significant resource investment and occupy large areas, necessitating intelligent and efficient management to improve the construction efficiency of high-speed railway bridges. With the widespread application of engineering construction management platforms, the production scheduling of precast beam yards has begun to be managed through information technology.

[0003] Specifically regarding beam storage scheduling, current mainstream beam yard scheduling decisions and schemes are often highly subjective, relying mainly on the production experience of project engineers and some relevant standard documents, lacking data-supported quantitative analysis and intelligent simulation optimization methods. In recent years, in response to problems such as frequent beam storage scheduling conflicts, insufficient site utilization, and long scheduling times in traditional beam yards, some scheduling schemes based on swarm optimization algorithms have emerged. However, these schemes suffer from problems such as a single optimization objective and a weak connection with the on-site construction process, making it difficult to cope with changes in uncertainties on-site.

[0004] Optimizing the scheduling of precast beams in railway precast beam yards and improving optimization efficiency are technical problems that need to be solved. Summary of the Invention

[0005] This invention provides a method and apparatus for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning, in order to overcome the deficiencies in the existing technology.

[0006] This invention provides a reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards, comprising the following steps:

[0007] Obtain a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field;

[0008] Based on a predetermined objective function, the operational behavior parameters are optimized using a deep neural network based on reinforcement learning, and the beam storage scheduling of the target beam yard is optimized based on the optimized operational behavior parameters.

[0009] According to the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards provided by the present invention, before obtaining the pre-established simulation model of the target beam yard, the method further includes:

[0010] Obtain a pre-established business logic model of the target beam yard, and construct simulation operation logic for the business logic model to obtain a simulation model of the target beam yard;

[0011] The business logic model includes a beam fabrication area business logic model and a beam storage area business logic model; the beam fabrication area business logic model is used to represent the business situation of the beam fabrication area in the target beam yard, and the beam storage area business logic model is used to represent the construction situation of the beam storage area in the target beam yard.

[0012] According to the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards, the step of constructing a simulation operation logic for the business logic model to obtain a simulation model of the target beam yard includes:

[0013] Based on the aforementioned business logic model, an object-oriented approach is used to construct the business modules for the target beam yard; wherein, the business modules include a production module, a storage module, and an erection module;

[0014] Based on the behavior and interaction methods of each module of the target beam yard, the operation behavior parameters of the target beam yard are determined and displayed in order to establish a simulation model of the target beam yard.

[0015] According to the present invention, a method for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning is provided, wherein the target beam yard further includes a beam-moving machine and a bridge-erecting machine;

[0016] After determining and displaying the operational behavior parameters of the target beam yard based on the behavior and interaction methods of each module of the target beam yard to establish a simulation model of the target beam yard, the method further includes:

[0017] Based on the beam yard shape of the target beam yard, motion rule modeling is performed on the beam moving machine and the bridge erecting machine.

[0018] According to the present invention, a method for optimizing the storage and scheduling of precast beams in railway precast beam yards based on reinforcement learning is provided, wherein the objective function includes: beam moving machine movement distance, beam yard utilization rate, and manual labor utilization rate;

[0019] The step of optimizing the operational behavior parameters using a deep neural network based on reinforcement learning, based on a pre-determined objective function, to optimize the beam storage scheduling of the target beam yard based on the optimized operational behavior parameters, includes:

[0020] Based on the moving distance of the beam transfer machine, the utilization rate of the beam yard, and the utilization rate of manpower, a deep neural network based on reinforcement learning is used to optimize the operation behavior parameters, so as to optimize the beam storage schedule of the target beam yard based on the optimized operation behavior parameters.

[0021] According to the present invention, a method for optimizing the storage and scheduling of precast beams in railway precast beam yards based on reinforcement learning is provided. The method involves optimizing the operational behavior parameters using a deep neural network based on reinforcement learning, based on a pre-determined objective function, to optimize the storage and scheduling of the target beam yard based on the optimized operational behavior parameters. The method includes:

[0022] The predetermined objective function is used as the main evaluation parameter, and the operational behavior parameters of the target beam yard are used as auxiliary evaluation parameters.

[0023] The main evaluation parameters and the auxiliary evaluation parameters are input into a deep neural network based on reinforcement learning to obtain optimized parameters, and the beam storage schedule of the target beam yard is optimized based on the optimized parameters.

[0024] This invention also provides a reinforcement learning-based device for optimizing the scheduling of precast beams in railway precast beam yards, comprising the following units:

[0025] An acquisition unit is used to acquire a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field;

[0026] An optimization unit is used to optimize the operational behavior parameters using a deep neural network based on reinforcement learning, based on a pre-determined objective function, so as to optimize the beam storage schedule of the target beam yard based on the optimized operational behavior parameters.

[0027] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the reinforcement learning-based railway precast beam yard storage scheduling optimization method as described above.

[0028] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards as described above.

[0029] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards as described above.

[0030] This invention provides a method and apparatus for optimizing beam storage scheduling in railway precast beam yards based on reinforcement learning. It acquires a pre-established simulation model of the target beam yard. The simulation model includes operational behavior parameters representing the behavior and interaction of each module within the target beam yard. Based on a pre-determined objective function, a deep neural network based on reinforcement learning is used to optimize the operational behavior parameters, thereby optimizing the beam storage scheduling of the target beam yard. Therefore, this invention establishes a corresponding simulation model with convenient parameter adjustment, and allows for real-time intervention in the simulation process, enabling closer interaction with the beam yard's operational status. By adjusting the model's behavior using a deep neural network based on reinforcement learning, based on a pre-determined objective function, it overcomes the limitations of beam storage scheduling optimization techniques under constraints related to resources such as platform size, equipment, and personnel. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0032] Figure 1 This is a flowchart illustrating the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards, as provided by this invention.

[0033] Figure 2 This is a schematic diagram of the beam yard spatial planar modeling provided by the present invention.

[0034] Figure 3 This is a schematic diagram of the overall framework of reinforcement learning provided by the present invention.

[0035] Figure 4 This is a schematic diagram of the structure of the railway precast beam yard storage scheduling optimization device based on reinforcement learning provided by the present invention.

[0036] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0038] The following is combined with Figures 1-5 This invention describes a method and apparatus for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning.

[0039] Figure 1 This is a flowchart illustrating the reinforcement learning-based method for optimizing the scheduling of precast railway beams in railway yards, as provided in this invention. Figure 1 As shown, the method includes the following:

[0040] Step 100: Obtain a pre-established simulation model of the target beam yard; wherein, the simulation model includes operational behavior parameters of the target beam yard, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam yard.

[0041] It should be noted that railway precast beam yards are crucial facilities in railway construction, primarily responsible for the manufacturing, curing, storage, and transportation of precast box girders. Through standardized production processes, precast beam yards can efficiently produce beams that meet design requirements, providing key components for railway bridge construction. Beam storage scheduling is a critical step in the railway precast beam yard production process, involving the storage, management, and transportation planning of precast beams.

[0042] Specifically, prior to step 100, the method further includes:

[0043] Step S100: Obtain the pre-established business logic model of the target beam yard, and construct the simulation operation logic of the business logic model to obtain the simulation model of the target beam yard;

[0044] The business logic model includes a beam fabrication area business logic model and a beam storage area business logic model; the beam fabrication area business logic model is used to represent the business situation of the beam fabrication area in the target beam yard, and the beam storage area business logic model is used to represent the construction situation of the beam storage area in the target beam yard.

[0045] Figure 2 This is a schematic diagram of the beam yard spatial planar modeling provided by the present invention. The following is in conjunction with... Figure 2 To provide a specific example, in one embodiment:

[0046] 1. Establish a business logic model for the beam fabrication area. Based on the business situation of the beam fabrication area in the beam yard, determine the data model, behavior model, and identification scheme for the beam fabrication area.

[0047] 1.1 Design of the beam fabrication area's morphological characteristics. The beam fabrication area is modeled using a matrix structure, consisting of inner molds, outer molds, and empty spaces. Each position is aligned longitudinally with the platform in the beam storage area. The beam fabrication area is configured with 3 sets of beam fabrication units, each set consisting of 2 inner molds and 4 outer molds. The finished product is output from the outer mold platform. The spatial modeling of the beam fabrication area is as follows: Figure 2 As shown.

[0048] 1.2 Design of beam fabrication area data and production time. The data for each beam in the beam fabrication area includes the beam number and the number of days remaining until completion; the production sequence of beams strictly follows the bridge erection sequence; a maximum of one beam from each unit can enter the production process each day, and the four outer membrane positions of each unit are used in rotation; the total production time for beams is 4 days, after which they can be sent to the beam storage area.

[0049] 2. Establish a business logic model for the beam storage area. Based on the construction status of the beam storage area in the beam yard, determine the data model, behavior model, and identification scheme for the beam storage area.

[0050] 2.1 Design of the beam storage area's morphological characteristics. The beam storage area is modeled using a matrix structure, with length and width parameters determined by the beam yard conditions. In this embodiment, it consists of 100 beam storage platforms arranged in 10 rows and 10 columns. The spatial modeling of the beam storage area is as follows: Figure 2 As shown, a maximum of 2 beams can be stacked on each pedestal, for a total maximum of 200 beams.

[0051] 2.2 Design of beam storage area data and production time. The data for each beam in the beam storage area includes beam number, remaining curing days, and remaining bridge erection days. Beams entering the beam storage area must first undergo a 14-day curing process, during which they cannot be moved or stacked. After the curing is completed, they must be stored in the beam storage area for a total of 30 days (including the curing time) before bridge erection can be carried out.

[0052] 2.3 Design special operating rules for the beam storage area. After the beam yard is full, if new beams continue to be stored, the beam storage area will trigger a beam relocation operation to stack the cured beams, thereby freeing up new beam storage platforms. When the bridge erecting machine turns around, all beams within a specific rectangular area in the beam storage area must be cleared. This operation needs to be planned in advance. When the beam yard is full and stacking is not possible, the beam fabrication area can be temporarily suspended for one day to alleviate the pressure on the beam storage area.

[0053] Further, in step S100, the simulation operation logic of the business logic model is constructed to obtain the simulation model of the target beam yard, including:

[0054] Step S110: Based on the business logic model, construct the business module of the target beam yard using object-oriented method; wherein, the business module includes a production module, a storage module and an erection module.

[0055] Step S120: Based on the behavior and interaction methods of each module of the target beam yard, determine and display the operation behavior parameters of the target beam yard in order to establish a simulation model of the target beam yard.

[0056] The following is a detailed description with reference to specific embodiments. Based on the above embodiments, in one embodiment:

[0057] 3. Simulation Logic Construction. Using an object-oriented approach, the production, storage, and erection business modules of the beam yard are constructed respectively. Based on the beam yard business design, the behavior and interaction methods of each module are defined, and the important parameters of the simulation process (i.e., the above-mentioned operational behavior parameters) are statistically analyzed and displayed.

[0058] 3.1 Beam Fabrication Module Behavior and Interface. The beam fabrication module is responsible for detecting the beam fabrication plan, adding beams to be produced to the production area according to the plan, monitoring beam production, counting completed beams, and submitting a beam storage application to the beam storage module; receiving stop signals from the beam storage area, and stopping production for 1 day for each signal received.

[0059] 3.2 Beam Storage Module Behavior and Interface. The beam storage module is responsible for monitoring and managing the timing operations such as curing of beams in the beam storage area; receiving beam storage requests from the beam fabrication module and processing beam storage operations, and moving and stacking beams when necessary; sending a stop command to the beam fabrication module when the beam storage area is nearly full to prevent the beam yard from becoming too full; receiving beam erection requests from the beam erection module, and being responsible for locating the beams to be erected and moving them to the bridge access area (e.g., Figure 2 (The row of beam storage areas closest to the railway bridge); receiving applications for turning around the beam erection modules and being responsible for clearing the designated area required for the bridge erection machine to turn around.

[0060] 3.3 Behavior and Interface of the Girder Erection Module. The girder erection module is responsible for detecting the girder erection plan, requesting girder erection from the girder storage module according to the plan; monitoring the erection mileage and direction of the girder, controlling the bridge erection speed, and requesting the bridge erection machine to turn around from the girder storage area when necessary.

[0061] 3.4 Control Module Behavior. The system cycles daily, calling relevant modules in the order of beam fabrication, beam storage, and beam erection to manage various common parameters during the simulation process, and to statistically analyze and display the beam yard's status and key simulation parameters.

[0062] Furthermore, the target beam yard also includes a beam-moving machine and a bridge-erecting machine.

[0063] After step S120, the method further includes:

[0064] Based on the beam yard shape of the target beam yard, motion rule modeling is performed on the beam moving machine and the bridge erecting machine.

[0065] The following is a detailed description with reference to specific embodiments. Based on the above embodiments, in one embodiment:

[0066] 4. Modeling the motion rules of the beam transfer machine. All movement of beams within the beam yard requires the beam transfer machine to complete. Based on the shape of the beam yard, the movement of the beam transfer machine has its specific rules.

[0067] 4.1 Beam Transfer Machine Movement Specifications. There is usually a lateral movement zone between the beam fabrication area and the beam storage area in the beam yard (sometimes between the beam storage area and the railway bridge). Any lateral movement of the beam transfer machine must be carried out within this lateral movement zone, but longitudinal movement is unrestricted. Figure 2 The numbers with circles in the middle mark four typical movement routes.

[0068] 4.2 Specifications for Beam Moving Machine Losses. The movement of the beam moving machine generates significant fuel consumption, which is a crucial optimization indicator in beam yard optimization. Due to its substantial weight, the fuel loss from moving the beam moving machine a specific distance unloaded is equivalent to 70% of the fuel loss from lifting beams.

[0069] 5. Modeling the motion rules of the bridge erecting machine. The bridge erecting operation is carried out by the bridge erecting machine, which has its own unique movement and turning methods, which have a significant impact on the beam yard and need to be modeled and described.

[0070] 5.1 Rules for Bridge Erection Machine Turning Around. The beam yard is usually built in the middle of the railway line. Bridge erection begins in one direction, and after completion, the bridge erecting machine turns around to proceed in the other direction until all work is completed. The bridge erecting machine can only turn around in the beam yard. Due to its large length, turning around requires clearing all beams from a specific area of ​​the storage area. Figure 2 The area enclosed by a 7*3 dotted frame in the middle of the beam storage area.

[0071] 5.2 Bridge Erection Speed ​​Rules. The number of bridge beams that the bridge erecting machine can erect per day depends on the distance between the erection location and the beam yard: 3 beams per day within 8km, and 2 beams per day within 8-15km. Considering the optimal beam storage at the beam yard, bridge erection operations are carried out approximately 30 days later than beam fabrication operations.

[0072] Step 200: Based on a predetermined objective function, optimize the operational behavior parameters using a deep neural network based on reinforcement learning, and optimize the beam storage schedule of the target beam yard based on the optimized operational behavior parameters.

[0073] It should be noted that the objective function includes: the moving distance of the beam transfer machine, the utilization rate of the beam yard, and the utilization rate of manpower.

[0074] Specifically, step 200 includes:

[0075] Based on the moving distance of the beam transfer machine, the utilization rate of the beam yard, and the utilization rate of manpower, a deep neural network based on reinforcement learning is used to optimize the operation behavior parameters, so as to optimize the beam storage schedule of the target beam yard based on the optimized operation behavior parameters.

[0076] Figure 3 This is a schematic diagram of the overall framework of reinforcement learning provided by the present invention. The following is in conjunction with... Figure 3 To provide a specific explanation, based on the above embodiments, in one embodiment:

[0077] 6. Establish a multi-objective optimization scheme. Using the total movement distance of the beam-moving machine, beam yard utilization rate, and labor utilization rate as objective functions, optimize and adjust the operational behavior parameters of each module during the simulation process to form a multi-objective optimized beam storage scheduling scheme. Figure 3 It demonstrates the overall framework of reinforcement learning.

[0078] 6.1 Beam Moving Distance. During the simulation, the distance of each beam movement is statistically analyzed. This includes the distance d1 from the beam moving machine to the beam starting position from its current unloaded position and the distance d2 from the beam lifting position to the beam target position. The sum of these two distances is the total moving distance for this beam movement. The total beam moving distance for all operations throughout the entire lifecycle of the beam yard is calculated and used as one of the objective functions.

[0079] 6.2 Beam Yard Utilization. A beam yard that is too large will lead to resource waste, while one that is too small will result in excessively frequent beam movement and stacking. To promote full utilization of the site area, the ratio between the daily beam storage quantity and the maximum beam storage quantity is statistically analyzed during the simulation and accumulated, serving as one of the objective functions.

[0080] 6.3 Labor Utilization Rate. The beam fabrication area may sometimes suspend or slow down its work due to the beam storage area being full, but the various expenses related to workers will not stop. During the simulation, the cumulative number of downtime days of the beam fabrication area throughout its entire life cycle is counted as one of the objective functions for measuring the working efficiency of the beam yard.

[0081] Specifically, step 200 also includes:

[0082] Step 210: Use the predetermined objective function as the main evaluation parameter and the operational behavior parameters of the target beam field as auxiliary evaluation parameters.

[0083] Step 220: Input the main evaluation parameters and the auxiliary evaluation parameters into a deep neural network based on reinforcement learning to obtain optimized parameters, and optimize the beam storage schedule of the target beam yard based on the optimized parameters.

[0084] The following is a detailed description with reference to specific embodiments. Based on the above embodiments, in one embodiment:

[0085] 7. Neural Network Model Construction. A deep neural network based on reinforcement learning is used as the optimization algorithm to construct a reinforcement learning framework covering four aspects: simulation engine, neural network, preference parameters, and optimization parameters.

[0086] 7.1 Preference Parameters. Preference parameters are the input to the simulation engine and the output of the neural network. In this embodiment, the neural network does not directly control the operation of a specific beam, but it adjusts the behavior of the beam storage module through preference parameters. There are 10 preference parameters that control various selection tendencies in the beam storage and movement processes, which can be found in [reference missing]. Figure 3 Preference parameters section.

[0087] 7.2 Evaluation Parameters. Evaluation parameters are the output of the simulation engine and also the input of the neural network. The neural network uses these evaluation parameters to determine whether the current preference parameter settings need adjustment. The evaluation parameters include the three main evaluation parameters from step 6, and four auxiliary evaluation parameters statistically derived during the simulation. See [link to relevant documentation]. Figure 3 Evaluation parameters section.

[0088] 7.3 Neural Network Model Design. The neural network in this embodiment has a total of 4 layers: an input layer with 7 neurons, a middle layer with 128 neurons, a second middle layer with 64 neurons, and an output layer with 10 neurons. The neurons are connected using the tansig activation function, such as... Figure 3 The neural network section is shown.

[0089] The above describes the steps of the reinforcement learning-based method for optimizing the storage and scheduling of precast railway beams in railway precast beam yards provided by this invention. As can be seen from the above description, the reinforcement learning-based method for optimizing the storage and scheduling of precast railway beams in railway precast beam yards provided by this invention obtains a pre-established simulation model of the target beam yard. The simulation model includes operational behavior parameters of the target beam yard, which represent the behavior and interaction methods of each module in the target beam yard. Based on a pre-determined objective function, a deep neural network based on reinforcement learning is used to optimize the operational behavior parameters, thereby optimizing the storage and scheduling of the target beam yard based on the optimized operational behavior parameters. Therefore, this invention establishes a corresponding simulation model, which is convenient for adjusting model parameters. Furthermore, the simulation process can be intervened at any time during operation, allowing for closer interaction with the beam yard's working conditions. Based on a pre-determined objective function, the deep neural network based on reinforcement learning is used to adjust the model behavior, overcoming the limitations of beam storage scheduling optimization technology under constraints from multiple resources such as piers, equipment, and personnel.

[0090] The following describes the reinforcement learning-based railway precast beam yard storage scheduling optimization device provided by the present invention. The reinforcement learning-based railway precast beam yard storage scheduling optimization device described below can be referred to in correspondence with the reinforcement learning-based railway precast beam yard storage scheduling optimization method described above.

[0091] Figure 4This is a schematic diagram of the reinforcement learning-based railway precast beam yard storage scheduling optimization device provided by the present invention, as shown below. Figure 4 As shown, the reinforcement learning-based railway precast beam yard storage scheduling optimization device provided by the present invention includes:

[0092] The acquisition unit 401 is used to acquire a pre-established simulation model of the target beam field; wherein, the simulation model includes the operation behavior parameters of the target beam field, and the operation behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field;

[0093] The optimization unit 402 is used to optimize the operation behavior parameters using a deep neural network based on reinforcement learning based on a predetermined objective function, so as to optimize the beam storage schedule of the target beam yard based on the optimized operation behavior parameters.

[0094] This invention provides a reinforcement learning-based device for optimizing beam storage scheduling in railway precast beam yards. It acquires a pre-established simulation model of the target beam yard. This simulation model includes operational behavior parameters representing the behavior and interaction of each module within the target beam yard. Based on a pre-determined objective function, a deep neural network based on reinforcement learning is used to optimize the operational behavior parameters, thereby optimizing the beam storage scheduling of the target beam yard. Therefore, this invention establishes a corresponding simulation model with convenient parameter adjustment. Furthermore, the simulation process can be intervened at any time during operation, allowing for closer interaction with the beam yard's operational status. By adjusting the model's behavior using a deep neural network based on reinforcement learning, based on a pre-determined objective function, this invention overcomes the limitations of beam storage scheduling optimization techniques constrained by resources such as platform size, equipment, and personnel.

[0095] Based on the above embodiments, in this embodiment, the device further includes a simulation unit, specifically used for:

[0096] Before obtaining the pre-established simulation model of the target beam yard, the pre-established business logic model of the target beam yard is obtained, and the simulation operation logic is constructed on the business logic model to obtain the simulation model of the target beam yard.

[0097] The business logic model includes a beam fabrication area business logic model and a beam storage area business logic model; the beam fabrication area business logic model is used to represent the business situation of the beam fabrication area in the target beam yard, and the beam storage area business logic model is used to represent the construction situation of the beam storage area in the target beam yard.

[0098] Based on the above embodiments, in this embodiment, the simulation unit is specifically used for:

[0099] Based on the aforementioned business logic model, an object-oriented approach is used to construct the business modules for the target beam yard; wherein, the business modules include a production module, a storage module, and an erection module;

[0100] Based on the behavior and interaction methods of each module of the target beam yard, the operation behavior parameters of the target beam yard are determined and displayed in order to establish a simulation model of the target beam yard.

[0101] Based on the above embodiments, in this embodiment, the target beam yard further includes a beam moving machine and a bridge erecting machine;

[0102] The device further includes a modeling unit, specifically used for:

[0103] After determining and displaying the operational behavior parameters of the target beam yard based on the behavior and interaction methods of each module of the target beam yard, and establishing a simulation model of the target beam yard, motion rule modeling is performed on the beam moving machine and the bridge erecting machine based on the beam yard morphology of the target beam yard.

[0104] Based on the above embodiments, in this embodiment, the objective function includes: beam moving machine movement distance, beam yard utilization rate, and labor utilization rate;

[0105] The optimization unit 402 is specifically used for:

[0106] Based on the moving distance of the beam transfer machine, the utilization rate of the beam yard, and the utilization rate of manpower, a deep neural network based on reinforcement learning is used to optimize the operation behavior parameters, so as to optimize the beam storage schedule of the target beam yard based on the optimized operation behavior parameters.

[0107] Based on the above embodiments, in this embodiment, the optimization unit 402 is specifically used for:

[0108] The predetermined objective function is used as the main evaluation parameter, and the operational behavior parameters of the target beam yard are used as auxiliary evaluation parameters.

[0109] The main evaluation parameters and the auxiliary evaluation parameters are input into a deep neural network based on reinforcement learning to obtain optimized parameters, and the beam storage schedule of the target beam yard is optimized based on the optimized parameters.

[0110] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions from the memory 530 to execute a reinforcement learning-based optimization method for precast beam storage scheduling in railway precast beam yards. This method includes:

[0111] Obtain a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field;

[0112] Based on a predetermined objective function, the operational behavior parameters are optimized using a deep neural network based on reinforcement learning, and the beam storage scheduling of the target beam yard is optimized based on the optimized operational behavior parameters.

[0113] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the reinforcement learning-based railway precast beam yard storage scheduling optimization method provided by the above methods, the method comprising:

[0115] Obtain a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field;

[0116] Based on a predetermined objective function, the operational behavior parameters are optimized using a deep neural network based on reinforcement learning, and the beam storage scheduling of the target beam yard is optimized based on the optimized operational behavior parameters.

[0117] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the reinforcement learning-based railway precast beam yard storage scheduling optimization method provided by the above methods, the method comprising:

[0118] Obtain a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field;

[0119] Based on a predetermined objective function, the operational behavior parameters are optimized using a deep neural network based on reinforcement learning, and the beam storage scheduling of the target beam yard is optimized based on the optimized operational behavior parameters.

[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning, characterized in that, include: Obtain a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field; Based on a predetermined objective function, the operational behavior parameters are optimized using a deep neural network based on reinforcement learning, and the beam storage scheduling of the target beam yard is optimized based on the optimized operational behavior parameters; before obtaining the pre-established simulation model of the target beam yard, the method further includes: Obtain a pre-established business logic model of the target beam yard, and construct simulation operation logic for the business logic model to obtain a simulation model of the target beam yard; The business logic model includes a beam fabrication area business logic model and a beam storage area business logic model. The beam fabrication area business logic model represents the business situation of the beam fabrication area in the target beam yard, and the beam storage area business logic model represents the construction situation of the beam storage area in the target beam yard. The step of constructing the simulation operation logic of the business logic models to obtain the simulation model of the target beam yard includes: Based on the aforementioned business logic model, an object-oriented approach is used to construct the business modules for the target beam yard; wherein, the business modules include a production module, a storage module, and an erection module; Based on the behavior and interaction methods of each module of the target beam yard, the operation behavior parameters of the target beam yard are determined and displayed to establish a simulation model of the target beam yard; the target beam yard also includes a beam moving machine and a bridge erecting machine; After determining and displaying the operational behavior parameters of the target beam yard based on the behavior and interaction methods of each module of the target beam yard to establish a simulation model of the target beam yard, the method further includes: Based on the beam yard shape of the target beam yard, motion rule modeling is performed on the beam moving machine and the bridge erecting machine.

2. The method for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning according to claim 1, characterized in that, The objective function includes: beam moving machine travel distance, beam yard utilization rate, and labor utilization rate; The step of optimizing the operational behavior parameters using a deep neural network based on reinforcement learning, based on a pre-determined objective function, to optimize the beam storage scheduling of the target beam yard based on the optimized operational behavior parameters, includes: Based on the moving distance of the beam transfer machine, the utilization rate of the beam yard, and the utilization rate of manpower, a deep neural network based on reinforcement learning is used to optimize the operation behavior parameters, so as to optimize the beam storage schedule of the target beam yard based on the optimized operation behavior parameters.

3. The method for optimizing the scheduling of precast beams in railway precast beam yards based on reinforcement learning according to claim 1, characterized in that, The step of optimizing the operational behavior parameters using a deep neural network based on reinforcement learning, based on a pre-determined objective function, to optimize the beam storage scheduling of the target beam yard based on the optimized operational behavior parameters, includes: The predetermined objective function is used as the main evaluation parameter, and the operational behavior parameters of the target beam yard are used as auxiliary evaluation parameters. The main evaluation parameters and the auxiliary evaluation parameters are input into a deep neural network based on reinforcement learning to obtain optimized parameters, and the beam storage schedule of the target beam yard is optimized based on the optimized parameters.

4. A reinforcement learning-based device for optimizing the scheduling of precast beams in railway prefabrication yards, characterized in that, include: An acquisition unit is used to acquire a pre-established simulation model of the target beam field; wherein, the simulation model includes operational behavior parameters of the target beam field, and the operational behavior parameters are used to represent the behavior and interaction mode of each module of the target beam field; An optimization unit is used to optimize the operational behavior parameters using a deep neural network based on reinforcement learning, based on a predetermined objective function, so as to optimize the beam storage schedule of the target beam yard based on the optimized operational behavior parameters. The device further includes a simulation unit, specifically used for: Before obtaining the pre-established simulation model of the target beam yard, the pre-established business logic model of the target beam yard is obtained, and the simulation operation logic is constructed on the business logic model to obtain the simulation model of the target beam yard. The business logic model includes a beam fabrication area business logic model and a beam storage area business logic model; the beam fabrication area business logic model is used to represent the business situation of the beam fabrication area in the target beam yard, and the beam storage area business logic model is used to represent the construction situation of the beam storage area in the target beam yard. The simulation unit is specifically used for: Based on the aforementioned business logic model, an object-oriented approach is used to construct the business modules for the target beam yard; wherein, the business modules include a production module, a storage module, and an erection module; Based on the behavior and interaction methods of each module of the target beam yard, the operation behavior parameters of the target beam yard are determined and displayed in order to establish a simulation model of the target beam yard; The target beam yard also includes beam moving machines and bridge erecting machines; The device further includes a modeling unit, specifically used for: After determining and displaying the operational behavior parameters of the target beam yard based on the behavior and interaction methods of each module of the target beam yard, and establishing a simulation model of the target beam yard, motion rule modeling is performed on the beam moving machine and the bridge erecting machine based on the beam yard morphology of the target beam yard.

5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards as described in any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards as described in any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the reinforcement learning-based method for optimizing the scheduling of precast beams in railway precast beam yards as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • BIM-based beam yard digital management system and method

    CN116843844A