Steel box girder three-dimensional model cost prediction method based on reinforcement learning
Through the method of reinforcement learning, deep mining and analysis of the three-dimensional model data of steel box girders is constructed, and the deep reinforcement learning model is solved, which is difficult to take into account the accuracy and efficiency of traditional cost estimation methods, and high-precision and intelligent cost prediction effects are achieved.
Patent Information
- Application Number
- CN202510152841.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Traditional steel box girder cost estimation methods are difficult to take into account both prediction accuracy and efficiency, especially when facing complex designs, they are prone to cost overruns or resource waste.
The cost prediction method of the three-dimensional model of steel box girder based on reinforcement learning is adopted, and the three-dimensional model data is deeply mined through intelligent algorithms, a deep reinforcement learning model is constructed, and a multi-dimensional information such as geometric features, material information and structural complexity is used to make high-precision cost prediction.
It significantly improves the accuracy and reliability of cost prediction, realizes intelligent and efficient cost prediction, can dynamically adjust parameters, gradually approach actual costs, and intuitively display high-cost areas through visualization tools to assist in engineering design optimization and resource allocation.
Smart Images

Figure CN119941340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to cost prediction technology in bridge engineering, and in particular to a cost prediction method for a three-dimensional steel box girder model based on reinforcement learning. Background Art
[0002] As an important load-bearing structure of bridges, steel box girders are widely used in modern bridge engineering due to their advantages such as light weight, high torsional rigidity and beautiful appearance. However, the manufacturing process of steel box girders is complex, involving a variety of materials, processes and components, and its manufacturing cost is affected by many factors such as geometric characteristics, material properties, and structural complexity. Traditional cost estimation methods mainly rely on empirical formulas or expert judgment. This method often fails to balance prediction accuracy and efficiency when faced with complex steel box girder designs, and is prone to cost overruns or waste of resources.
[0003] With the popularization of building information modeling and computer-aided design technology, three-dimensional digital models have gradually become an important tool for bridge engineering design and construction. The model contains rich data such as the geometric characteristics, material properties and connection relationships of steel box girders, which provides a data basis for cost prediction based on digital models. However, due to the diversity of data formats, high model complexity and nonlinear characteristics of cost influencing factors, how to use these data efficiently and accurately for cost prediction is still a difficult problem that needs to be solved.
[0004] In recent years, artificial intelligence technology, especially reinforcement learning algorithms, has demonstrated powerful learning and prediction capabilities in complex decision-making problems. By building an interactive model between the agent and the environment, reinforcement learning can dynamically adjust strategies and approach the optimal solution, which is very suitable for solving multi-dimensional and nonlinear problems. In the field of steel box girder cost prediction, combining reinforcement learning with three-dimensional digital models can fully mine the multi-dimensional information in the model and establish a high-precision cost prediction model, thereby providing a scientific basis for engineering design optimization and resource allocation.
[0005] Therefore, it is necessary to provide a cost prediction method for the three-dimensional model of steel box girders to deeply mine and analyze the three-dimensional model data through intelligent algorithms, improve the accuracy and efficiency of prediction, and solve the above problems. Summary of the invention
[0006] The present invention aims to provide a cost prediction method for a three-dimensional model of a steel box girder based on reinforcement learning. By deeply mining the data in the three-dimensional model through an intelligent algorithm, the multi-dimensional information such as geometric features, material information and structural complexity is fully utilized to achieve high-precision cost prediction.
[0007] In order to achieve the above-mentioned purpose, the present invention provides a cost prediction method for a three-dimensional model of a steel box girder based on reinforcement learning, the main contents of which include preprocessing the three-dimensional model data of the steel box girder, extracting and analyzing features, building and training a reinforcement learning model to optimize and evaluate the model, predicting the cost of the model, displaying the prediction results, and periodically updating the model training data.
[0008] The preprocessing of the three-dimensional model data is to convert the three-dimensional model data of the steel box girder from the original format, namely the BIM and CAD format, into a standardized neutral format, namely the IFC standard format, to ensure the integrity of the geometric information, material properties and connection relationships.
[0009] The extraction and analysis features are to extract the geometric features of the steel box girder, including plate thickness, size, and volume, extract material properties including material type, density, and tensile strength, extract structural complexity including weld length and node density, and then, according to the main structure, including top plate, bottom plate, and web plate, and according to the secondary structure, including U ribs and stiffeners, the model data is stored in a hierarchical manner in a relational database and a graph database to generate a multi-dimensional input feature vector for reinforcement learning.
[0010] The construction and training of the reinforcement learning model refers to using a deep reinforcement learning algorithm to construct a model, taking the characteristics of each component of the steel box girder as the state space and the cost prediction value of the component as the action space, and then designing a reward function to minimize the prediction error based on the difference between the actual manufacturing cost and the predicted cost. The reinforcement learning algorithm uses a deep Q learning (DQN) algorithm to train the model, optimizes the prediction path through experience replay, and makes the model gradually approach the actual cost.
[0011] The optimization and evaluation model mentioned above refers to using historical project data to train the model, generate prediction costs, and calculate error indicators through the validation set, including mean square error (MSE) and mean absolute error (MAE). If the error exceeds the threshold, the model parameters and reward function are adjusted to continuously optimize the model.
[0012] The key to cost prediction by the model is to input new three-dimensional model data of steel box girders into the trained model, output the predicted manufacturing cost of each component and the whole, superimpose the predicted results on the three-dimensional model through visualization tools, and intuitively display high-cost areas in the form of color coding to assist engineering design optimization and resource allocation. Based on the data fed back from actual projects, the model training data is periodically updated.
[0013] A hierarchical storage structure is used in the three-dimensional model data preprocessing process, including a component table, which stores the geometric characteristics, material information and complexity index of each component, a relationship table, which stores the connection relationship between components, including the connection type and connection position, and a feature vector table, which quantizes the features of all components to generate feature data for input into the reinforcement learning model.
[0014] The reinforcement learning model uses a deep Q learning algorithm, including constructing a deep neural network consisting of an input layer, a hidden layer and an output layer. The input layer receives a multidimensional feature vector of a steel box girder component, including geometric features, material properties and structural complexity. The output layer corresponds to the Q value of each possible action, which is used to guide the prediction of component costs. The target network is introduced and a fixed parameter update interval is maintained with the online network. The parameters of the target network are updated through the soft update rule of the online network. The formula is: , where θ is the online network parameter, θ′ is the target network parameter, and τ is the soft update factor. The sampled historical training data is stored in the experience pool. Training is performed by randomly sampling small batches of data to break the time correlation of data and improve training efficiency. The sampling probability is allocated according to the importance of historical samples. Samples with higher rewards or larger errors are trained first to improve the model's learning ability for key training data. The ε-greedy strategy is used to dynamically adjust the exploration rate ε. In the early stage, random actions are selected with a higher probability to increase exploration diversity. In the later stage, ε is gradually reduced, and high Q-value actions are selected to improve decision-making accuracy.
[0015] The design of the reward function aims to minimize the error between the actual cost and the predicted cost. The specific formula is: R = -|C 预测 -C 实际 |, where Cactual is the actual manufacturing cost and Cpredicted is the model predicted cost.
[0016] The following optimization steps are used during the model training process: Step 1: Use the experience replay mechanism to randomly sample historical data to prevent the model from falling into the local optimum due to data correlation during training. Step 2: Exploration and Utilization Mechanism: initially set a higher exploration rate to randomly explore different estimation paths, and then gradually reduce the exploration rate to improve the accuracy of cost prediction; Step three, hyperparameter adjustment, dynamically adjust the learning rate and discount factor so that the model can balance the learning speed and convergence stability in different training stages.
[0017] Compared with the related art, the present invention has the following beneficial effects: (1) The present invention conducts in-depth mining of multi-dimensional information such as geometric features, material information, and structural complexity of the three-dimensional model, and comprehensively analyzes the cost influencing factors of the steel box girder, which can more accurately reflect the actual manufacturing cost and significantly improve the accuracy and reliability of cost prediction; (2) The deep reinforcement learning (DQN) algorithm is used to take the characteristics of each component of the steel box girder as the state space and the cost prediction value of the component as the action space. The prediction path is optimized through the reward function. The model has self-learning ability and can dynamically adjust parameters to gradually approach the actual cost, thus realizing intelligent and efficient cost prediction. (3) Using hierarchical storage structure and feature vectorization technology, the geometric information, material properties and complexity indicators of steel box girder components are standardized and structured for storage, which facilitates data management and model input, improves data processing efficiency, and ensures the integrity and consistency of information; (4) Through continuous training and periodic updating of historical project data, the model has strong generalization ability, can adapt to the needs of different steel box girder design scenarios, maintain high prediction accuracy, and meet the diverse requirements in actual engineering applications; (5) Mechanisms such as experience replay, importance sampling, and dynamic adjustment of exploration rate are used to break the time correlation of training data, prioritize learning of key samples, improve model training efficiency and stability, and ensure rapid convergence in scenarios of different complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is an overall flow chart of a method for predicting cost of a three-dimensional steel box girder model based on reinforcement learning according to the present invention; Figure 2 A diagram showing the preprocessing process of the three-dimensional model data of the present invention; Figure 3 A diagram showing the process of training a reinforcement learning model for the present invention; Figure 4 Optimization and evaluation model process diagram for the present invention; Figure 5 A diagram of the application process of the model of the present invention for cost prediction; Figure 6 This is a schematic diagram of a hierarchical storage structure used in the data preprocessing process of the present invention; Figure 7 A schematic diagram of a deep Q learning algorithm used in the reinforcement learning model of the present invention; DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0020] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0021] like Figure 1-2 As shown, the preprocessing of the three-dimensional model data is to convert the three-dimensional model data of the steel box girder from the original format, namely the BIM and CAD format, into a standardized neutral format, namely the IFC standard format, to ensure the integrity of geometric information, material properties and connection relationships. Through feature extraction and analysis, the geometric features of the steel box girder are extracted, including plate thickness, size, and volume. The material properties are extracted, including material type, density, and tensile strength. The structural complexity is extracted, including weld length and node density. In the next step, the model data is hierarchically stored in relational databases and graph databases according to the main structure, including the top plate, bottom plate, and web plate, and also according to the secondary structure, including U ribs and stiffeners, to generate a multi-dimensional input feature vector for reinforcement learning.
[0022] like Figure 3 As shown, a reinforcement learning model is constructed and trained for steel box girders. The model is constructed using a deep reinforcement learning algorithm. The characteristics of each component of the steel box girder are used as the state space, and the cost prediction value of the component is used as the action space. A reward function is designed to minimize the prediction error based on the difference between the actual manufacturing cost and the predicted cost. The reinforcement learning algorithm uses a deep Q learning (DQN) algorithm to train the model, and optimizes the prediction path through experience replay, so that the model gradually approaches the actual cost.
[0023] like Figure 4 As shown in the figure, in the process of optimizing and evaluating the model, the historical project data is used to train the model. After the prediction cost is generated, the error indicators are calculated through the validation set, including the mean square error (MSE) and the mean absolute error (MAE). If the error exceeds the threshold, the model parameters and reward function are adjusted to continuously optimize the model so that it has strong generalization ability and the ability to adapt to different steel box girder designs.
[0024] like Figure 5 As shown in the figure, in the process of applying the model for cost prediction, new steel box girder three-dimensional model data is input into the trained model, and the manufacturing cost prediction values of each component and the whole are output. The prediction results are superimposed on the three-dimensional model through visualization tools, and high-cost areas are intuitively displayed in the form of color coding to assist engineering design optimization and resource allocation. Based on the data fed back from actual projects, the model training data is periodically updated to further improve the prediction accuracy.
[0025] like Figure 6 As shown in the figure, a hierarchical storage structure is used in the preprocessing of the 3D model data. There is a steel box girder component table, which stores the geometric characteristics, material information and complexity index of each component. There is also a steel box girder relationship table, which stores the connection relationship between components, including connection type and connection position. There is also a steel box girder feature vector table, which quantizes the features of all components to generate feature data for input into the reinforcement learning model.
[0026] like Figure 7 As shown, the reinforcement learning model uses a deep Q learning algorithm, including constructing a deep neural network consisting of an input layer, a hidden layer, and an output layer. The input layer receives a multidimensional feature vector of the steel box girder component, including geometric features, material properties, and structural complexity. The output layer corresponds to the Q value of each possible action, which is used to guide the prediction of the component cost. The target network is introduced to maintain a fixed parameter update interval with the online network to reduce the instability during the training process. The parameters of the target network are updated through the soft update rules of the online network. The formula is: , where θ is the online network parameter, θ′ is the target network parameter, and τ is the soft update factor. The sampled historical training data is stored in the experience pool. Training is performed by randomly sampling small batches of data to break the time correlation of data and improve training efficiency. The sampling probability is allocated according to the importance of historical samples. Samples with higher rewards or larger errors are trained first to improve the model's learning ability for key training data. The ε-greedy strategy is used to dynamically adjust the exploration rate ε. In the early stage, random actions are selected with a higher probability to increase exploration diversity. In the later stage, ε is gradually reduced, and high Q-value actions are selected to improve decision-making accuracy.
[0027] The design of the pre-training reward function aims to minimize the error between the actual cost and the predicted cost. The specific formula is: R = -|C 预测 -C 实际 |, where Cactual is the actual manufacturing cost and Cpredicted is the model predicted cost. The following optimization steps are used during training: Step 1: Use the experience replay mechanism to randomly sample historical data to prevent the model from falling into the local optimum due to data correlation during training. Step 2: Exploration and Utilization Mechanism: initially set a higher exploration rate to randomly explore different estimation paths, and then gradually reduce the exploration rate to improve the accuracy of cost prediction; Step three, hyperparameter adjustment, dynamically adjust the learning rate and discount factor so that the model can balance the learning speed and convergence stability in different training stages.
[0028] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the above description is only for explaining the principles of the present invention, and that various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention, and these changes and improvements all fall within the scope of the present invention claimed for protection.
[0029] The protection scope of the present invention is defined by the following claims and their equivalents.
Claims
1. A cost prediction method for a three-dimensional steel box girder model based on reinforcement learning, characterized in that: Preprocess the 3D steel box girder model data, extract and analyze features, build and train reinforcement learning models, optimize and evaluate models, make cost predictions for the models, display the prediction results, and periodically update model training data.
2. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 1, characterized in that: The preprocessing of the three-dimensional model data is to convert the three-dimensional model data of the steel box girder from the original format, namely the BIM and CAD format, into a standardized neutral format, namely the IFC standard format, to ensure the integrity of the geometric information, material properties and connection relationships.
3. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 1, characterized in that: The extraction and analysis features are to extract the geometric features of the steel box girder, including plate thickness, size, and volume, extract material properties including material type, density, and tensile strength, extract structural complexity including weld length and node density, and then, according to the main structure, including top plate, bottom plate, and web plate, and according to the secondary structure, including U ribs and stiffeners, the model data is stored in a hierarchical manner in a relational database and a graph database to generate a multi-dimensional input feature vector for reinforcement learning.
4. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 1, characterized in that: The construction and training of the reinforcement learning model refers to using a deep reinforcement learning algorithm to construct a model, taking the characteristics of each component of the steel box girder as the state space and the cost prediction value of the component as the action space, and then designing a reward function to minimize the prediction error based on the difference between the actual manufacturing cost and the predicted cost. The reinforcement learning algorithm uses a deep Q learning (DQN) algorithm to train the model, optimizes the prediction path through experience replay, and makes the model gradually approach the actual cost.
5. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 1, characterized in that: The optimization and evaluation model mentioned above refers to using historical project data to train the model, generate prediction costs, and calculate error indicators through the validation set, including mean square error (MSE) and mean absolute error (MAE). If the error exceeds the threshold, the model parameters and reward function are adjusted to continuously optimize the model.
6. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 1, characterized in that: The key to cost prediction by the model is to input new three-dimensional model data of steel box girders into the trained model, output the predicted manufacturing cost of each component and the whole, superimpose the predicted results on the three-dimensional model through visualization tools, and intuitively display high-cost areas in the form of color coding to assist engineering design optimization and resource allocation. Based on the data fed back from actual projects, the model training data is periodically updated.
7. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning as claimed in claim 2, characterized in that: A hierarchical storage structure is used in the three-dimensional model data preprocessing process, including a component table, which stores the geometric characteristics, material information and complexity index of each component, a relationship table, which stores the connection relationship between components, including the connection type and connection position, and a feature vector table, which quantizes the features of all components to generate feature data for input into the reinforcement learning model.
8. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning as claimed in claim 4, characterized in that: The reinforcement learning model uses a deep Q learning algorithm, including constructing a deep neural network consisting of an input layer, a hidden layer and an output layer. The input layer receives a multidimensional feature vector of a steel box girder component, including geometric features, material properties and structural complexity. The output layer corresponds to the Q value of each possible action, which is used to guide the prediction of component costs. The target network is introduced and a fixed parameter update interval is maintained with the online network. The parameters of the target network are updated through the soft update rule of the online network. The formula is: , where θ is the online network parameter, θ′ is the target network parameter, and τ is the soft update factor. The sampled historical training data is stored in the experience pool. Training is performed by randomly sampling small batches of data to break the time correlation of data and improve training efficiency. The sampling probability is allocated according to the importance of historical samples. Samples with higher rewards or larger errors are trained first to improve the model's learning ability for key training data. The ε-greedy strategy is used to dynamically adjust the exploration rate ε. In the early stage, random actions are selected with a higher probability to increase exploration diversity. In the later stage, ε is gradually reduced, and high Q-value actions are selected to improve decision-making accuracy.
9. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 4, characterized in that: The design of the reward function aims to minimize the error between the actual cost and the predicted cost. The specific formula is: R = -|C 预测 -C 实际 |, where Cactual is the actual manufacturing cost and Cpredicted is the model predicted cost.
10. The method for predicting the cost of a three-dimensional steel box girder model based on reinforcement learning according to claim 4, characterized in that: The following optimization steps are used during the model training process: Step 1: Use the experience replay mechanism to randomly sample historical data to prevent the model from falling into the local optimum due to data correlation during training. Step 2: Exploration and Utilization Mechanism: initially set a higher exploration rate to randomly explore different estimation paths, and then gradually reduce the exploration rate to improve the accuracy of cost prediction; Step three, hyperparameter adjustment, dynamically adjust the learning rate and discount factor so that the model can balance the learning speed and convergence stability in different training stages.
Citation Information
Patent Citations
Federal learning-based steel price prediction method and device, equipment and medium
CN113112307A
Engineering cost management system and method based on BIM technology
CN117670400A
Highway bridge engineering cost prediction method and system
CN117708963A
Building project cost control method and device based on BIM
CN118365278A
Mold cost analysis method, system and equipment and storage medium
CN118552268A