Intelligent construction method based on data driving
By using support vector machines and stacked integrated learning strategies in intelligent construction projects, combined with non-dominant sorting genetic algorithms, the multi-objective construction model is optimized, and the problem of inefficient multi-objective optimization in intelligent construction project management is solved, achieving more accurate predictions and more effective management decisions.
Patent Information
- Application Number
- CN202510143363.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
There is a problem of low multi-objective optimization efficiency in intelligent construction project management. Traditional methods are difficult to systematically weigh management indicators such as project progress, quality, cost, and safety, and a large amount of multi-source heterogeneous data cannot be effectively converted into management knowledge.
Using a data-driven intelligent construction method, a multi-objective prediction model is constructed through the support vector machine algorithm, combining stacked ensemble learning strategies and non-dominant sorting genetic algorithms, the multi-objective model is optimized and solved, and the optimal model parameter combination takes into account different management goals.
It improves the prediction accuracy of project management indicators, achieves effective trade-offs between multiple goals, and improves the digital and intelligent level of project management.
Smart Images

Figure CN120013484A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent construction, and in particular to a data-driven intelligent construction method. Background Art
[0002] As the scale and complexity of construction projects continue to increase, the traditional project management model can no longer meet the development needs of the modern construction industry. In order to improve the level of construction project management, the concept of intelligent construction has emerged in recent years. Through the application of emerging technologies such as BIM, Internet of Things, and big data, it realizes the informatization, digitization, and intelligent management of the entire life cycle of engineering projects, so as to improve project implementation efficiency and ensure construction quality and safety.
[0003] However, in the practice of intelligent construction project management, there are still potential conflicts and balance issues between multiple management objectives. Management indicators such as project progress, quality, cost, and safety affect each other and change dynamically. Optimizing a single goal may harm the performance of other goals, making it difficult to maximize the overall benefits of the project. Traditional project management methods mainly rely on the experience and qualitative decision-making of managers, lack quantitative analysis and scientific optimization methods, and are difficult to make systematic trade-offs among multiple goals.
[0004] At the same time, smart construction projects generate a large amount of multi-source heterogeneous management data, including BIM model parameters, construction site monitoring data, collaborative information of all participants, etc. These data contain important characteristics and inherent laws of the project implementation process, and are of great value for accurately assessing project status, predicting management risks, and optimizing control measures. However, in actual applications, due to the lack of effective data fusion and mining technology, a large amount of data has not been converted into useful management knowledge, and the digitalization and intelligence level of project management needs to be further improved. Summary of the invention
[0005] In view of the low efficiency of multi-objective optimization in intelligent construction project management in the prior art, this application provides a data-driven intelligent construction method, which constructs a multi-objective prediction model through a support vector machine algorithm, and combines the prediction results of each target using a stacked ensemble learning strategy, thereby improving the prediction accuracy of project management indicators. In the optimization solution of the multi-objective model, a non-dominated sorting genetic algorithm and a simulated annealing strategy are introduced to obtain the optimal model parameter combination that takes into account different management objectives in the form of a Pareto optimal solution set.
[0006] The present application provides a data-driven intelligent construction method, including: collecting multi-source heterogeneous data of intelligent construction projects; the multi-source heterogeneous data includes environmental data, material and equipment data, personnel organization data and construction process data; using a recursive feature elimination (RFE) algorithm to perform feature selection on the collected multi-source heterogeneous data to obtain a feature subset reflecting the project progress, quality and safety; based on the feature subset, a multi-objective model for predicting the project progress, quality and safety is constructed through a support vector machine (SVM) algorithm; using a non-dominated sorting genetic algorithm (NSGA-II) to solve the multi-objective model, introducing a simulated annealing strategy in the iterative search process of NSGA-II to jump out of the local optimal solution, and outputting an optimized Pareto optimal solution set; selecting an optimal parameter combination based on the Pareto optimal solution set, using the optimal parameter combination to adjust the BIM model of the current intelligent construction project, and using the adjusted BIM model to guide construction.
[0007] Furthermore, a recursive feature elimination (RFE) algorithm is used to perform feature selection on the collected multi-source heterogeneous data to obtain a feature subset, including: converting the collected multi-source heterogeneous data into a structured data set; extracting indicators reflecting project progress, quality and safety based on the structured data set to construct an original feature set; using a recursive feature elimination (RFE) algorithm to screen the original feature set; using a sigmoid function to perform nonlinear transformation on the screened features, and outputting a standardized feature subset.
[0008] Furthermore, the recursive feature elimination (RFE) algorithm is used to screen the original feature set, including: using a feature importance evaluation method based on a decision tree to calculate the information gain of each feature in the original feature set as the feature importance weight; sorting the original features in descending order according to the feature importance weight, removing one or a group of features with the lowest weight, and obtaining a feature subset after dimensionality reduction; using the feature subset after dimensionality reduction as input to construct a support vector machine (SVM) classifier, using the K-fold cross-validation method to divide the training set and the validation set, evaluating the classification accuracy of the SVM classifier on the validation set, and using the classification accuracy as the quality score of the feature subset; repeatedly executing the quality score evaluation step to obtain the feature subset with the highest quality score in each round of iteration, and when the change in the optimal quality score of multiple consecutive rounds of iterations is less than a threshold, stopping the iteration and outputting the optimal feature subset.
[0009] Furthermore, based on the feature subset, a multi-objective model for predicting project progress, quality and safety is constructed by support vector machine (SVM) algorithm, including: dividing the feature subset into a training set, a validation set and a test set; wherein the training set is used for parameter learning of the SVM regression model, the validation set is used for hyperparameter optimization of the SVM regression model, and the test set is used for performance evaluation of the SVM regression model; independent SVM regression sub-models for project progress, quality and safety are constructed respectively, and radial basis kernel function is used as the kernel function of the SVM regression sub-model, and the hyperparameters of each SVM regression sub-model are optimized by grid search and K-fold cross-validation methods; the hyperparameters include penalty coefficients and kernel function parameters; based on the optimized SVM progress prediction sub-model, SVM quality prediction sub-model and SVM safety prediction sub-model, a stacking-based ensemble learning strategy is adopted to construct a multi-objective prediction model.
[0010] Furthermore, according to the optimized SVM progress prediction sub-model, SVM quality prediction sub-model and SVM safety prediction sub-model, a stacking-based ensemble learning strategy is adopted to construct a multi-objective prediction model, including: taking the SVM progress prediction sub-model, SVM quality prediction sub-model and SVM safety prediction sub-model as base learners, using the validation set to train each sub-model respectively, and evaluating the prediction performance of each sub-model on the test set; taking the prediction results of each SVM sub-model on the test set as new training data, using the real progress, quality and safety indicators of the project as labels, training the meta-learner, using the decision tree algorithm to construct the meta-learner, and optimizing the hyperparameters of the meta-learner through grid search and K-fold cross validation; combining the base learner and the meta-learner in a hierarchical structure to form a complete multi-objective prediction model; wherein the base learner outputs the preliminary prediction results of each target, and the meta-learner performs nonlinear combination on the preliminary prediction results to output the final project progress, quality and safety prediction values.
[0011] Furthermore, the non-dominated sorting genetic algorithm NSGA-II is used to solve the multi-objective model, including: according to the established multi-objective prediction model, the penalty coefficient C of the base learner SVM, the kernel function parameter γ, and the tree depth d, node splitting criterion s and the minimum number of leaf node samples m of the meta-learner decision tree are used as optimization variables , forming a D-dimensional decision space; selecting the mean absolute percentage error MAPE(x) of the multi-objective prediction model on the training set, the root mean square error RMSE(x) on the test set, and the time complexity T(x) of the base learner and meta learner as the optimization objectives, and constructing the multi-objective function: , x∈D; the established multi-objective function F(x) is solved by using the non-dominated sorting genetic algorithm NSGA-II to obtain the Pareto optimal solution set P; during the NSGA-II iteration process, a simulated annealing operation is introduced every t generations to jump out of the local optimal solution; the Pareto optimal solution set P obtained after iterative optimization is output, and each optimal solution in the solution set P is The corresponding set of parameter combinations , substitute into the built multi-objective prediction model.
[0012] Furthermore, the established multi-objective function F(x) is solved by using the non-dominated sorting genetic algorithm NSGA-II to obtain the Pareto optimal solution set P, including: The code is a binary string, and an initial population of size N is randomly generated, where each individual corresponds to a feasible solution of the multi-objective function F(x); the objective function values MAPE(x), RMSE(x) and T(x) of each individual in the initial population are calculated, and the initial population is quickly non-dominated sorted according to the dominance relationship of the objective function values to obtain the Pareto level and crowding distance of the individuals, and N individuals are selected based on the crowding distance to form the first-generation parent population P0; where the dominance relationship is determined by comparing the advantages and disadvantages of the individuals in each objective function; for the individuals in the t-th generation parent population Pt, two of them are randomly formed into a pair of parent individuals , for each pair of paternal Perform adaptive simulated binary crossover operation to obtain offspring individuals , there are N offspring individuals in total, forming a offspring population ; for the offspring population Each individual in , perform directional polynomial mutation to obtain the mutated offspring individuals , the parent population and the mutated offspring population Merge to form a new population of size 2N ; for populations Repeat the fast non-dominated sort to obtain a non-dominated solution set ;from At the beginning, select individuals from each solution set in turn to join the next generation population , until the number of individuals reaches N; if the last layer of solution set The number of individuals exceeds , then in Select the first layer with the largest crowding distance. Individuals incorporated , and get the next generation population of size N . Iterate until the preset maximum number of generations is reached , output the final population The corresponding non-dominated solution set P; the solution set P is the Pareto optimal solution set of the multi-objective function F (x), and each optimal solution in the set Represents a set of optimized parameter combinations of multi-objective prediction models.
[0013] Furthermore, for the t-th generation parent population The individuals in the , for each pair of paternal Perform adaptive simulated binary crossover operation to obtain offspring individuals , including: Setting adaptive crossover probability : ,in, and are the upper and lower bounds of the crossover probability, t is the current iteration number, is the maximum number of iterations; for the parent population Each pair of individuals , generate a random number r in [0, 1]: if , then for this pair of paternal individuals Perform a simulated binary crossover to generate two offspring individuals ; Otherwise, directly copy the pair of parent individuals as offspring individuals, that is, .
[0014] The generation and optimization of population individuals are achieved through adaptive simulation of binary crossover and polynomial mutation. Adaptive simulation of binary crossover generates new offspring individuals in continuous space by simulating the single-point crossover process of binary coding. The crossover probability is adaptively adjusted according to the evolutionary algebra, maintaining a high crossover probability in the early stage of evolution to increase population diversity, and reducing the crossover probability in the later stage of evolution to maintain the convergence of the population. Polynomial mutation introduces new search directions and diversity by perturbing the decision variables of individuals, preventing the algorithm from converging to the local optimum too early. The mutation amplitude is controlled by polynomial distribution, so that the mutated individuals search near the original individuals, balancing the ability of local search and global exploration. Through the synergy of crossover and mutation, the NSGA-II algorithm can effectively explore and optimize the solution space of multi-objective problems and generate high-quality Pareto optimal solution sets. At the same time, the parameter settings in the algorithm, such as crossover probability, mutation probability, distribution index, etc., can be adjusted according to the characteristics and needs of the problem to balance the convergence speed of the algorithm and the quality of the solution.
[0015] Furthermore, for the offspring population Each individual in , perform directional polynomial mutation to obtain the mutated offspring individuals , including: Each component , with probability Perform polynomial mutation; in polynomial mutation, if the individual components Selected for mutation, based on the corresponding component of the parent individual Introducing the preference factor λ: ,in, For the parent individual The value of component i, k is the power index of the preference factor, is a symbolic function; the preference factor λ is introduced into the perturbation process of polynomial variation, and the individual components are updated as follows: , where δ is the random perturbation of the multinomial distribution; The components of Mutate in sequence and finally obtain the mutated offspring individuals .
[0016] Through directional polynomial mutation, a preference factor is introduced on the basis of standard polynomial mutation, so that the mutation process has a certain directionality. The preference factor guides the mutation direction according to the relative position of the offspring individuals and the parent individuals in the decision space. When the offspring individuals deviate greatly from the parent individuals in a certain component, the effect of the preference factor is more significant, prompting the mutation to proceed in the direction away from the parent individuals, increasing the diversity and exploration ability of the population. At the same time, the power exponent k of the preference factor controls the preference intensity. The larger the k value, the stronger the preference effect and the more obvious the mutation direction. By adjusting the k value, the exploration and utilization capabilities of the algorithm can be balanced. The parent population and the mutated offspring population are merged to form a new population of size 2N. The merged population contains excellent individuals of the parent generation and new individuals of the offspring, reflecting the elite retention strategy. The new population will be subject to selection, crossover and mutation operations in the next generation of evolution, continuously optimized and evolved, and finally converged to the Pareto optimal solution set of the problem. By merging the parent population and the child population, the NSGA-II algorithm maintains the quality and diversity of the solution during population evolution, avoids premature convergence, and improves the global search capability and convergence speed of the algorithm.
[0017] Furthermore, a simulated annealing strategy is introduced in the iterative search process of NSGA-II to escape from the local optimal solution, including: setting the initial temperature of simulated annealing and cooling coefficient α, let the current temperature ; In each iteration of NSGA-II, a new population is obtained Afterwards, from Randomly select an individual from the non-dominated solution set , for its decision variables Perturb and get new individuals ; The perturbation method is to add random noise to each component of x; calculate the new individual The objective function value of , and the current individual The objective function value of According to the Metropolis criterion, if Dominate , then accept the new solution ,use replace Otherwise, with probability Accept inferior solution ; Update temperature , if T is lower than the preset threshold , then stop simulated annealing; otherwise, return to continue iteration; the new solution obtained by simulated annealing Joining the next generation population of NSGA-II As a new search starting point, continue to execute the iterative process in step S42 until the maximum number of iterations is reached. .
[0018] By introducing simulated annealing operation in the NSGA-II iteration process, the global search ability of the algorithm can be improved and the local optimal solution can be jumped out. By accepting inferior solutions, simulated annealing allows the algorithm to accept some temporary inferior solutions during the search process, so as to have the opportunity to jump out of the local optimal area and explore a wider solution space. As the temperature decreases, the probability of accepting inferior solutions gradually decreases, the algorithm gradually tends to local search, and finally converges to the global optimal solution. After each round of NSGA-II iteration, the execution frequency of simulated annealing operation can be controlled by judging whether the current generation needs to perform simulated annealing. Randomly selecting individuals from the non-dominated solution set as the current solution can increase the randomness and diversity of the search. Perturbing the decision variables can generate new solutions in the neighborhood of the current solution and expand the search range. By comparing the objective function values of the new solution and the current solution, and deciding whether to accept the new solution according to the Metropolis criterion, a balance can be achieved between the quality and diversity of the solution. The temperature update and the judgment of the termination condition control the convergence speed and termination timing of the simulated annealing process. By adding the new solution obtained by simulated annealing to the next generation population of NSGA-II, the search results of simulated annealing can be combined with the evolution process of NSGA-II to improve the overall performance of the algorithm. By continuing to execute NSGA-II iterations, optimization can be continued at the new search starting point until the maximum iteration number is reached or other termination conditions are met. The combination of simulated annealing and NSGA-II can introduce the global search capability of simulated annealing while maintaining the multi-objective optimization and non-dominated sorting characteristics of NSGA-II, thereby improving the performance and efficiency of the algorithm on complex multi-objective optimization problems.
[0019] Compared with the prior art, the advantages of this application are: The recursive feature elimination algorithm is used to select features from the collected multi-source heterogeneous data, and the optimal feature subset reflecting the project progress, quality and safety is obtained through iterative search. This method uses feature importance weights to characterize the correlation between the original features and the project management goals, and combines the classification performance of the support vector machine to evaluate and screen the feature subsets, quantifying the advantages and disadvantages of the features from two levels: data dimension and machine learning performance, which not only reduces the complexity of the data, but also ensures the prediction accuracy of the subsequent model.
[0020] The support vector machine uses kernel functions to map the original features to high-dimensional space, and realizes nonlinear regression of project management indicators through the maximum margin hyperplane. Compared with traditional statistical prediction methods, the support vector machine can capture complex patterns in data and has strong nonlinear fitting and generalization capabilities. At the same time, the model hyperparameters are optimized through grid search and cross-validation, further improving the accuracy and stability of multi-objective prediction.
[0021] Aiming at multiple project management objectives, a stacking-based ensemble learning strategy is proposed, which realizes the fusion of prediction results of each objective through the combined optimization of two-layer learners. In the base learner layer, independent support vector machine sub-models are constructed for progress, quality and safety to give full play to the prediction advantages of each sub-model; in the meta-learner layer, the decision tree algorithm is used to nonlinearly combine the outputs of each sub-model, and the meta-learner parameters are optimized to adapt to the association constraints between different objectives. Stacking integration can comprehensively utilize the complementary information of various project management indicators, balance the prediction bias and variance, and obtain more comprehensive and accurate multi-objective prediction results than a single model.
[0022] In the multi-objective model optimization, the non-dominated sorting genetic algorithm NSGA-II is used to search for the optimal parameter combination. NSGA-II introduces non-dominated sorting and crowding distance, constructs an elite retention mechanism, optimizes and balances various objectives simultaneously during population evolution, overcomes conflicts between objectives, and obtains a Pareto frontier composed of multiple non-dominated solutions. By retaining non-dominated solutions at different levels, NSGA-II provides multiple solutions for the trade-off between progress, quality, and safety, providing a flexible and optional set of alternatives for project decision-making.
[0023] In order to escape from the local optimal solution, the simulated annealing strategy is introduced in the NSGA-II iteration process to perform probabilistic perturbations on the dominant solution individuals. Simulated annealing guides the algorithm to accept inferior solutions with a certain probability through the Metropolis criterion, expands the search range of the solution space, and controls the perturbation amplitude through annealing and cooling, and gradually converges to the global optimal area in the later search. The combination of simulated annealing and NSGA-II enhances population diversity, improves search efficiency, and finally obtains a high-quality Pareto solution set that takes into account multiple objectives. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein: Figure 1 is an exemplary flow chart of a data-driven intelligent construction method according to the present application; Figure 2 is an exemplary flow chart of obtaining a standardized feature subset according to the present application; Figure 3 is an exemplary flow chart for constructing a stacked ensemble learning model according to the present application; Figure 4 is an exemplary flow chart for obtaining a Pareto optimal solution set according to the present application. DETAILED DESCRIPTION
[0025] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0026] Figure 1 This is an exemplary flow chart of a data-driven intelligent construction method shown in the present application, including: collecting multi-source heterogeneous data of intelligent construction projects; the multi-source heterogeneous data includes environmental data, material and equipment data, personnel organization data and construction process data; using a recursive feature elimination (RFE) algorithm to perform feature selection on the collected multi-source heterogeneous data to obtain a feature subset reflecting the project progress, quality and safety; based on the feature subset, a multi-objective model for predicting project progress, quality and safety is constructed through a support vector machine (SVM) algorithm; using a non-dominated sorting genetic algorithm (NSGA-II) to solve the multi-objective model, introducing a simulated annealing strategy in the iterative search process of NSGA-II to jump out of the local optimal solution, and outputting an optimized Pareto optimal solution set; S5, selecting an optimal parameter combination based on the Pareto optimal solution set, using the optimal parameter combination to adjust the BIM model of the current intelligent construction project, and using the adjusted BIM model to guide construction.
[0027] Collect multi-source heterogeneous data of intelligent construction projects; multi-source heterogeneous data includes environmental data, material and equipment data, personnel organization data and construction process data; collect meteorological parameters such as temperature, humidity, wind speed, rainfall, light intensity, and environmental pollution indicators such as PM2.5 and noise at the construction site in real time through environmental monitoring sensors. At the same time, connect to the meteorological bureau platform to obtain weather forecast information for the next period of time. Environmental data reflects the changes in the external environment of the project. Building materials are identified and tracked by Internet of Things technologies such as RFID tags and QR codes, and intelligent terminals such as barcode scanners and RFID readers are used to collect information such as material specifications, quantities, on-site acceptance, and inventory status. For large construction equipment such as tower cranes and excavators, the online monitoring system collects operating parameters such as equipment working hours, fuel consumption levels, and fault alarms to evaluate equipment performance. Material and equipment data reflects the resource allocation and utilization efficiency of the project, which is directly related to engineering quality and cost control. Project management personnel, work teams, professional subcontractors and other organizational entities use mobile apps and smart wearable devices to collect information such as personnel attendance, working hours, operation behavior, labor intensity, etc., and combine positioning technologies such as face recognition and indoor GPS to monitor the real-time location and movement trajectory of personnel on site. Through the OA collaborative platform, business information such as material requirements, design changes, technical briefings, and safety training of all parties are obtained. Personnel organization data reflects the behavioral laws and management needs of all parties involved in the project, and is the basis for optimizing project organization and processes. The BIM model covering the entire life cycle of project design, construction, and operation and maintenance is the main source of construction process data. Engineering information such as component geometric parameters, material properties, schedules, and adjacent logic are extracted from the BIM model. At the same time, through laser scanners, smart cameras and other equipment, the actual completion of the building is reconstructed in three dimensions to obtain point cloud data reflecting the construction process. Computer vision algorithms are used to analyze on-site images and video streams to identify scene elements such as construction procedures, machine operation, and personnel behavior. The construction process data dynamically records the entire process of project implementation and is the core basis for evaluating project progress and deviations.
[0028] Figure 2According to the exemplary flow chart of obtaining standardized feature subsets shown in this application, the recursive feature elimination (RFE) algorithm is used to select features from the collected multi-source heterogeneous data to obtain feature subsets reflecting project progress, quality and safety; the collected multi-source heterogeneous data is converted into a structured data set; after obtaining the structured data set of the intelligent construction project, the original feature set reflecting the key indicators of project management is constructed by the feature engineering method. First, based on the knowledge of project management and data analysis experience, indicators closely related to project progress, quality and safety are extracted as candidate features, such as progress deviation, material qualification rate, equipment failure rate, number of personnel violations, etc. Then, these indicators are processed and represented by features, qualitative indicators are quantified into numerical variables, and quantitative indicators are normalized or standardized to form a normalized feature vector. The original feature set contains all potential factors that affect project management goals in massive high-dimensional data. In order to screen out the key features with the largest amount of information and the most relevant to the management goals from the original feature set, a feature selection algorithm based on recursive feature elimination (RFE) is adopted. RFE obtains the optimal feature subset by recursively training the model and deleting features with small weight coefficients.
[0029] The feature importance evaluation method based on decision tree is used to measure the contribution of each feature in the original feature set to the project management goal. The decision tree recursively selects the optimal splitting feature, divides the data set into different subsets, and constructs a tree-like classification or regression model. During the training process of the decision tree, the information gain ratio or Gini coefficient of each feature in the splitting process can be calculated as a quantitative indicator of feature importance. Information gain ratio: measures the degree to which a feature reduces the uncertainty of data set classification. For the original feature set D and feature A, D is divided into multiple subsets according to the value of A. , calculate the information entropy of each subset , then the information gain ratio of feature A is: ; Where Gain(D, A) is the information gain of feature A on data set D, and IV(A) is the intrinsic value of feature A. Information gain ratio penalizes the number of feature values and can better reflect the actual contribution of the feature. Gini coefficient: measures the degree of influence of features on the classification purity of the data set. For the binary classification problem, the Gini coefficient of data set D is defined as: ;in, is the proportion of the i-th class samples in D. The Gini coefficient reflects the impurity of the data set. The Gini coefficient of feature A is the weighted average of the Gini coefficients of each subset: ; The smaller the Gini coefficient, the better the splitting effect of the feature on the data set and the higher the purity.
[0030] Feature sorting: Sort the feature importance weight vector in descending order, with features with larger weights in front and features with smaller weights in the back. The sorted feature sequence reflects the relative importance of each feature and provides a reference for feature elimination. Recursive elimination: According to the sorted feature sequence, remove several features with the smallest weights each iteration to obtain a feature subset after dimensionality reduction. Different elimination strategies can be used, such as removing a fixed number of features each time, removing a fixed proportion of features each time, and removing features with weights less than a threshold each time.
[0031] The feature subset after dimensionality reduction is used as input to construct a support vector machine (SVM) classifier to evaluate the predictive performance of the feature subset for project management goals. The project management data set is divided into a feature matrix X and a target vector y. The feature matrix X contains the feature subset after dimensionality reduction, each row represents a sample, and each column represents a feature. The target vector y contains the corresponding project management target label (such as progress, quality, safety level). The feature matrix X is standardized to make the value range of each feature similar to avoid affecting the model performance due to large differences in feature scales. Commonly used standardization methods include minimum-maximum standardization and Z-score standardization. According to the characteristics of project management data, an appropriate kernel function is selected. Commonly used kernel functions include linear kernel, polynomial kernel and Gaussian kernel (RBF kernel). The linear kernel is suitable for linearly separable data, and the polynomial kernel and Gaussian kernel are suitable for nonlinear data. The Gaussian kernel can control the nonlinearity of the kernel function by adjusting the parameter γ, and has good flexibility.
[0032] Specifically, in this embodiment, a Gaussian kernel function is selected, and the form of the Gaussian kernel function is: , where x and y are two sample vectors, and γ is a parameter that controls the width of the kernel function. By adjusting the value of γ, the degree of nonlinearity of the Gaussian kernel function can be changed. The larger the γ value, the higher the nonlinearity of the kernel function and the more complex the decision boundary; the smaller the γ value, the lower the nonlinearity of the kernel function and the smoother the decision boundary. Selecting an appropriate γ value enables the SVM classifier to achieve good classification performance on the training data while avoiding overfitting. The objective function of SVM consists of two parts: a penalty term for classification errors and a maximization term for the support vector interval. The penalty term for classification errors uses the hinge loss function, and for each sample ,in is the eigenvector, is the corresponding label (+1 or -1), and the hinge loss is defined as: max(0, 1 - y_i(w^T x_i + b)), where w is the normal vector of the hyperplane and b is the bias term. The maximization term of the support vector interval is , the goal is to minimize , to obtain the maximum interval. The regularization parameter C is introduced to balance the classification error penalty and the interval maximization, and the objective function is: . Transform the objective function into a dual problem and introduce the Lagrange multiplier α to obtain the dual problem: ,in is a Gaussian kernel function that satisfies the constraints: and . Use the gradient descent algorithm to solve the dual problem and update the value of α. Initialize α to 0, and then iterate the following steps until convergence: Calculate the gradient: ; Update α: , where η is the learning rate; projecting α to the constraints: ; After obtaining the optimal α value, calculate the hyperplane parameters w and b: , . Randomly divide the data set into K subsets of similar size. Perform K iterations, each time selecting a subset as the validation set and the remaining K-1 subsets as the training set. In each iteration, use the training set to train the SVM classifier and evaluate the classification performance on the validation set, such as accuracy, precision, recall, and F1 score. Record the performance indicators of each iteration and calculate the average performance of K iterations as the overall quality score of the current feature subset and hyperparameter combination. Use the average performance indicator obtained by cross-validation to evaluate the prediction performance of the SVM classifier on the current feature subset. Use grid search or random search to find the optimal combination in the hyperparameter space (such as C, γ) to further improve the model performance. Select the hyperparameter combination with the best performance, retrain the SVM classifier on all training data, and obtain the final model. For new project management samples, extract the same feature subset and use the trained SVM classifier for prediction. Based on the prediction results, evaluate and make decisions on project management goals, such as progress warning, quality control, safety level judgment, etc.
[0033] Each round of iteration starts from the optimal feature subset of the previous round, sorts and recursively eliminates the features according to their importance weights, and obtains several new feature subsets. For each new feature subset, an SVM classifier is constructed and K-fold cross-validation is performed, and the classification accuracy is calculated as the quality score. The quality scores of all feature subsets are compared, and the subset with the highest score is selected as the optimal feature subset for the current round. The iterative optimization process continuously updates the optimal feature subset through multiple rounds of feature selection and performance evaluation. When the increase in the optimal quality score of several consecutive rounds of iterations is less than a certain threshold, it is considered that the feature subset has converged and reached the optimal state, and the iterative process is stopped. At this time, the optimal feature subset of the last round is output as the key feature of project management for subsequent multi-objective optimization.
[0034] Each feature in the filtered feature subset is normalized to scale the feature value to a uniform numerical range. The sigmoid function is applied to the normalized feature value for nonlinear transformation to map the feature value to a probability space between 0 and 1. Apply the sigmoid function to obtain the transformed eigenvalues : The transformed eigenvalues represent the relative importance of the features in the data set. The closer to 1, the greater the impact of the feature on the project management goal, and the closer to 0, the smaller the impact. The transformed eigenvalues are arranged in the original feature order to form a standardized feature matrix : , where d represents the dimension of the feature subset, that is, the number of key features. Each column represents a standardized feature, and each row represents a data sample.
[0035] Figure 3 It is an exemplary flowchart of building a stacked ensemble learning model as shown in the present application. According to the feature subset, a multi-objective model for predicting project progress, quality and safety is constructed by a support vector machine (SVM) algorithm; including: using a stratified sampling method to randomly divide the feature subset into three mutually exclusive subsets in a certain proportion. The data set is usually divided into a training set, a validation set and a test set in a ratio of 8:1:1. Stratified sampling ensures that the distribution of samples in each subset on target attributes such as project progress, quality and safety is consistent with the original data set, avoiding the deviation of data division. Training set: used for parameter learning of the SVM regression model, by minimizing the empirical risk function, estimating the weight coefficient and bias term of the model. The training set needs to cover various scenarios and states of project management and provide sufficient sample diversity. Validation set: used for hyperparameter optimization of the SVM regression model, through grid search and cross-validation methods, select the hyperparameter combination with the best model performance. The validation set is independent of the training set and is used to evaluate the generalization ability of the model on unknown data and guide the tuning of hyperparameters. Test set: used for performance evaluation of SVM regression model. After training and optimization, the prediction accuracy, generalization error and other indicators of the model are evaluated on the test set. The test set simulates the actual application scenario and is used to verify the actual effect of the model.
[0036] Independent SVM regression sub-models are constructed for different project management goals. Each sub-model takes a standardized feature subset as input and predicts the continuous value of the corresponding target. SVM regression introduces an ε-insensitive loss function, which allows the predicted value to deviate from the true value within a certain range, thereby improving the robustness and generalization ability of the model. The radial basis kernel function (RBF) is used as the kernel function of the SVM regression sub-model. The RBF kernel function maps the input features to a high-dimensional space, making the nonlinear pattern in the original space separable in the high-dimensional space. The RBF kernel function is defined as follows: , where x and x' represent two feature vectors, ||·|| represents the Euclidean distance, and γ represents the parameter of the kernel function, which controls the width of the Gaussian distribution. The performance of the SVM regression sub-model is affected by hyperparameters and needs to be optimized through data-driven methods. A combination of grid search and K-fold cross-validation is used to search for the optimal combination of hyperparameters. Hyperparameters include: Penalty coefficient C: controls the complexity of the model and the tolerance for misclassified samples. The larger C is, the higher the model complexity is and the stricter the penalty for misclassified samples is. Kernel function parameter γ: controls the width of the RBF kernel function and affects the effect of feature mapping to high-dimensional space. The larger γ is, the narrower the Gaussian distribution is and the higher the model complexity is. Grid search is performed on the validation set to enumerate different combinations of C and γ, and K-fold cross-validation is used to evaluate the performance of each combination. The hyperparameter combination with the smallest average validation error is selected as the optimal hyperparameter.
[0037] The SVM progress prediction submodel, SVM quality prediction submodel and SVM safety prediction submodel are used as base learners. The validation set is used to train each submodel, and the prediction performance of each submodel is evaluated on the test set. Submodel training: For each SVM submodel, the validation set is used as training data, and the model is retrained using the optimized hyperparameters. The training process estimates the weight coefficients and bias terms of the model by minimizing the ε-insensitive loss function and the structural risk minimization principle. Submodel evaluation: The prediction performance of each SVM submodel is evaluated on the test set, and evaluation indicators such as mean absolute error (MAE) and root mean square error (RMSE) are calculated. The evaluation results reflect the generalization ability and prediction accuracy of the submodel on unknown data.
[0038] Meta-learner training: Training data preparation: The prediction results of the SVM progress prediction sub-model on the test set are recorded as , forming an N-dimensional vector, where N is the number of test set samples. , the prediction result of the SVM quality prediction sub-model on the test set is recorded as , forming an N-dimensional vector. , the prediction result of the SVM security prediction sub-model on the test set is recorded as , forming an N-dimensional vector. ,Will , , Concatenate by columns to form new training data , It is an N*3 matrix. , using the project's real progress, quality and safety indicators as labels , It is an N*3 matrix. ,in, , , They represent the true progress, quality and safety indicators of the i-th test sample respectively.
[0039] Decision tree algorithm: Use the CART (Classification and Regression Tree) algorithm to build a meta-learner, and define the meta-learner as a regression decision tree. As input features, The regression decision tree is trained with y as the target variable. The decision tree recursively divides the feature space and selects the optimal split feature and split threshold at each node to minimize the mean square error of the divided child nodes. The recursive partitioning process continues until the predetermined stopping condition is met, such as reaching the maximum tree depth, the number of node samples is lower than the threshold, etc. For the regression decision tree, the output of the leaf node is the mean value of the target variable of the node sample.
[0040] Hyperparameter optimization: A combination of grid search and K-fold cross validation is used to optimize the hyperparameters of the decision tree meta-learner. Hyperparameters include: Maximum tree depth: controls the complexity of the decision tree. A larger depth can capture more complex feature interactions, but may lead to overfitting. Minimum number of leaf node samples: controls the growth of the tree. A larger value can reduce overfitting, but may lose some information. Maximum number of split features: controls the number of features considered for each split. A smaller value can reduce computational complexity. Set the candidate value range of the hyperparameter and enumerate all possible hyperparameter combinations through grid search. For each hyperparameter combination, K-fold cross validation is used to evaluate the performance of the decision tree on the training set. The training set is randomly divided into K mutually exclusive subsets. Each time, K-1 subsets are selected as training data, and the remaining 1 subset is used as validation data. Repeat K times, each time selecting a different subset as validation data, and obtain K validation results. Calculate the average of the K validation results as the performance metric for the hyperparameter combination. Select the hyperparameter combination with the best average validation performance as the optimal hyperparameter for the decision tree meta-learner.
[0041] Model combination: Base learner layer: deploy the trained SVM progress prediction sub-model, SVM quality prediction sub-model and SVM safety prediction sub-model in parallel. , each sub-model makes predictions independently. The SVM progress prediction sub-model outputs the progress prediction value The SVM quality prediction sub-model outputs the quality prediction value The SVM safety prediction sub-model outputs the safety prediction value .Will , , Prediction results of combined base learners . .
[0042] Meta-learner layer: The prediction results of the base learner Input into the trained decision tree meta-learner. The decision tree meta-learner traverses the tree structure and The optimal splitting path is selected by the feature value until the leaf node is reached. The output value of the leaf node is the mean of the target variable of the node sample, which is used as the prediction result of the meta-learner. . ,in, , , Represent the final predicted values of project progress, quality and safety respectively.
[0043] Figure 4According to the exemplary flow chart of obtaining the Pareto optimal solution set shown in the present application, the non-dominated sorting genetic algorithm NSGA-II is used to solve the multi-objective model, and the simulated annealing strategy is introduced in the iterative search process of NSGA-II to jump out of the local optimal solution, and the optimized Pareto optimal solution set is output; including: according to the established multi-objective prediction model, the key parameters of the model are used as optimization variables to form a decision space, and the model performance evaluation index is selected as the optimization target to construct a multi-objective optimization function. Decision space construction: The penalty coefficient C and the kernel function parameter γ of the base learner SVM are selected as optimization variables. C controls the fault tolerance and generalization performance of the SVM. A larger C will lead to overfitting, and a smaller C will lead to underfitting. γ controls the width of the RBF kernel function. A larger γ will increase the complexity of the model, and a smaller γ will make the model too simple. The tree depth d, the node splitting criterion s and the minimum number of leaf node samples m of the meta-learner decision tree are selected as optimization variables. d controls the complexity of the decision tree. A larger d will lead to overfitting, and a smaller d will lead to underfitting. s controls the standard for node splitting. Commonly used criteria include Gini index and information gain. m controls the minimum number of samples for leaf nodes. A larger m will cause the tree to stop early, and a smaller m will cause the tree to be too complex. Combine the optimization variables into a D-dimensional vector x = (C, γ, d, s, m), where D is the dimension of the optimization variable. Determine the D-dimensional decision space based on the value range of each optimization variable. The value range of C is (0, positive infinity), and can be sampled on a logarithmic scale, such as arrive The value range of γ is (0, positive infinity), and can be sampled on a logarithmic scale, such as arrive The logarithmic uniform sampling between . The value range of d is [1, maximum tree depth], and the maximum tree depth is usually set to the logarithmic value of the number of features or samples in the data set. The value range of s is [0, 1]. For the Gini index, the optimal value is 0; for information gain, the optimal value is 1. The value range of m is [1, maximum number of leaf node samples], and the maximum number of leaf node samples is usually set to between 1% and 10% of the number of samples in the training set.
[0044] Multi-objective function construction: The mean absolute percentage error MAPE (x) of the multi-objective prediction model on the training set is selected as one of the optimization objectives. MAPE (x) measures the model's fitting ability on the training set. A smaller MAPE (x) indicates that the model has a higher degree of fit to the training data. The calculation formula for MAPE (x) is: , where n is the number of samples in the training set, is the true value of the i-th training sample, is the prediction value of the model for the i-th training sample. The root mean square error RMSE (x) of the multi-objective prediction model on the test set is selected as one of the optimization objectives. RMSE (x) measures the generalization ability of the model on the test set. A smaller RMSE (x) indicates that the model has a stronger prediction ability for unknown data. The calculation formula of RMSE (x) is: , where m is the number of test set samples, is the true value of the jth test sample, is the model's prediction value for the jth test sample. The time complexity T(x) of the base learner and meta-learner is selected as one of the optimization objectives. T(x) measures the computational efficiency of the model. A smaller T(x) indicates a faster training and prediction speed of the model. For SVM, T(x) is related to the number of samples n, the number of features d, and the number of support vectors s, and is generally For decision trees, T(x) is related to the number of samples n, the number of features d, and the tree depth h, and is generally O(ndh). MAPE(x), RMSE(x), and T(x) are combined into a multi-objective function F(x): , x∈D, where D is the constructed D-dimensional decision space.
[0045] The established multi-objective function F(x) is solved by using the non-dominated sorting genetic algorithm NSGA-II to obtain the Pareto optimal solution set P. Initialization: Optimization variable encoding: Encode the D-dimensional optimization variable x=(C, γ, d, s, m) into a binary string. For continuous variables C and γ: Determine the value range of C and γ and According to the required accuracy, the value range is divided into and equally spaced discrete values. Use bit and The bits of binary numbers encode the discrete values of C and γ respectively. The codes of C and γ are concatenated in sequence to form a length of For discrete variables d, s, and m: determine the value set of d, s, and m , and .use , and The binary numbers encode the values of d, s and m respectively. The codes of d, s and m are concatenated in sequence to form a binary number of length The binary substrings of continuous variables and discrete variables are concatenated to form a complete individual code with a code length of .
[0046] Initial population generation: Randomly generate an initial population of size N, each individual corresponds to a feasible solution to the multi-objective function F(x). For each individual in the population: For continuous variables C and γ: and The values of C and γ are randomly generated within the range. The values of C and γ are quantized to the closest discrete values and converted into corresponding bit and For discrete variables d, s, and m: , and Randomly select the values of d, s and m in . Convert the selected values into the corresponding Bit, bit and Binary coding of continuous variables and discrete variables is concatenated to form a complete coding of individuals. Repeat the above process to generate N individuals and form the initial population.
[0047] Individual decoding: Decode the binary code of each individual in the initial population into the corresponding decision variable value. For continuous variables C and γ: bit and Convert binary substrings into decimal integers and . According to the range and precision, the integer and Mapped to the corresponding actual values of C and γ: , , for discrete variables d, s, and m: convert the individual codes into Bit, bit and Convert binary substrings into decimal integers , and . According to the value set, the integer , and Mapped to the corresponding actual values of d, s and m: , where d is the value set of d, , where s is the value set of s, , where m is the value set of m, and the decoded actual values of C, γ, d, s and m are combined into the decision variable combination corresponding to the individual .
[0048] Fast non-dominated sorting: Objective function value calculation: Calculate the objective function values MAPE(x), RMSE(x), and T(x) for each individual in the initial population. For each individual x in the population: Substitute individual x into the training set data and calculate the prediction results of the multi-objective prediction model. Based on the prediction results and the true value of the training set, calculate the MAPE(x) and RMSE(x) of the model corresponding to individual x. According to the computational complexity of the model and the number of samples, estimate the computational time T(x) of the model corresponding to individual x. Obtain the evaluation values of all individuals in the population on the three objective functions.
[0049] Non-dominated sorting: Perform a fast non-dominated sort on the initial population based on the dominance relationship of the objective function value. For each individual x in the population: Initialize the dominance set of individual x and the dominated number is an empty set. For each pair of individuals (x, y) in the population: Compare the evaluation values of x and y on the three objective functions MAPE(x), RMSE(x), T(x) and MAPE(y), RMSE(y), T(y). If x is better than or equal to y on all objective functions and strictly better than y on at least one objective function, then x dominates y. Add y to the dominating set of x If y is better than or equal to x in all objective functions and strictly better than x in at least one objective function, then y dominates x. Add 1. Find the individuals in the population whose number of dominance is 0 and mark them as the first-level non-dominated solution set F1. Remove the individuals in F1 from the population, repeat the non-dominated sorting on the remaining individuals, and obtain the second-level non-dominated solution set F2. Repeat the above process until all individuals in the population are assigned to a certain level of non-dominated solution set.
[0050] Crowding distance calculation: Calculate the crowding distance of each individual. For each non-dominated solution set :For each objective function (MAPE, RMSE and T): The individuals in The values are sorted in ascending order. For each individual x, calculate its The crowding distance :If x is The minimum or maximum individual on .otherwise, . and are x in The objective function values of the right and left neighbors on . and They are middle The maximum and minimum values of . Calculate The total crowding distance of each individual x in , which is the sum of the crowding distances of all objective functions.
[0051] Parent population selection: N individuals are selected based on the crowding distance to form the first generation parent population . According to the level of non-dominated solution set from low to high Consider in turn: If , then All individuals in .like , then sort them according to the congestion distance CD from large to small The individuals in Individuals join Finally, we get a parent population of size N. , individuals are sorted according to the non-dominated level and crowding distance. Genetic operation: Father selection: For individuals in the t-th generation parent population Pt, two of them are randomly formed into father individual pairs . The sire is selected by the Tournament Selection method: Randomly select k individuals from the Tournament Set to form a Tournament Set. Select the individual with the lowest non-dominated level from the Tournament Set as the parent individual If there are multiple individuals with the lowest non-dominated level in the Tournament Set, the individual with the largest crowding distance is selected as the parent individual. Repeat the above process and select another parent individual The paternal selection process is repeated until N / 2 pairs of paternal individuals are generated.
[0052] Adaptive simulated binary crossover: for each pair of parents Perform adaptive simulated binary crossover operation to obtain offspring individuals . Set the adaptive crossover probability . and are the upper and lower bounds of the crossover probability, t is the current evolutionary generation, is the maximum evolutionary generation. For each pair of paternal individuals , generate a random number r in [0, 1]. , then simulate binary crossover is performed on the pair of father individuals: for each decision variable (i=1, 2, ......., D), perform crossover operations respectively: generate random numbers in [0, 1] .according to Calculate the crossover factor :like ,but .otherwise, . is the distribution index, which controls the degree of aggregation of offspring individuals near the parent individuals. Calculate the decision variable value of the offspring individuals: ; ; Perform boundary processing on the decision variables of the offspring individuals to ensure that they are within the definition domain. Otherwise, directly copy the pair of parent individuals as offspring individuals, that is, . The N generated offspring individuals form a offspring population .
[0053] Polynomial mutation: for the offspring population Each individual in Perform polynomial mutation operation. Set mutation probability For each decision variable (i=1, 2, ..., D), generate a random number in [0, 1] .like , then perform polynomial mutation on the decision variable: generate a random number in [0, 1] .according to Calculate the factor of variation :like ,but .otherwise, . is the distribution index, which controls the degree of aggregation of individuals after mutation near the original individuals. Calculate the decision variable value after mutation: ; and The decision variables The upper and lower bounds of . Perform boundary processing on the mutated decision variables to ensure that they are within the domain. Put the mutated individuals back into the offspring population .
[0054] Mutation operation: Directional polynomial mutation: for the offspring population Each individual in , perform directional polynomial mutation to obtain the mutated offspring individuals . For offspring individuals Each component (i=1, 2, ..., D), with probability Perform polynomial mutation. In polynomial mutation, if the individual components Selected for mutation: Determining the parent individuals The corresponding amount .like It is the parent individual If crossover occurs, for and Zhongyu More similar individuals. Similarity can be determined by comparing Euclidean distance or decision variable values.
[0055] Calculate the preference factor λ: ; is a sign function. It returns 1 when the value in the bracket is greater than 0, returns -1 when it is less than 0, and returns 0 when it is equal to 0. k is the power exponent of the preference factor, which controls the preference strength and is usually 1 or 2. Generate a random perturbation δ of the polynomial distribution: Generate a random number in [0, 1] .according to calculate :like ,but .otherwise, . is the distribution index of the multinomial distribution, and Individual weight Update individual components : ; Perform boundary processing on the updated individual components to ensure that they are within the definition domain. The components of (i=1, 2, ......., D) mutate in sequence, and finally obtain the mutated offspring individuals . Population merging: merge the parent population and the mutated offspring population Merge to form a new population of size 2N . ; New population Contains all individuals in the parent population and the mutated offspring population.
[0056] Environment selection: Fast non-dominated sort: Repeat the fast non-dominated sort on the merged population Rt to obtain a non-dominated solution set . Use the same fast non-dominated sorting method as . Calculate the population The non-dominated rank and crowding distance of each individual in . Construct the next generation population: Initialize the next generation population is an empty set. From the non-dominated solution set At the beginning, select individuals from each solution set to join , until the number of individuals reaches N. , then All individuals in .like , then enter the crowding distance selection stage. Crowding distance selection: For the last layer of non-dominated solution set , if the number of individuals exceeds the remaining vacancies :calculate The crowding distance of each individual in . Sort the individuals in descending order. Select the one with the largest crowding distance. Individuals join . Update the population: Finally, we get the next generation population of size N . It contains elite individuals from both the parent and child populations, while maintaining the diversity of the population. Termination condition: Iteration process: Set the maximum number of evolution generations . Starting from the first generation, iterate until reaching For the current population Perform genetic operations, including parent selection, crossover, and mutation, to generate offspring populations . The parent population and progeny population Merge to form a combined population For the combined population Execute environment selection to obtain the next generation population . Let t=t+1 and enter the next generation of evolution.
[0057] Introducing simulated annealing operation in the NSGA-II iteration process: Parameter setting: initial temperature Setting: Initial temperature Determines the probability that the simulated annealing algorithm accepts inferior solutions in the initial stage. The value of should be large enough to ensure that the algorithm can accept inferior solutions with a high probability in the early stage and jump out of the local optimum. It can be determined by empirical formulas or experiments. The commonly used empirical formulas for the value of include: ,in is the maximum difference of the objective function value, is the initial acceptance probability (usually 0.8~0.9). ,in is the average difference of the objective function value, and K is the proportionality coefficient (usually 1 to 100). Example: Setting , indicating that the initial temperature is 100. Setting of the cooling coefficient α: The cooling coefficient α determines the rate of temperature drop and controls the convergence speed of the algorithm. The value range of α is (0, 1). A larger α value indicates a slower temperature drop and a slower algorithm convergence speed; a smaller α value indicates a faster temperature drop and a faster algorithm convergence speed. Commonly used α values include 0.8, 0.9, 0.95, etc., which can be adjusted according to the characteristics of the problem and the performance of the algorithm. Example: Setting α = 0.9 means that the temperature is multiplied by 0.9 each time the temperature drops. Minimum temperature threshold Setting: Minimum temperature threshold Determines the termination condition of the simulated annealing algorithm. When the temperature T drops to When , the algorithm stops searching and considers that the global optimal solution has been reached. The value of should be small enough to ensure that the algorithm can terminate after finding the global optimal solution. It can be determined based on the characteristics of the problem and the convergence of the algorithm. The value of is usually a smaller value between 0.1 and 1. Example: Setting , indicating that when the temperature drops below 0.1, the algorithm stops searching. Determination of the simulated annealing execution frequency t: The simulated annealing execution frequency t determines how many generations the simulated annealing operation is introduced during the NSGA-II iteration process. The value of t should be determined based on factors such as the complexity of the problem, the convergence speed of the algorithm, and the limitations of computing resources. A larger t value means a lower frequency of simulated annealing operations, and the algorithm is more inclined to local search; a smaller t value means a higher frequency of simulated annealing operations, and the algorithm is more inclined to global exploration. The appropriate t value can be determined through experiments and comparisons, usually an integer value between 5 and 20. Example: Setting t = 10 means that a simulated annealing operation is introduced every 10 generations.
[0058] Simulated annealing process: Determine whether to perform simulated annealing: Get a new population in each round of NSGA-II iteration After that, calculate the current iteration number t. Determine whether t is an integer multiple of the simulated annealing execution frequency, that is, t % simulated annealing execution frequency == 0. If so, enter the simulated annealing process; otherwise, skip the simulated annealing and directly enter the next round of NSGA-II iteration. Select individuals: from Randomly select an individual from the non-dominated solution set As the current solution. The non-dominated solution set can be obtained by fast non-dominated sorting or other non-dominated sorting algorithms. Random selection can be achieved by random number generator or other randomization strategies. Perturbed decision variables: The decision variables x = (C, γ, d, s, m) are disturbed to obtain new individuals The perturbation method is to add random noise to each component of x, that is: ,in for Random value within . ,in for Random value within . ,in for Random value within . ,in for Random value within . ,in for The range of random noise can be set according to the range of decision variables and the characteristics of the problem. Random noise can be generated by a random number generator, such as uniform distribution, Gaussian distribution, etc.
[0059] Boundary processing: perturbed decision variables Perform boundary processing to ensure that it is within the domain. If a component exceeds the boundary of the domain, the following processing methods can be adopted: Truncation to boundary value: If it exceeds the upper boundary, take the upper boundary value; if it exceeds the lower boundary, take the lower boundary value. Reflection: If it exceeds the boundary, reflect the excess part back into the domain. Periodic mapping: If it exceeds the boundary, map it to the other end of the domain. Calculate the objective function value: Calculate the new individual The objective function value of . represents the mean absolute percentage error, represents the root mean square error, Represents the model training time. The objective function value can be obtained by substituting into the model and evaluating the model performance.
[0060] Compare the objective function value: the new individual The objective function value of With the current individual The objective function value of For comparison. Dominate , that is, better than or equal to , then accept the new solution ,use replace Otherwise, accept the inferior solution with a certain probability , the probability calculation formula is: , where T is the current temperature. Generate a random number rand between 0 and 1. If rand < P, accept the inferior solution. ,use replace Otherwise, keep the current solution Update temperature: Update the current temperature T according to the cooling coefficient α, that is The range of the temperature reduction coefficient α is (0, 1), and commonly used values are 0.8, 0.9, 0.95, etc. Determine whether to terminate simulated annealing: Determine whether the current temperature T is lower than the preset minimum temperature threshold .if , the simulated annealing process is stopped; otherwise, the simulated annealing process is continued: if the current temperature T is still higher than , then return to the simulated annealing process. Update the population: The new solution obtained by simulated annealing Joining the next generation population of NSGA-II as a new search starting point. Individuals randomly selected from The individuals with smaller crowding distance. Continue NSGA-II iteration: The population updated by simulated annealing operation As a new generation population. Continue to perform the NSGA-II iteration process until the maximum number of iterations is reached. .
[0061] Output the optimal solution set and substitute it into the multi-objective prediction model: After the NSGA-II iterative optimization and simulated annealing optimization process, the final population P is obtained. Perform non-dominated sorting on the population P to obtain the Pareto optimal solution set . Pareto optimal solution set Contains all non-dominated solutions in the population P, that is, solutions that are not dominated by other solutions on multiple objective functions. Each solution in the solution set P Represents a set of optimized parameter combinations, which is an optimal solution to the multi-objective optimization problem. For each optimal solution in the Pareto optimal solution set P :Combining the optimized parameters Substitute the constructed multi-objective prediction model. The multi-objective prediction model is trained using the training set data and the initial parameters before optimization. Substituting the optimized parameter combination into the multi-objective prediction model is equivalent to retraining or fine-tuning the model with the optimized parameters.
[0062] Select the optimal parameter combination based on the Pareto optimal solution set, use the optimal parameter combination to adjust the BIM model of the current intelligent construction project, and use the adjusted BIM model to guide construction. Select the optimal solution: According to the preset performance threshold, select the optimal solution that meets the conditions from the Pareto optimal solution set P. The preset performance thresholds include: Mean absolute percentage error threshold , represents the acceptable prediction error range. Root mean square error threshold , which indicates the acceptable prediction error size. Calculation time threshold , represents the acceptable model training and prediction time. For each optimal solution , to determine whether the corresponding objective function value meets the performance threshold: If , ,and , then the solution is used as the candidate optimal solution. If multiple solutions meet the conditions, one of them can be selected as the final optimal solution according to actual needs. If no solution satisfies the condition, the performance threshold can be appropriately relaxed, or the solution with the best objective function value can be selected from the candidate solutions as the optimal solution. .
[0063] Substitute into the multi-objective prediction model: the optimal solution The parameter combination in Substitute the constructed multi-objective prediction model. The multi-objective prediction model consists of SVM progress prediction sub-model, SVM quality prediction sub-model, SVM safety prediction sub-model and decision tree meta-learner. Use the optimal parameter combination Retrain or fine-tune each sub-model and meta-learner: For the SVM progress prediction sub-model, use As the parameters of SVM, use As the data preprocessing parameters. For the SVM quality prediction sub-model and the SVM safety prediction sub-model, similar parameter setting methods are used. For the decision tree meta-learner, use As the parameters of data preprocessing, the output of the sub-model is used as the input feature of the meta-learner. Through the retrained or fine-tuned multi-objective prediction model, an optimized joint prediction of project progress, quality and safety is formed.
[0064] Dynamically adjust the BIM model: Use the optimized SVM progress prediction submodel to predict the planned duration and completion percentage of each process of the project. According to the prediction results, calculate the progress deviation of each process, that is, the difference between the actual progress and the planned progress. Associate the progress deviation information with the construction status parameters of the corresponding components in the BIM model, and dynamically update the construction status of the components. Use the optimized SVM quality prediction submodel to predict the probability of occurrence of key quality problems in the project. According to the prediction results, identify high-risk quality problems and associate them with the quality attribute parameters of the corresponding components in the BIM model. For components predicted to be high-risk, add quality warnings or quality control measures to their quality attribute parameters. Use the optimized SVM safety prediction submodel to predict high-risk operation areas and the probability of accidents. According to the prediction results, identify high-risk operation areas and associate them with the spatial positions of the corresponding components in the BIM model. For operation areas predicted to be high-risk, add safety warnings or safety control measures to the attributes of the corresponding components. Through dynamic association and update, the progress, quality and safety prediction information are integrated into the BIM model to form an information-enhanced dynamic BIM model.
[0065] Integrated application in intelligent construction management platform: Import the dynamically adjusted BIM model into the intelligent construction management platform as the core data foundation of the digital twin system. In the intelligent construction management platform, the optimized multi-objective prediction model is integrated to realize the real-time prediction and early warning functions of project progress, quality and safety. Based on the BIM model and prediction results, functional modules such as progress management, quality management, safety management and resource optimization are integrated in the management platform. Progress management module: According to the predicted progress deviation, the progress warning is automatically generated, and decision support for progress correction measures is provided. Quality management module: According to the predicted probability of quality problems, the quality warning is automatically generated, and decision support for quality control measures is provided. Safety management module: According to the predicted high-risk operation areas and accident probability, the safety warning is automatically generated, and decision support for safety control measures is provided. Resource optimization module: According to the predicted progress, quality and safety conditions, the project resource allocation is optimized and the resource utilization efficiency is improved. By integrating multi-source heterogeneous data and intelligent algorithms, a digital twin system covering the entire process of design, construction, operation and maintenance is formed. The visualization, prediction, early warning and decision support functions provided by the digital twin system are used to guide all project participants to carry out intelligent construction in a collaborative manner. Design stage: Use BIM models and prediction models to optimize design solutions and improve design quality and constructability. Construction stage: Use dynamically adjusted BIM models and prediction and early warning functions to achieve real-time monitoring and intelligent management of the construction process. Operation and maintenance stage: Use digital twin systems to achieve full life cycle management of facilities, improve operation and maintenance efficiency and asset value. Through the collaborative function of the intelligent construction management platform, promote information sharing and collaborative work among all project participants, and improve project management efficiency and decision-making level.
Claims
1. A data-driven intelligent construction method, characterized in that: include: Collect multi-source heterogeneous data of intelligent construction projects; multi-source heterogeneous data includes environmental data, material and equipment data, personnel organization data and construction process data; The recursive feature elimination (RFE) algorithm is used to select features from the collected multi-source heterogeneous data to obtain feature subsets that reflect project progress, quality, and safety. Based on the feature subset, a multi-objective model for predicting project progress, quality and safety is constructed using the support vector machine (SVM) algorithm. The non-dominated sorting genetic algorithm NSGA-II is used to solve the multi-objective model. The simulated annealing strategy is introduced in the iterative search process of NSGA-II to jump out of the local optimal solution and output the optimized Pareto optimal solution set. The optimal parameter combination is selected according to the Pareto optimal solution set, the BIM model of the current intelligent construction project is adjusted using the optimal parameter combination, and the adjusted BIM model is used to guide the construction.
2. The data-driven intelligent construction method according to claim 1, characterized in that: The recursive feature elimination (RFE) algorithm is used to select features from the collected multi-source heterogeneous data, including: Convert the collected multi-source heterogeneous data into structured data sets; According to the structured data set, the indicators reflecting the project progress, quality and safety are extracted to construct the original feature set; the original feature set is screened by using the recursive feature elimination (RFE) algorithm; The filtered features are transformed nonlinearly based on the sigmoid function, and the standardized feature subset is output.
3. The data-driven intelligent construction method according to claim 2, characterized in that: The original feature set is screened based on the recursive feature elimination (RFE) algorithm, including: The feature importance evaluation method based on decision tree is used to calculate the information gain of each feature in the original feature set as the feature importance weight; According to the feature importance weights, the original features are sorted in descending order, and the feature or a group of features with the lowest weights are removed to obtain the feature subset after dimensionality reduction; The feature subset after dimensionality reduction is used as input to construct a support vector machine (SVM) classifier. The K-fold cross-validation method is used to divide the training set and the validation set. The classification accuracy of the SVM classifier is evaluated on the validation set, and the classification accuracy is used as the quality score of the feature subset. Repeat the quality score evaluation step to obtain the feature subset with the highest quality score in each iteration. When the change in the optimal quality score of multiple consecutive iterations is less than the threshold, stop the iteration and output the optimal feature subset.
4. The data-driven intelligent construction method according to claim 2, characterized in that: A multi-objective model for predicting project progress, quality, and safety is built using the support vector machine (SVM) algorithm, including: The feature subset is divided into a training set, a validation set, and a test set; the training set is used for parameter learning of the SVM regression model, the validation set is used for hyperparameter optimization of the SVM regression model, and the test set is used for performance evaluation of the SVM regression model; Separate SVM regression sub-models are constructed for project progress, quality, and safety. The radial basis kernel function is used as the kernel function of the SVM regression sub-model. The hyperparameters of each SVM regression sub-model are optimized through grid search and K-fold cross-validation methods. The hyperparameters include penalty coefficients and kernel function parameters. According to the optimized SVM progress prediction sub-model, SVM quality prediction sub-model and SVM safety prediction sub-model, a multi-objective prediction model is constructed by adopting a stacking-based ensemble learning strategy.
5. The data-driven intelligent construction method according to claim 4, characterized in that: A stacking-based ensemble learning strategy is used to build a multi-objective prediction model, including: The SVM progress prediction sub-model, SVM quality prediction sub-model and SVM safety prediction sub-model are used as base learners. The validation set is used to train each sub-model, and the prediction performance of each sub-model is evaluated on the test set. The prediction results of each SVM sub-model on the test set are used as new training data, and the actual progress, quality and safety indicators of the project are used as labels to train the meta-learner. The meta-learner is constructed using the decision tree algorithm, and the hyperparameters of the meta-learner are optimized through grid search and K-fold cross validation. The base learner and the meta learner are combined in a hierarchical structure to form a complete multi-objective prediction model; the base learner outputs the preliminary prediction results of each target, and the meta learner performs nonlinear combination on the preliminary prediction results to output the final project progress, quality and safety prediction values.
6. The data-driven intelligent construction method according to claim 5, characterized in that: The non-dominated sorting genetic algorithm NSGA-II is used to solve the multi-objective model, including: According to the established multi-objective prediction model, the penalty coefficient C of the base learner SVM, the kernel function parameter γ, the tree depth d, the node splitting criterion s and the minimum number of leaf node samples m of the meta-learner decision tree are used as optimization variables. , forming a D-dimensional decision space; Select the mean absolute percentage error MAPE(x) of the multi-objective prediction model on the training set, the root mean square error RMSE(x) on the test set, and the time complexity T(x) of the base learner and meta learner as the optimization objectives, and construct the multi-objective function: , x∈D The established multi-objective function F(x) is solved by using the non-dominated sorting genetic algorithm NSGA-II to obtain the Pareto optimal solution set P; During the NSGA-II iteration process, simulated annealing operation is introduced every t generations to jump out of the local optimal solution; Output the Pareto optimal solution set P obtained after iterative optimization, and convert each optimal solution in the solution set P into The corresponding set of parameter combinations , substitute into the multi-objective prediction model constructed in step S33.
7. The data-driven intelligent construction method according to claim 6, characterized in that: The non-dominated sorting genetic algorithm NSGA-II is used to solve the problem and obtain the Pareto optimal solution set P, including: D-dimensional optimization variables Encoded as a binary string, randomly generate an initial population of size N, where each individual corresponds to a feasible solution of the multi-objective function F(x); Calculate the objective function values MAPE(x), RMSE(x) and T(x) of each individual in the initial population, perform fast non-dominated sorting on the initial population according to the dominance relationship of the objective function values, obtain the Pareto level and crowding distance of the individuals, and select N individuals based on the crowding distance to form the first generation parent population P0; the dominance relationship is determined by comparing the advantages and disadvantages of the individuals in each objective function; For the individuals in the t-th generation parent population Pt, two of them are randomly formed into pairs of paternal individuals. , for each pair of paternal Perform adaptive simulated binary crossover operation to obtain offspring individuals , there are N offspring individuals in total, forming a offspring population ; For offspring population Each individual in , perform directional polynomial mutation to obtain the mutated offspring individuals , the parent population and the mutated offspring population Merge to form a new population of size 2N ; For populations Repeat the fast non-dominated sort to obtain a non-dominated solution set ;from At the beginning, select individuals from each solution set in turn to join the next generation population , until the number of individuals reaches N; if the last layer of solution set The number of individuals exceeds , then in Select the first layer with the largest crowding distance. Individuals incorporated , and get the next generation population of size N ; Iterate until the maximum number of generations is reached. , output the final population The corresponding non-dominated solution set P; the solution set P is the Pareto optimal solution set of the multi-objective function F (x), and each optimal solution in the set Represents a set of optimized parameter combinations of multi-objective prediction models.
8. The data-driven intelligent construction method according to claim 7, characterized in that: For each pair of paternal Performs adaptive analog binary crossover operations including: Setting the adaptive crossover probability : ; in, and are the upper and lower bounds of the crossover probability, t is the current iteration number, is the maximum number of iterations; For parent population Each pair of individuals , generate a random number r in [0, 1]: like , then for the corresponding pair of paternal individuals Perform a simulated binary crossover to generate two offspring individuals ; Otherwise, directly copy the corresponding parent individuals as offspring individuals, that is, .
9. The data-driven intelligent construction method according to claim 8, characterized in that: Get the mutated offspring individuals ,include: For offspring Each component , with probability Perform polynomial mutation; In polynomial mutation, if the individual components Selected for mutation, based on the corresponding component of the parent individual Introducing the preference factor λ: ; in, For the parent individual The value of component i, k is the power index of the preference factor, is a symbolic function; Introducing the preference factor λ into the perturbation process of polynomial mutation, the individual components are updated as follows: ; Among them, δ is the random perturbation of the multinomial distribution; For individuals The components of Mutate in sequence and finally obtain the mutated offspring individuals .
10. The data-driven intelligent construction method according to claim 7, characterized in that: The simulated annealing strategy is introduced into the iterative search process of NSGA-II, including: Set the initial temperature of simulated annealing and cooling coefficient α, let the current temperature ; In each iteration of NSGA-II, a new population is obtained Afterwards, from Randomly select an individual from the non-dominated solution set , for its decision variables Perturb and get new individuals ; The perturbation method is to add random noise to each component of x; Calculate new individuals The objective function value of , and the current individual The objective function value of According to the Metropolis criterion, if Dominate , then accept the new solution ,use replace Otherwise, with probability Accept inferior solution ; Update Temperature , if T is lower than the preset threshold , then stop simulated annealing; otherwise, return to continue iteration; The new solution obtained by simulated annealing Joining the next generation population of NSGA-II As the new search starting point, continue the iterative process until the maximum number of iterations is reached .
Citation Information
Cited By
Color master batch processing energy consumption and efficiency collaborative optimization system based on production data
CN120317840A
Operator mapping method, device and equipment of heterogeneous system and medium
CN120654786A
Machine learning method and system for dry-type transformer manufacturing process optimization
CN120706226A
Engineering construction whole process digital management method and system
CN120806232A
Electric power project budget calculation method and system
CN120873504A