Component performance prediction method based on data-driven generalization enhancement and related device
By converting the input features of civil engineering models into empirical formula-based factorial forms and combining representative sampling and incremental learning, the problem of accuracy degradation of machine learning models outside the database range is solved, achieving broader and more accurate component performance prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGAN UNIV
- Filing Date
- 2022-11-15
- Publication Date
- 2026-04-28
AI Technical Summary
In civil engineering, machine learning-based models can typically only guarantee accuracy within the training database. When the point to be predicted falls outside the database range, its accuracy may drop significantly, resulting in poor prediction performance.
A data-driven generalization enhancement approach is adopted. This approach involves converting the original model input features into the basic factor form of an empirical formula, collecting measured data and converting it to the same form, generating a dataset using a representative sampling strategy, pre-training the model using a classic data-driven model, and then training the model through incremental learning to establish a generalized enhancement prediction model.
It achieves accurate predictions outside the training database, improves the model's interpolation and extrapolation capabilities, and enhances the accuracy and coverage of component performance predictions.
Smart Images

Figure CN115719037B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of structural component performance prediction, specifically a component performance prediction method and related apparatus based on data-driven generalization enhancement. Background Technology
[0002] Civil engineering has developed alongside human society, aiming to construct various types of spaces needed for human production and daily life. The structures created by civil engineering must be able to withstand various natural or man-made forces, providing reliable shelter for human production and daily life; therefore, the safety of civil engineering is of paramount importance.
[0003] In recent years, data-driven models based on machine learning have become increasingly common in structural performance prediction. They can quickly and easily learn the inherent patterns in existing data and exhibit good performance. However, machine learning models typically only guarantee accuracy within the range of the training database used. When the points to be predicted fall outside the database range, their accuracy may drop significantly. This is detrimental to the application of machine learning in civil engineering because there will always be a gap between the range of collected data and the range of data to be predicted. Specifically, most of the experimental data that can be collected is obtained on scaled-down models, which are much smaller than those actually used in engineering. Summary of the Invention
[0004] The purpose of this invention is to provide a method and related apparatus for predicting component performance based on data-driven generalization enhancement, in order to solve the problem that machine learning-based models can usually only guarantee accuracy within the range of the training database used, while their accuracy may drop significantly when the point to be predicted falls outside the range of the database.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A data-driven generalization enhancement-based component performance prediction method, including
[0007] Determine the original model input features for predicting component performance;
[0008] Select an empirical formula model and transform the input features of the original model into the basic factor form of the empirical formula according to the structure of the empirical formula;
[0009] Collect the measured dataset of structural component performance prediction and convert it into the basic factor form of empirical formula to establish the dataset of measured data; the basic factor of empirical formula is used as the input of the model and the measured structural performance index is used as the output of the model.
[0010] A representative sampling strategy is selected to generate a representative dataset; its input features are the basic factor forms of empirical formulas, and its output is the calculated value of the empirical formulas.
[0011] Based on a representative dataset, a pre-trained model is obtained by training the model on a classic data-driven model.
[0012] The measured dataset is introduced into the pre-trained model, and incremental learning is used to train the model again to obtain a data-driven generalization enhancement-based structural component performance prediction model for prediction.
[0013] Furthermore, the original model input features include: concrete compressive strength, tensile and compressive strength of steel reinforcement in the tension and compression zones, tensile strength of prestressed tendons in the tension zone, cross-sectional area of longitudinal ordinary steel reinforcement in the tension and compression zones, cross-sectional area of longitudinal prestressed tendons in the tension zone, cross-sectional width, cross-sectional height, distance from the resultant point of longitudinal ordinary steel reinforcement in the compression zone to the compression edge of the cross-section, and coefficients when simplifying the equivalent rectangular stress diagram of concrete in the compression zone; the polynomial composition of the empirical formula includes the following basic factors: constant terms or exponential terms.
[0014] Furthermore, the collected measured dataset of structural component performance is not required to cover all potential data value ranges.
[0015] Furthermore, when sampling based on empirical formulas and prior knowledge, the sampling space should cover all potential data value ranges, and the sampling strategies used include SOBOL sampling and Latin hypercube sampling.
[0016] Furthermore, random search, grid search, and Bayesian optimization algorithms are used for hyperparameter optimization and training of the data-driven model.
[0017] Furthermore, when incrementally training the established model, the performance metrics of the model on datasets D1 and D2 are considered simultaneously; the weight ratio of the metrics on the two datasets is determined based on the performance of the empirical formula when predicting component performance. When the performance of the empirical formula is better, the weight on dataset D2 is increased accordingly.
[0018] Furthermore, pre-trained models w0 and θ0 are the model's weights and hyperparameters, respectively. The prediction process is as follows: collect the characteristic indicators of the structural components to be predicted, transform the input form based on the basic factor form of the empirical formula, input the trained data to drive the model, and obtain the prediction results.
[0019] Furthermore, a component performance prediction system based on data-driven generalization enhancement includes:
[0020] The feature determination module is used to determine the original model input features for predicting component performance;
[0021] The conversion module is used to select an empirical formula model and convert the original model input features into the basic factor form of the empirical formula according to the structure of the empirical formula.
[0022] The dataset creation module is used to collect the measured dataset of structural component performance prediction and convert it into the basic factor form of empirical formulas to establish the dataset of measured data; the basic factors of empirical formulas are used as input terms of the model, and the measured structural performance indicators are used as output terms of the model.
[0023] The input module is used to select a representative sampling strategy to generate a representative dataset; its input features are the basic factor forms of the empirical formula, and its output is the calculated value of the empirical formula.
[0024] The prediction module is used to train a model on a classic data-driven model based on a representative dataset to obtain a pre-trained model. The actual test dataset is then introduced into the pre-trained model, and incremental learning is used to train the model again to obtain a data-driven generalization enhancement-based structural component performance prediction model for prediction.
[0025] Furthermore, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a component performance prediction method based on data-driven generalization enhancement.
[0026] Furthermore, a computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of a component performance prediction method based on data-driven generalization enhancement.
[0027] Compared with the prior art, the present invention has the following technical effects:
[0028] The data-driven generalization enhancement-based structural component performance prediction method proposed in this invention utilizes a dataset based on empirical formulas for model training to obtain a pre-trained model, thus fully leveraging the physical knowledge within existing empirical formulas. During the learning process of the physical knowledge in the empirical formulas, machine learning algorithms are employed, combining the acquired physical knowledge with a data-driven prediction framework. The method proposed in this invention can be used to obtain accurate component performance prediction models. This method uses machine learning algorithms with good training results, thus exhibiting excellent interpolation capabilities. Furthermore, because this method first uses a dataset based on empirical formulas during the training phase, its coverage is extremely broad, ensuring excellent extrapolation capabilities. Therefore, the performance of the method proposed in this invention surpasses that of traditional empirical formulas and typical machine learning algorithms.
[0029] The method proposed in this invention can directly predict structural performance (such as the bending and shear bearing capacity of beams), and is supplemented by structural monitoring and test data, making this method useful for structural performance evaluation, remaining service life prediction, durability analysis and structural maintenance. Attached Figure Description
[0030] Figure 1 This is a schematic flowchart of the structural component performance prediction method based on data-driven generalization enhancement of the present invention;
[0031] Figure 2 This is a schematic diagram of the specific training process of the pre-trained model G;
[0032] Figure 3 This is a schematic diagram of the specific training process for incremental learning on a pre-trained model G.
[0033] Figure 4 This is the model usage process of the present invention when predicting the performance of structural components after model training is completed;
[0034] Figure 5 Taking the prediction of the flexural bearing capacity of a beam's normal section as an example, the diagram shows the performance of the pre-trained model G on the training set and test set when it performs incremental learning.
[0035] Figure 6 Taking the prediction of the flexural bearing capacity of a beam's normal section as an example, the results of three prediction methods—the method proposed in this invention, the empirical formula, and the classic machine learning algorithm—are plotted. Detailed Implementation
[0036] The present invention will be further described below with reference to the accompanying drawings:
[0037] Please see Figures 1 to 6 ,
[0038] S1: Determine the original model input features x = [x1, x2, ... x] for predicting component performance. t ].
[0039] S2: Select an empirical formula model, and set the initial input features x = [x1, x2, ... x t Transform the empirical formula into its basic factor form according to its structure.
[0040] S3: Collect the measured dataset of structural component performance prediction and convert it into the basic factor form of the empirical formula, thereby establishing the measured data dataset D1. The basic factors of the empirical formula are used as the input terms of the model, and the measured structural performance indicators are used as the output terms of the model.
[0041] S4: To fully embed the prior knowledge of the corresponding empirical formulas into the data-driven model, a representative dataset D2 is generated based on a representative sampling strategy selected from the empirical formulas. Its input features are the basic factor forms of the empirical formulas, and its output is the calculated value of the empirical formulas.
[0042] S5: Based on dataset D2, train the model on the classic data-driven model to obtain the pre-trained model. w0 and θ0 are the model's weights and hyperparameters, respectively.
[0043] S6: Introduce the measured dataset into the pre-trained model G, and use incremental learning to train the model again to obtain a data-driven generalization enhancement-based structural component performance prediction model.
[0044] In step S2, the polynomial composition of the empirical formula may include, but is not limited to, the following basic factors: constant term, exponent term, etc.
[0045] In step S3, the collected measured dataset of structural component performance is not required to cover all potential data value ranges.
[0046] In step S4, when sampling based on prior knowledge of empirical formulas, the sampling space should cover all potential data value ranges. Sampling strategies that can be used include, but are not limited to, Sobol sampling, Latin hypercube sampling, etc.
[0047] In steps S5 and S6, when optimizing and training the hyperparameters of the data-driven model, methods such as random search, grid search, and Bayesian optimization algorithms can be used.
[0048] In step S6, when incrementally training the established model, the performance metrics of the model on datasets D1 and D2 should be considered simultaneously. The weight ratio of the metrics on the two datasets can be determined based on the performance of the empirical formula when predicting component performance. When the performance of the empirical formula is better, the weight on dataset D2 can be increased accordingly to learn more knowledge from it.
[0049] Example:
[0050] like Figure 1 As shown, the specific steps of this invention are as follows:
[0051] S1: Based on the currently mature method for predicting the flexural capacity of beam sections, the relevant parameters for predicting the flexural capacity of beam sections are determined as x = [f c ,f y ,f′ y ,f py A s ,A′ s A p,b,h0,a′ s ,α1].
[0052] Among them, f c f is the compressive strength of concrete. y ,f′ y These represent the tensile strength and compressive strength of ordinary steel reinforcement in the tension zone and compression zone, respectively; f py A represents the tensile strength of the prestressed tendons in the tension zone. s ,A′ s These are the cross-sectional areas of the longitudinal ordinary reinforcing bars in the tension zone and compression zone, respectively; A p a is the cross-sectional area of the longitudinal prestressing tendons in the tension zone; b is the cross-sectional width; h0 is the effective height of the cross-section; a′ s α1 is the distance from the resultant point of the longitudinal ordinary steel reinforcement in the compression zone to the compression edge of the section; α1 is the coefficient when the concrete in the compression zone is simplified into an equivalent rectangular stress diagram.
[0053] S2: The formula for predicting the flexural capacity of a beam's normal section is specified in GB50010-2010 Code for Design of Concrete Structures as follows:
[0054]
[0055] α1f c bx = f y A s -f y 'A' s +f py A p (2)
[0057] Where M is the predicted value of the flexural bearing capacity of the beam's normal section, and x is the height of the concrete compression zone in the equivalent rectangular stress diagram.
[0058] Based on the above formula, the selected initial feature x = [f c ,f y ,f′ y ,f py A s ,A′ s A p ,b,h0,a′ s Convert α1] to its basic factor form:
[0059]
[0060] S3: Collect 174 sets of measured scaled data of the bending capacity of the beam cross section, convert the beam characteristic values according to formula (3), and then establish the dataset D1 of the measured data. The form shown in formula (3) is used as the input item, and the measured bending capacity of the beam cross section is used as the output item.
[0061] S4: In order to fully embed the prior knowledge of the prediction of the flexural bearing capacity of beam cross sections in GB50010-2010 into the data-driven model, a sampling space is formed by all potential data value ranges. The SOBO sampling strategy is selected, and 200 groups are sampled from it to generate a representative dataset D2. Its input features are in the form shown in formula (3), and the output is the calculated value of the empirical formula.
[0062] S5: As Figure 2 As shown, a neural network was chosen as the original data-driven model. Based on dataset 22, hyperparameter optimization and training of the model were performed on it, ultimately resulting in a pre-trained model. w0 and θ0 are the weights and hyperparameters of the pre-trained model, respectively.
[0063] S6: As Figure 3 As shown, the measured dataset D1 is introduced into the pre-trained model G, and incremental learning is used to retrain the model. During training, the model is considered in the dataset D1 simultaneously, ultimately resulting in a data-driven generalization enhancement-based prediction model for the flexural capacity of beam sections. The model's performance on the training and test sets is as follows. Figure 5 As shown.
[0064] S7: Collect 12 sets of full-scale measured data on the flexural capacity of the beam's cross-section. Use the prediction model obtained in S6 to predict its capacity. The process is as follows: Figure 4 As shown. Empirical formula prediction values were calculated using both formulas (1) and (2). Furthermore, a classic machine learning algorithm neural network was optimized and trained on dataset D1, and the optimized model was used to predict the bearing capacity of full-scale beams. The results of these three prediction methods are plotted as follows. Figure 6 As shown, the method proposed in this invention has the highest prediction accuracy and the lowest dispersion, and its performance exceeds that of traditional empirical formulas and typical machine learning algorithms.
[0065] This invention is applicable to the performance prediction of various structural components, including but not limited to the prediction of bending, shear, and torsional resistance of structural components such as slabs, beams, columns, and walls. In use, it only requires that existing empirical formulas for structural performance calculations be available, and that a certain amount of measured data on structural performance be collected. This method enables more accurate predictions of structural component performance. Compared to the predictions of classical machine learning algorithms and empirical formulas, the prediction accuracy of this method is significantly improved. Furthermore, the method proposed in this patent can be used to assist in structural performance assessment, remaining service life prediction, durability analysis, and structural maintenance. In the future, the data-driven structural component performance prediction method established based on this patent can become a new approach for developing data-driven models and promote new developments in civil engineering.
[0066] In another embodiment of the present invention, a component performance prediction system based on data-driven generalization enhancement is provided, which can be used to implement the above-mentioned component performance prediction method based on data-driven generalization enhancement. Specifically, the system includes:
[0067] The feature determination module is used to determine the original model input features for predicting component performance;
[0068] The conversion module is used to select an empirical formula model and convert the original model input features into the basic factor form of the empirical formula according to the structure of the empirical formula.
[0069] The dataset creation module is used to collect the measured dataset of structural component performance prediction and convert it into the basic factor form of empirical formulas to establish the dataset of measured data; the basic factors of empirical formulas are used as input terms of the model, and the measured structural performance indicators are used as output terms of the model.
[0070] The input module is used to select a representative sampling strategy to generate a representative dataset; its input features are the basic factor forms of the empirical formula, and its output is the calculated value of the empirical formula.
[0071] The prediction module is used to train a model on a classic data-driven model based on a representative dataset to obtain a pre-trained model. The actual test dataset is then introduced into the pre-trained model, and incremental learning is used to train the model again to obtain a data-driven generalization enhancement-based structural component performance prediction model for prediction.
[0072] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0073] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions from the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a component performance prediction method based on data-driven generalization enhancement.
[0074] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the component performance prediction method based on data-driven generalization enhancement in the above embodiments.
[0075] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0076] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0077] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0078] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A component performance prediction method based on data-driven generalization enhancement, characterized in that, include Determine the original model input features for predicting component performance; Select an empirical formula model and transform the input features of the original model into the basic factor form of the empirical formula according to the structure of the empirical formula; Collect the measured dataset of structural component performance prediction and convert it into the basic factor form of empirical formulas to establish the dataset of measured data. The model uses the basic factors of empirical formulas as input terms and the measured structural performance indicators as output terms. Select a representative sampling strategy to generate a representative dataset; Its input features are the basic factor forms of empirical formulas, and its output is the calculated value of the empirical formulas; Based on a representative dataset, a pre-trained model is obtained by training the model on a classic data-driven model. The measured dataset is introduced into the pre-trained model, and incremental learning is used to train the model again to obtain a data-driven generalization enhancement-based structural component performance prediction model for prediction.
2. The component performance prediction method based on data-driven generalization enhancement according to claim 1, characterized in that, The original model input features include: concrete compressive strength, tensile and compressive strength of steel reinforcement in the tension and compression zones, tensile strength of prestressed tendons in the tension zone, cross-sectional area of longitudinal ordinary steel reinforcement in the tension and compression zones, cross-sectional area of longitudinal prestressed tendons in the tension zone, cross-sectional width, cross-sectional height, distance from the resultant point of longitudinal ordinary steel reinforcement in the compression zone to the compression edge of the cross-section, and coefficients when simplifying the equivalent rectangular stress diagram of concrete in the compression zone; the polynomial composition of the empirical formula includes the following basic factors: constant terms or exponential terms.
3. The component performance prediction method based on data-driven generalization enhancement according to claim 1, characterized in that, The collected measured dataset of structural component performance is not required to cover all potential data ranges.
4. The component performance prediction method based on data-driven generalization enhancement according to claim 1, characterized in that, When sampling based on empirical formulas and prior knowledge, the sampling space should cover all potential data value ranges. Sampling strategies used include SOBO sampling and Latin hypercube sampling.
5. The component performance prediction method based on data-driven generalization enhancement according to claim 1, characterized in that, When optimizing and training the hyperparameters of the data-driven model, random search, grid search, and Bayesian optimization algorithms are used.
6. The component performance prediction method based on data-driven generalization enhancement according to claim 1, characterized in that, When incrementally training the established model, the performance metrics of the model on datasets D1 and D2 are considered simultaneously. The weight ratio of the metrics on the two datasets is determined based on the performance of the empirical formula when predicting component performance. When the performance of the empirical formula is better, the weight on dataset D2 is increased accordingly.
7. The component performance prediction method based on data-driven generalization enhancement according to claim 1, characterized in that, pre-trained model w0 and θ0 are the model's weights and hyperparameters, respectively. The prediction process is as follows: collect the characteristic indicators of the structural components to be predicted, transform the input form based on the basic factor form of the empirical formula, input the trained data to drive the model, and obtain the prediction results.
8. A component performance prediction system based on data-driven generalization enhancement, characterized in that, include: The feature determination module is used to determine the original model input features for predicting component performance; The conversion module is used to select an empirical formula model and convert the original model input features into the basic factor form of the empirical formula according to the structure of the empirical formula. The dataset creation module is used to collect the measured dataset of structural component performance prediction and convert it into the basic factor form of empirical formulas to establish the dataset of measured data. The model uses the basic factors of empirical formulas as input terms and the measured structural performance indicators as output terms. The input module is used to select a representative sampling strategy to generate a representative dataset; Its input features are the basic factor forms of empirical formulas, and its output is the calculated value of the empirical formulas; The prediction module is used to train a model on a classic data-driven model based on a representative dataset to obtain a pre-trained model. The actual test dataset is then introduced into the pre-trained model, and incremental learning is used to train the model again to obtain a data-driven generalization enhancement-based structural component performance prediction model for prediction.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the component performance prediction method based on data-driven generalization enhancement as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the component performance prediction method based on data-driven generalization enhancement as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Improvements in electric lighting fittings
GB540055A
Polypropylene melt index predicating method based on multiple priori knowledge mixed model
CN102609593A
Conformance test method for estimating probability fatigue life of structure dangerous point
CN112597682A