Construction project cost data disaster recovery method and disaster recovery management system
By using neural networks to predict and allocate disaster recovery cost data of construction projects, the traditional methods have solved the shortcomings in real-time, completeness and intelligent decision-making of data, and achieved efficient and accurate data disaster recovery processing.
Patent Information
- Application Number
- CN202510020626.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
When traditional data disaster recovery methods face large-scale and high-complex construction project cost data, it is difficult to ensure the real-time and integrity of the data, and the lack of intelligent decision-making mechanisms, resulting in waste or insufficient backup resources.
By obtaining the sample construction project cost data sequence and initializing the neural network, disaster recovery backup prediction is carried out, training disaster recovery backup allocation space data is generated, and valid sample data is screened through global network learning errors for network parameter learning, and the trained disaster recovery backup model is obtained.
It realizes accurate and efficient disaster recovery backup and processing of construction project cost data, improves the accuracy and efficiency of data disaster recovery, and ensures the security and reliability of data.
Smart Images

Figure CN119938405A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a construction engineering cost data disaster recovery method and a disaster recovery management system. Background Art
[0002] In the management process of construction projects, cost data is crucial information, which is directly related to the cost budget, fund allocation and economic benefit evaluation of the project. However, due to various force majeure factors, such as natural disasters, system failures or human errors, cost data is at risk of loss or damage. Once these data are lost or damaged, it will cause great trouble to the smooth progress of the project and may even lead to serious economic losses.
[0003] Traditional data disaster recovery methods often rely on simple backup strategies, such as regularly copying data to other storage devices or remote servers. However, these methods often seem to be inadequate when faced with large-scale, highly complex construction project cost data. On the one hand, simple backup strategies are difficult to ensure the real-time and integrity of data, and data may be lost or erroneous during the backup process; on the other hand, traditional disaster recovery methods lack intelligent decision-making mechanisms and cannot flexibly allocate backups according to the actual needs and importance of data, resulting in waste or insufficiency of backup resources. Summary of the invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a construction project cost data disaster recovery method, the method comprising:
[0005] Acquire a sample construction project cost data sequence and initialize a neural network; the sample construction project cost data sequence includes a plurality of candidate sample construction project cost data; each candidate sample construction project cost data carries disaster recovery backup allocation space data annotation data;
[0006] Perform disaster recovery backup prediction on each candidate sample construction project cost data according to the initialized neural network, and generate training disaster recovery backup allocation space data for each candidate sample construction project cost data;
[0007] For each of the candidate sample construction project cost data, determine the global network learning error according to the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data, and extract valid sample construction project cost data from the sample construction project cost data sequence according to the global network learning error of each of the candidate sample construction project cost data;
[0008] Performing network parameter learning on the initialized neural network based on the extracted valid sample construction project cost data, and outputting a trained disaster recovery backup model;
[0009] A construction project cost data set to be processed for data disaster recovery is obtained, and disaster recovery backup processing is performed on the construction project cost data set according to the trained disaster recovery backup model and the predicted disaster recovery backup allocation space data.
[0010] On the other hand, an embodiment of the present invention also provides a disaster recovery management system, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0011] Based on the above aspects, the embodiment of the present application performs disaster recovery backup prediction on multiple candidate data in the sample construction project cost data sequence by initializing a neural network, and generates training disaster recovery backup allocation space data. The global network learning error is determined by comparing the predicted data with the labeled data, thereby screening out valid sample data for network parameter learning to obtain a trained disaster recovery backup model. This method can accurately and efficiently perform disaster recovery backup processing on construction project cost data, improve the accuracy and efficiency of data disaster recovery, and ensure the security and reliability of construction project cost data. In practical applications, the trained disaster recovery backup model is used to perform disaster recovery backup processing on the construction project cost data set to be processed, effectively avoiding the risk of data loss or damage. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the execution flow of the construction project cost data disaster recovery method provided by an embodiment of the present invention.
[0013] Figure 2 It is a schematic diagram of the hardware architecture of the disaster recovery management system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 It is a flow chart of a construction project cost data disaster recovery method provided by an embodiment of the present invention. The construction project cost data disaster recovery method is introduced in detail below.
[0015] Step S110, obtaining a sample construction project cost data sequence and initializing a neural network. The sample construction project cost data sequence includes a plurality of candidate sample construction project cost data. Each candidate sample construction project cost data carries disaster recovery backup allocation space data annotation data.
[0016] In this embodiment, in the scenario of construction project cost management, a certain construction project database stores cost data of many past construction projects. These cost data cover information from the purchase cost of basic building materials, the rental and purchase costs of various equipment, human resource costs to various management costs. Some of the data in this database constitutes a sample construction project cost data sequence. For example, the cost data of one of the construction projects includes the cost of materials such as concrete, steel, and wood, the labor costs of the construction team, the rental costs of large machinery and equipment, and for this project, based on past experience or evaluation after actual operation, a more reasonable disaster recovery backup allocation space data labeling data has been determined, such as setting the space required for disaster recovery backup for various types of cost data according to a certain ratio, such as concrete cost data may be allocated 1% of the total cost backup space, and labor cost data may be allocated 3% of the backup space. The cost data of this construction project becomes a candidate sample construction project cost data. At the same time, the construction of the initialization neural network is based on a specific algorithm architecture, such as a multi-layer perceptron (MLP) structure, including an input layer, several hidden layers, and an output layer. The number of neurons in the input layer is determined according to the number of features in the sample construction project cost data. For example, the number of neurons in the input layer is determined according to the type of material, type of cost, etc. The number of neurons and the number of layers in the hidden layer are preliminarily set based on experience or preliminary experiments. The output layer is used to output the prediction results related to the disaster recovery backup allocation space data.
[0017] Step S120, performing disaster recovery backup prediction on each candidate sample construction project cost data according to the initialized neural network, and generating training disaster recovery backup allocation space data for each candidate sample construction project cost data.
[0018] Taking the construction project cost data mentioned above as an example, this candidate sample construction project cost data is input into the initialized neural network. Inside the neural network, the data first enters the input layer, and each neuron receives the corresponding data feature value. Then, the data undergoes complex calculations and transformations in the hidden layer. Assuming that the hidden layer uses the sigmoid activation function, when the data is transmitted between neurons in the hidden layer, it will be linearly combined according to the connection weights between neurons, and then nonlinearly transformed by the sigmoid function. For the output layer, it further calculates based on the output results of the hidden layer to generate training disaster recovery backup allocation space data for this candidate sample construction project cost data. This training disaster recovery backup allocation space data may contain logical relationships such as the priority ranking of different cost components during disaster recovery backup, and the backup space allocation ratio of different cost data after risk assessment. For example, for the material cost part, the neural network gives a backup space allocation logic structure classified by different materials based on factors such as the risk of material price fluctuations and the stability of material supply. For example, for rare metal materials with large price fluctuations and unstable supply, a larger proportion of backup space with a higher priority is allocated; for common building materials with relatively stable prices and sufficient supply, a smaller proportion of backup space with a lower priority is allocated.
[0019] Step S130, for each of the candidate sample construction project cost data, determine the global network learning error based on the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data annotation data, and extract valid sample construction project cost data from the sample construction project cost data sequence based on the global network learning error of each of the candidate sample construction project cost data.
[0020] Continuing with the construction project cost data as an example, for each candidate example, first determine the first network learning error corresponding to each initialized neural network. For example, for a candidate example, the backup space ratio of a certain material cost given in its training disaster recovery backup allocation space data is 20%, while the backup space ratio of the material cost in its disaster recovery backup allocation space data annotation data is 15%. This part of the error is calculated by a specific error calculation function (such as a mean square error function), and then the errors of all cost components are combined to obtain the first network learning error corresponding to this initialized neural network. If there are multiple initialized neural networks, for example two, calculate the first network learning error corresponding to each neural network respectively. Next, determine the relative learning error between each two training disaster recovery backup allocation space data. Assume that there is a difference in the backup space allocation of equipment rental costs in the training disaster recovery backup allocation space data of the same candidate example for two initialized neural networks, and obtain the relative learning error by calculating the quantitative value of this difference. The sum of each first network learning error and relative learning error is determined as the global network learning error.
[0021] When extracting valid sample construction project cost data, it is assumed that there are multiple extraction methods. One method is to mark the sample construction project cost data with a set percentage of the smallest global network learning error as valid sample construction project cost data. For example, from 100 candidate samples, select 20% (i.e. 20) of the samples with the smallest global network learning error as valid samples. Another method is to mark the sample construction project cost data with a global network learning error less than the set error as valid sample construction project cost data. For example, if the set error is 0.1, if the global network learning error of a candidate sample is 0.08, then this sample is marked as a valid sample. These valid samples are considered to be more valuable data in the neural network learning process after comparison with the labeled data and error analysis among many candidate samples.
[0022] Step S140 , performing network parameter learning on the initialized neural network according to the extracted valid sample construction project cost data, and outputting a trained disaster recovery backup model.
[0023] Taking the valid sample construction project cost data selected previously as an example, network parameter learning is performed on multiple initialized neural networks based on these valid sample construction project cost data. In the process of network parameter learning, the connection weights and other parameters of the neural network are adjusted through the back propagation algorithm. For example, for a certain construction project cost data in the valid sample, according to its input features and the corresponding disaster recovery backup allocation space data annotation data, the error between the current neural network output result and the annotation data is calculated, and then this error is back-propagated from the output layer to the hidden layer and the input layer. During the propagation process, the connection weights between neurons are adjusted according to a certain learning rate. Assuming that for a certain neuron connection, its current weight is 0.5, and the amount to be adjusted is calculated to be -0.1 based on the error and the learning rate, then the adjusted weight becomes 0.4.
[0024] After learning the network parameters of multiple initialized neural networks, multiple initialized neural networks are generated. Then, the performance of these neural networks is verified. Performance verification can be performed in many ways, such as using new sample construction project cost data that has not participated in the training to input into these neural networks, and comparing the errors between their prediction results and the actual labeled data. Assuming that one of the neural networks has a small prediction error for most samples in the test of new samples, while the prediction error of another neural network is large, then the neural network with a small prediction error is determined from these multiple initialized neural networks as the disaster recovery backup model after training.
[0025] Step S150, obtaining a construction project cost data set to be processed for data disaster recovery, and performing disaster recovery processing on the construction project cost data set based on the trained disaster recovery backup model and the predicted disaster recovery backup allocation space data.
[0026] In this embodiment, it is assumed that there is a new large-scale construction project, and its construction engineering cost data set contains detailed cost information. First, the trained disaster recovery backup model is used to process the construction engineering cost data set.
[0027] The disaster recovery backup model is used to extract the characteristic knowledge distribution data of the construction project cost data set. The specific process is as follows:
[0028] The construction project cost data set is subjected to data cleaning operations to remove possible noise data (e.g., abnormally high or low cost data of a certain item due to data entry errors), erroneous data (e.g., data errors caused by incorrectly writing material price units), and incomplete data (e.g., a certain cost item has only partial cost records but lacks other relevant cost records). The construction project cost data set that has undergone data cleaning operations is then standardized, for example, data of different units are unified into the same standard unit, and then the standardized construction project cost data set is encoded to generate a data encoding result of the construction project cost data set. The data encoding result is input into the initial feature extraction layer of the disaster recovery backup model. Assuming that there are 10 neurons in the initial feature extraction layer, the data encoding result is linearly combined by these 10 neurons to obtain an intermediate feature extraction result. For example, the outputs of the 10 neurons are added according to certain weights to obtain an intermediate result, and then the intermediate feature extraction result is activated (e.g., using a ReLU activation function) to generate an initial feature set.
[0029] The feature crossover method is used to combine the features in the initial feature set to generate a combined feature sequence. For example, the material cost feature and the labor cost feature are cross-combined to obtain a new combined feature. Then, the combined feature sequence is screened based on the feature importance evaluation index, which is obtained by calculating the correlation between each feature and the disaster recovery backup target. Assuming that after calculating the correlation, it is found that some combined features have a low correlation with the disaster recovery backup target, these features are removed from the combined feature sequence to obtain a screened combined feature sequence.
[0030] The screened combined feature sequence is input into the deep feature mining layer in the disaster recovery backup model, and the deep feature mining layer consists of 3 hidden layers. The screened combined feature sequence is input into the first layer of the deep feature mining layer, which has 8 neurons. The screened combined feature sequence is linearly combined through these 8 neurons, and the linear combination operation result is activated (such as using tanh activation function), and the output result is passed to the next hidden layer to repeat the above linear combination and activation function processing process, and the output result of the last hidden layer is used as the deep feature set.
[0031] Perform semantic analysis on the deep features in the deep feature set, for example, analyze the semantic meaning of material cost features in different disaster recovery scenarios. After grouping the deep features with similar semantics according to the results of the semantic analysis, perform knowledge fusion operations on the grouped deep features to generate feature knowledge subsets. For example, the deep features related to price fluctuations of different building materials are fused into a feature knowledge subset. Combine all feature knowledge subsets to generate an integrated feature knowledge set. Determine the dimensions of feature knowledge distribution, assuming that 5 dimensions are determined, corresponding to different cost components and disaster recovery related factors. Perform statistics on the integrated feature knowledge set according to the determined dimensions, and based on the statistical results, use a matrix data structure to construct the structure of the feature knowledge distribution data, where rows represent different dimensions and columns represent specific attributes of the feature knowledge under each dimension.
[0032] Determine the characteristic focusing factor of the characteristic knowledge distribution data, and extract the focused knowledge distribution data of the characteristic knowledge distribution data according to the characteristic focusing factor. Specifically as follows:
[0033] Determine the feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension. First, analyze the knowledge content related to the disaster recovery knowledge dimension classification in the feature knowledge distribution data, such as the knowledge content related to the backup frequency, backup storage location, data recovery strategy, etc. Take each feature in the knowledge content classified under the disaster recovery knowledge dimension as a node to build a feature association network. For example, for features such as the material cost backup frequency and the equipment rental cost backup frequency under the backup frequency knowledge dimension, specifically analyze the direct association relationship between each feature (such as the material cost backup frequency and the equipment rental cost backup frequency may be directly associated due to the association with the project progress) and indirect association relationship (such as indirect association through the overall project risk assessment), and connect the relevant feature nodes with connecting edges according to the direct and indirect association relationship between the features to generate the feature association network, which is used to reflect the relationship between the various features in the knowledge content classified under the disaster recovery knowledge dimension.
[0034] For each node in the feature association network, based on the degree value of the node (e.g., a node is connected to 3 other nodes, and the degree value is 3), betweenness centrality (indicates the frequency of a node appearing on the shortest path between other nodes, for example, a node often appears on the shortest path between other nodes, and the betweenness centrality is high) and closeness centrality (reflects the average distance from a node to other nodes, such as a node with a shorter average distance from other nodes, and the closeness centrality is high), evaluate the importance of the node in the feature association network and generate a first feature importance evaluation result. Obtain a pre-set disaster recovery strategy, such as using high-frequency backup and safe storage location strategies for important cost data, and adjust the first feature importance evaluation result according to the disaster recovery strategy to generate an adjusted feature importance evaluation result. Determine a benchmark value, such as the average value of all feature importances. For each node, a calculation rule is constructed based on the relationship between the adjusted importance evaluation result of the node and the benchmark value, and different dimension weights are set for the backup frequency knowledge dimension, the backup storage location knowledge dimension, and the data recovery strategy knowledge dimension, such as the backup frequency knowledge dimension weight is 0.4, the backup storage location knowledge dimension weight is 0.3, and the data recovery strategy knowledge dimension weight is 0.3, and a calculation framework is constructed based on the calculation rules and dimension weights. The calculation framework is called to calculate the feature focus factor of each node under the disaster recovery knowledge dimension based on the adjusted feature importance evaluation result and the dimension weight of the disaster recovery knowledge dimension, and the feature focus factor of each node under the disaster recovery knowledge dimension is summarized to determine the feature focus factor of the feature knowledge distribution data in the disaster recovery knowledge dimension.
[0035] The focused knowledge distribution data is determined based on the feature knowledge distribution data and the feature focusing factor of the disaster recovery knowledge dimension. First, the data structure of the feature knowledge distribution data is parsed, and based on the data structure of the feature knowledge distribution data, the feature elements associated with the disaster recovery knowledge dimension are searched, for example, in the data of the matrix structure, the elements in the rows and columns related to the backup frequency, the backup storage location, and the data recovery strategy are searched, and all the feature elements related to the disaster recovery knowledge dimension are extracted by searching and matching in the parsed data structure to generate a feature element set. The feature focusing factor of the disaster recovery knowledge dimension is used to assign weights to each feature element in the feature element set to generate a weight sequence of the feature element set, and the feature focusing factor is used to measure the index of the importance of each feature element under the disaster recovery knowledge dimension. According to a preset weight threshold or weight ratio, according to the weight sequence of the feature element set, the feature element set is screened to generate an optimized feature element subset. The data structure is reconstructed based on the screened feature element subset to generate the focused knowledge distribution data.
[0036] Disaster recovery backup prediction is performed on the focused knowledge distribution data to generate the confidence of the construction project cost data set corresponding to multiple reference disaster recovery backup allocation space data. For example, for a certain cost component, the confidence under different backup space allocation schemes may be obtained, such as the confidence of allocating 10% backup space is 0.6, the confidence of allocating 15% backup space is 0.8, etc. According to the confidence of the construction project cost data set corresponding to multiple reference disaster recovery backup allocation space data, the predicted disaster recovery backup allocation space data of the construction project cost data set is determined. This predicted disaster recovery backup allocation space data comprehensively considers the backup space allocation schemes of different cost components under different confidence levels. For example, for the material cost part, according to its confidence under different backup space allocation schemes, a backup space allocation logic structure divided by material category and risk level is determined; for the labor cost part, a backup space allocation logic structure based on job importance and personnel turnover risk is also determined. Finally, disaster recovery backup processing is performed on the construction project cost data set according to the predicted disaster recovery backup allocation space data, and the data is backed up to a corresponding disaster recovery storage location according to a determined logical structure.
[0037] Based on the above steps, the embodiment of the present application performs disaster recovery backup prediction on multiple candidate data in the sample construction project cost data sequence by initializing the neural network, and generates training disaster recovery backup allocation space data. The global network learning error is determined by comparing the predicted data with the labeled data, so as to screen out the effective sample data for network parameter learning and obtain the trained disaster recovery backup model. This method can accurately and efficiently perform disaster recovery backup processing on the construction project cost data, improve the accuracy and efficiency of data disaster recovery, and ensure the security and reliability of the construction project cost data. In practical applications, the trained disaster recovery backup model is used to perform disaster recovery backup processing on the construction project cost data set to be processed, which effectively avoids the risk of data loss or damage.
[0038] In a possible implementation, the initialized neural network includes two or more.
[0039] Step S120 includes: loading each candidate sample construction project cost data into multiple initialized neural networks respectively, and generating training disaster recovery backup allocation space data corresponding to each initialized neural network for each candidate sample construction project cost data.
[0040] For each candidate sample construction project cost data, step S130 includes:
[0041] Step S131, determining a first network learning error corresponding to each of the initialized neural networks according to the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data corresponding to each of the initialized neural networks.
[0042] Step S132: determining the global network learning error of the candidate sample construction project cost data based on the first network learning error corresponding to each of the initialized neural networks.
[0043] In a possible implementation, step S140 includes:
[0044] Step S141, performing network parameter learning on the multiple initialized neural networks respectively according to the valid sample construction project cost data to generate multiple initialized neural networks.
[0045] Step S142, determining the trained disaster recovery backup model from the multiple initialized neural networks by performing performance verification on the multiple initialized neural networks.
[0046] In a possible implementation, step S132 includes:
[0047] Step S1321, determining the relative learning error between every two of the training disaster recovery backup allocation space data.
[0048] Step S1322: Determine the sum of each of the first network learning errors and the relative learning error as the global network learning error.
[0049] In this embodiment, taking the cost data of a large-scale construction project as an example, it is assumed that there are three initialized neural networks, namely neural network A, neural network B and neural network C. For a specific construction project cost data, it includes various detailed cost information such as the cost of building structure materials, the cost of electromechanical equipment, and the salary of construction personnel. When the candidate sample construction project cost data is loaded into neural network A, the input layer of neural network A receives the characteristic values of these cost data, and then calculates through its internal hidden layer. The neurons in the hidden layer linearly combine the input data according to the preset weights, and after the nonlinear transformation of the activation function, finally obtain the training disaster recovery backup allocation space data for this candidate sample construction project cost data at the output layer. This training disaster recovery backup allocation space data is a logical structure data related to the disaster recovery backup allocation space obtained based on the internal algorithm and weight calculation of neural network A, and may include different allocation ratios and priority relationships of different cost parts in disaster recovery backup. Similarly, when the candidate sample construction project cost data is loaded into neural network B and neural network C, the corresponding training disaster recovery backup allocation space data will be obtained respectively. Due to the differences in the initial weights and structures of each neural network, the training disaster recovery backup allocation space data obtained by the three neural networks may be different in structure and value.
[0050] For each candidate sample construction project cost data, the global network learning error is determined based on the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data annotation data. For the candidate sample construction project cost data mentioned above, the first network learning error corresponding to each initialized neural network is determined based on the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data annotation data corresponding to each initialized neural network. For example, for neural network A, in the training disaster recovery backup allocation space data output by it, the disaster recovery backup allocation ratio of the building structure material cost is 10%, and the disaster recovery backup allocation ratio of the building structure material cost in the disaster recovery backup allocation space data annotation data is 8%. The error of this part is calculated by a specific error calculation function (such as a mean square error function), and then the errors of all cost components are comprehensively calculated to obtain the first network learning error corresponding to neural network A. Similarly, similar calculations are performed on neural network B and neural network C to obtain their respective first network learning errors. Then, based on the first network learning errors corresponding to each initialized neural network, the global network learning error of the candidate sample construction project cost data is determined. Specifically, the relative learning error between each two training disaster recovery backup allocation space data is determined. Assuming that the proportion of disaster recovery backup allocation for building structure material costs of neural network A and neural network B is 10% and 12% respectively, the relative learning error of this part is obtained by calculating the difference between the two and quantifying the difference according to certain rules. Then the sum of the first network learning error of neural network A, neural network B and neural network C and the calculated relative learning error is determined as the global network learning error of this candidate sample construction project cost data.
[0051] Based on the extracted valid sample construction project cost data, the network parameters of the initialized neural network are learned to generate a trained disaster recovery backup model. In the previous operation, the valid sample construction project cost data has been extracted according to the global network learning error. Now, based on these valid sample construction project cost data, the network parameters of multiple initialized neural networks are learned respectively to generate multiple initialized neural networks. Taking one of the construction project cost data in these valid samples as an example, for neural network A, this valid sample construction project cost data is input into neural network A, and according to the error between its output result and the corresponding disaster recovery backup allocation space data labeling data, the parameters such as the neuron connection weight in neural network A are adjusted through the back propagation algorithm. For example, the initial weight of a certain neuron connection in neural network A is 0.3, and the weight change amount to be adjusted according to the error calculation is -0.05, then the adjusted weight becomes 0.25. The same operation is performed on neural network B and neural network C, thereby generating multiple initialized neural networks after network parameter learning. Then, by performing performance verification on multiple initialized neural networks, the trained disaster recovery backup model is determined from multiple initialized neural networks. During the performance verification process, a new set of construction project cost data that has not participated in the previous training process is selected as the verification set. These verification set data are respectively input into neural network A, neural network B, and neural network C after network parameter learning, and the errors between their respective prediction results and the actual disaster recovery backup allocation space data annotation data are calculated. Assuming that the average error of neural network A on the verification set data is 0.08, the average error of neural network B is 0.12, and the average error of neural network C is 0.1, then since the average error of neural network A is the smallest, neural network A is determined as the trained disaster recovery backup model.
[0052] When determining the global network learning error of the candidate sample construction project cost data, continue to take the previous neural network A, neural network B and neural network C as examples. For the disaster recovery backup allocation space data of the building mechanical and electrical equipment cost, the output of neural network A is 15%, and the output of neural network B is 13%. The difference between the two is calculated to be 2%, and then according to certain quantitative rules (such as factors such as the proportion of this part of the cost in the entire cost data), this difference is converted into a relative learning error value. The same operation is performed for other cost components to obtain the relative learning error between each two training disaster recovery backup allocation space data. Then, the sum of each first network learning error and these relative learning errors is determined as the global network learning error. This global network learning error comprehensively considers the error of each initialized neural network itself and the relative error caused by the output difference between them, and can more comprehensively reflect the learning effect of the candidate sample construction project cost data in multiple initialized neural networks. This method helps to accurately extract valid sample construction project cost data in subsequent steps, thereby improving the accuracy and reliability of the trained disaster recovery backup model.
[0053] In a possible implementation, step S140 may also include: performing iterative network parameter learning on the initialized neural network based on the extracted valid sample construction project cost data to generate a first temporary neural network, and if the first temporary neural network meets the network convergence requirements, generating the trained disaster recovery backup model based on the first temporary neural network.
[0054] If the first temporary neural network does not meet the network convergence requirement, the first temporary neural network after the network parameter learning is used as an iterative neural network, and the following steps are iteratively performed until the generated second temporary neural network meets the network convergence requirement, and the trained disaster recovery backup model is generated according to the second temporary neural network that meets the network convergence requirement:
[0055] Step A110, obtaining an iterative sample construction project cost data sequence.
[0056] Step A120, performing disaster recovery backup prediction on each candidate sample construction project cost data in the iterative sample construction project cost data sequence according to the initialized neural network, and generating training disaster recovery backup allocation space data for each candidate sample construction project cost data.
[0057] Step A130: for each candidate sample construction project cost data, determine the global network learning error according to the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data.
[0058] Step A140: extracting valid sample construction project cost data from the sample construction project cost data sequence based on the global network learning error of each candidate sample construction project cost data.
[0059] Step A150, iteratively learning network parameters of the initialized neural network based on the extracted valid sample construction project cost data to generate a second temporary neural network. If the second temporary neural network does not meet the network convergence requirements, the second temporary neural network is used as an iterative neural network.
[0060] In this embodiment, taking the construction project cost data mentioned above as an example, it is assumed that valid sample data has been extracted from many candidate sample construction project cost data. For parameters such as the neuron connection weight in the initialized neural network, these valid sample construction project cost data are used for iterative adjustment. For example, a certain construction project cost data in the valid sample contains detailed information such as infrastructure costs, decoration costs, and equipment purchase costs, and this valid sample construction project cost data is input into the initialized neural network. The neural network processes the data according to the structure and algorithm of its input layer, hidden layer, and output layer. During the processing, the neuron connection weight is adjusted according to the error between the output result and the actual disaster recovery backup allocation space data labeling data through the back propagation algorithm. For the connection weight between a certain neuron in the hidden layer and the output layer neuron, the initial weight may be 0.4, and the amount of weight adjustment obtained by the error calculation is -0.03, so after one iteration, the weight becomes 0.37. After multiple such iterations, after processing multiple valid sample construction project cost data, a first temporary neural network is generated. At this time, it is necessary to determine whether the first temporary neural network meets the network convergence requirements. The network convergence requirement can be defined in many ways, such as the network output error after several consecutive iterations is less than a certain set threshold, or the change in network parameters is within a certain range. If the first temporary neural network meets the network convergence requirement, a trained disaster recovery backup model is generated based on the first temporary neural network. This trained disaster recovery backup model will be used for subsequent disaster recovery backup prediction operations such as the construction project cost data set.
[0061] If the first temporary neural network does not meet the network convergence requirements, then the first temporary neural network after the network parameter learning is used as the iterative neural network, and the subsequent steps are iteratively executed. First, an iterative sample construction project cost data sequence is obtained. This iterative sample construction project cost data sequence can be part of the candidate sample construction project cost data that has not been fully utilized before, or newly collected data related to the construction project cost. For example, some new construction project cost data of different building types (such as residential buildings, commercial buildings, etc.) or different regions are newly acquired, and these data constitute an iterative sample construction project cost data sequence.
[0062] Next, the disaster recovery backup prediction is performed on each candidate sample construction project cost data in the iterative sample construction project cost data sequence according to the initialized neural network, and the training disaster recovery backup allocation space data of each candidate sample construction project cost data is generated. Take a commercial building project cost data in the iterative sample construction project cost data sequence as an example, this data contains cost information such as land acquisition cost, commercial facility construction cost, and marketing cost. It is input into the initialized neural network, and the initialized neural network is calculated according to its own structure and algorithm. After the input layer receives the characteristic values of these cost data, the neurons in the hidden layer are linearly combined and activated. The function is processed, and finally the training disaster recovery backup allocation space data for this candidate sample construction project cost data is generated at the output layer, which contains the allocation logic of different cost parts in disaster recovery backup, such as the land acquisition cost determines the corresponding disaster recovery backup allocation ratio and priority according to its importance and risk factors in the project (such as land policy change risk, etc.).
[0063] For each candidate sample construction project cost data, the global network learning error is determined based on the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data annotation data. For example, for the commercial facility construction cost part of the above-mentioned commercial building project cost data, the allocation ratio in the training disaster recovery backup allocation space data is 12%, and the allocation ratio in the disaster recovery backup allocation space data annotation data is 10%. The error of this part is calculated by a specific error calculation function (such as a mean square error function), and then the error of all cost components is combined to obtain the error of the candidate sample construction project cost data corresponding to the initialization neural network. According to the method mentioned above, the relative learning error between each two training disaster recovery backup allocation space data is determined, and the sum of each first network learning error and the relative learning error is determined as the global network learning error.
[0064] Valid sample construction project cost data are extracted from the sample construction project cost data sequence based on the global network learning error of each candidate sample construction project cost data. For example, a method can be adopted in which the sample construction project cost data with a set proportion of the smallest global network learning error is marked as valid sample construction project cost data. Assuming that there are 100 candidate samples in the iterative sample construction project cost data sequence, and the set proportion is 20%, then the 20 candidate samples with the smallest global network learning error are marked as valid sample construction project cost data. Or a method is adopted in which the sample construction project cost data with a global network learning error less than the set error is marked as valid sample construction project cost data, such as setting the error to 0.1, if the global network learning error of a candidate sample is 0.08, then this sample is marked as valid sample construction project cost data.
[0065] According to the extracted valid sample construction project cost data, the initialized neural network is iteratively learned to generate a second temporary neural network. Taking a residential construction project cost data from these newly extracted valid sample construction project cost data as an example, it is input into the initialized neural network, and according to the error between the output result and the actual disaster recovery backup allocation space data labeling data, the neural network's neuron connection weight and other parameters are adjusted by the back propagation algorithm according to the network parameter learning method mentioned above. After multiple iterations, multiple valid sample construction project cost data are processed to generate a second temporary neural network. If the second temporary neural network does not meet the network convergence requirements, the second temporary neural network is used as an iterative neural network, and the above steps are continued to be iteratively executed until the generated temporary neural network meets the network convergence requirements, and the trained disaster recovery backup model is generated according to the temporary neural network that meets the network convergence requirements. This process continuously optimizes the parameters of the neural network, so that it can more accurately predict the disaster recovery backup allocation space data, thereby improving the accuracy and reliability of the entire disaster recovery backup model. Through this iterative method, different sample construction project cost data can be fully utilized, and the parameters of the neural network can be continuously adjusted to adapt to various complex construction project cost situations, ensuring the accuracy and effectiveness of disaster recovery backup prediction.
[0066] In a possible implementation, step A140 includes at least one of the following:
[0067] Step A141, marking the sample construction project cost data with a set proportion and the smallest global network learning error as the valid sample construction project cost data.
[0068] In this embodiment, it is assumed that there is a large sample construction project cost data sequence, which contains cost data of many construction projects, and these projects cover different types of buildings (such as residential buildings, commercial buildings, industrial buildings, etc.) and different construction scales. For each candidate sample construction project cost data, its global network learning error has been calculated. For example, there are 100 candidate sample construction project cost data, and the proportion is set to 20%, that is, 20 samples are selected as valid samples. First, these 100 candidate samples are sorted according to the global network learning error. The global network learning error reflects the degree of deviation of each candidate sample in the neural network learning process. This error is a comprehensive consideration of the difference between the prediction result of the neural network for the sample and the actual disaster recovery backup allocation space data annotation data, the relative difference between the prediction results of different neural networks, and other factors. After sorting, select the top 20 candidate sample construction project cost data with the smallest global network learning error, and mark them as valid sample construction project cost data. These data marked as valid samples are of great significance in subsequent operations such as neural network parameter learning, because they show smaller deviations from the actual labeled data in the previous learning process, which can provide more accurate learning samples for the neural network and help improve the accuracy of the neural network.
[0069] Step A142, marking the sample construction project cost data whose global network learning error is less than the set error as the valid sample construction project cost data.
[0070] Similarly, in this sequence containing many candidate sample construction project cost data, an error value is pre-set, for example, the error is set to 0.1. Then the global network learning error of each candidate sample construction project cost data is checked one by one. Taking the cost data of one of the commercial building projects as an example, the cost data of this project includes multiple parts such as building structure cost, decoration cost, equipment installation cost, etc. The global network learning error is obtained after the error calculation between the prediction result of the neural network for this sample and the actual disaster recovery backup allocation space data labeling data. If the global network learning error of the cost data of this commercial building project is 0.08, since 0.08 is less than the set error 0.1, then the cost data of this commercial building project is marked as a valid sample construction project cost data. The other candidate samples are judged in the same way. As long as their global network learning error is less than the set error, they are marked as valid sample construction project cost data. The valid sample construction project cost data selected in this way meet specific requirements in terms of global network learning error, can provide a reliable data basis for further learning of the neural network, and help improve the accuracy and reliability of the disaster recovery backup model, so that the disaster recovery backup model finally trained can more accurately predict and process the construction project cost data set for disaster recovery backup.
[0071] In a possible implementation, step S150 includes:
[0072] Step S151, using the disaster recovery backup model, extracting characteristic knowledge distribution data of the construction project cost data set, determining a characteristic focusing factor of the characteristic knowledge distribution data, and extracting focused knowledge distribution data of the characteristic knowledge distribution data according to the characteristic focusing factor.
[0073] Step S152, performing disaster recovery backup prediction on the focused knowledge distribution data, and generating confidence levels of the construction project cost data set corresponding to a plurality of reference disaster recovery backup allocation space data.
[0074] Step S153, determining predicted disaster recovery backup allocation space data of the construction project cost data set according to the confidence levels of the plurality of reference disaster recovery backup allocation space data corresponding to the construction project cost data set.
[0075] In this embodiment, a construction project cost data set containing cost information of multiple construction projects is taken as an example. This construction project cost data set contains detailed cost data such as various construction material costs, labor costs, equipment rental costs, etc. The disaster recovery backup model first processes this data set and extracts characteristic knowledge distribution data through calculation. For example, for the cost of construction materials, the model will analyze the proportion of the cost of different materials (such as steel, cement, etc.) in the total cost, the fluctuation situation, and the correlation with other cost factors (such as construction progress, market supply, etc.), so as to construct the characteristic knowledge distribution data part about the cost of construction materials in this construction project cost data set. For labor costs, the corresponding characteristic knowledge distribution data part will be constructed by considering factors such as the salary level, working hours, and seasonal changes in manpower demand of different types of work (such as masons, electricians, etc.). After constructing the characteristic knowledge distribution data of the entire construction project cost data set, the characteristic focusing factor of the characteristic knowledge distribution data is then determined. Taking the cost of building materials as an example, we analyze its key factors in the disaster recovery backup scenario. For example, the supply stability of some rare materials is more important to disaster recovery backup, so these factors will be given higher weights in the calculation of the feature focus factor. The feature focus factor of the building material cost part is calculated based on these weight relationships. Similar calculations are performed for other parts of the entire construction project cost data set (such as labor costs, equipment rental costs, etc.) to obtain their respective feature focus factors. Then, based on these feature focus factors, the focused knowledge distribution data of the feature knowledge distribution data is extracted. For example, in the building material cost part, the knowledge data related to the factors that have a greater impact on disaster recovery backup (such as the stability of rare material supply, the price fluctuation trend of key materials, etc.) are selected according to the feature focus factors to form the focused knowledge distribution data of the building material cost part. The same operation is performed on the other parts of the entire construction project cost data set, and finally the complete focused knowledge distribution data is obtained.
[0076] Next, after obtaining the focused knowledge distribution data, predictions are made for different reference disaster recovery backup allocation space data. For example, for the construction material cost part, it is assumed that there are three cases for the reference disaster recovery backup allocation space data: 10%, 15% and 20% of the total cost as disaster recovery backup space. The disaster recovery backup model analyzes and calculates the construction material cost-related factors in the focused knowledge distribution data (such as material price fluctuations, supply stability, etc.), and obtains a confidence level of 0.6 in the case of 10% disaster recovery backup allocation space, a confidence level of 0.8 in the case of 15% disaster recovery backup allocation space, and a confidence level of 0.9 in the case of 20% disaster recovery backup allocation space. Similar predictions are made for other parts of the construction project cost data set (such as labor costs, equipment rental costs, etc.) for these three reference disaster recovery backup allocation space data to obtain their respective confidence levels.
[0077] Finally, taking the entire construction project cost data set as a whole, the confidence of each part such as building material cost, labor cost, and equipment rental cost under different reference disaster recovery backup allocation space data is comprehensively considered. For example, the confidence of building material cost under 15% disaster recovery backup allocation space is higher, the confidence of labor cost under 10% disaster recovery backup allocation space is higher, and the confidence of equipment rental cost under 20% disaster recovery backup allocation space is higher. However, since the overall disaster recovery backup effect of the entire construction project cost data set needs to be considered, it is necessary to weigh the confidence of different reference disaster recovery backup allocation space data according to factors such as the proportion of each part in the total cost and their mutual relationship. After calculation and weighing, the predicted disaster recovery backup allocation space data of the entire construction project cost data set is determined. This predicted disaster recovery backup allocation space data is not a simple value, but a logical structure data that comprehensively considers the characteristics of each cost part, the mutual relationship, and the confidence under different disaster recovery backup allocation spaces, which can provide an accurate basis for the disaster recovery backup processing of the construction project cost data set.
[0078] In a possible implementation, step S151 includes:
[0079] Step S1511, performing a data cleaning operation on the construction project cost data set to remove noise data, erroneous data and incomplete data in the construction project cost data set.
[0080] Step S1512, after the construction project cost data set that has undergone the data cleaning operation is standardized, data encoding is performed on the standardized construction project cost data set to generate a data encoding result of the construction project cost data set.
[0081] Step S1513, input the data encoding result into the initial feature extraction layer of the disaster recovery backup model, perform a linear combination operation on the data encoding result through multiple neurons in the initial feature extraction layer to obtain an intermediate feature extraction result, and perform activation function processing on the intermediate feature extraction result to generate an initial feature set.
[0082] Step S1514, using a feature crossover method, combining the features in the initial feature set to generate a combined feature sequence, and filtering the combined feature sequence based on a feature importance evaluation index to obtain a filtered combined feature sequence, wherein the importance evaluation index is obtained by calculating the correlation between each feature and the disaster recovery backup target.
[0083] Step S1515, input the screened combined feature sequence into the deep feature mining layer in the disaster recovery backup model, the deep feature mining layer is composed of multiple hidden layers, the screened combined feature sequence is input into the first layer of the deep feature mining layer, and the screened combined feature sequence is linearly combined through the neurons of the first layer, and the linear combination operation result is activated. After processing, the output result is passed to the next hidden layer to repeat the above-mentioned linear combination and activation function processing process, and the output result of the last hidden layer is used as the deep feature set.
[0084] Step S1516, performing semantic analysis on the deep features in the deep feature set, and grouping the deep features with similar semantics according to the semantic analysis results, performing knowledge fusion operation on the grouped deep features to generate feature knowledge subsets, and combining all feature knowledge subsets to generate an integrated feature knowledge set.
[0085] Step S1517, determine the dimension of feature knowledge distribution, perform statistics on the integrated feature knowledge set according to the determined dimension, and construct the structure of feature knowledge distribution data using a matrix data structure or a vector data structure based on the statistical results. If a matrix structure is used, the rows represent different dimensions, and the columns can represent the specific attributes of the feature knowledge under each dimension.
[0086] Step S1518, determining the feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension.
[0087] Step S1519: determining the focused knowledge distribution data according to the feature knowledge distribution data and the feature focusing factor of the disaster recovery knowledge dimension.
[0088] In a possible implementation, step S1518 includes:
[0089] Step S1518-1, analyzing the knowledge content related to the disaster recovery knowledge dimension classification in the characteristic knowledge distribution data.
[0090] Step S1518-2, with each feature in the knowledge content classified under the disaster recovery knowledge dimension as a node, construct a feature association network, wherein the direct association relationship and indirect association relationship between each feature are specifically analyzed, and according to the direct and indirect association relationship between the features, the relevant feature nodes are connected with connecting edges to generate the feature association network, which is used to reflect the mutual relationship between the various features in the knowledge content classified under the disaster recovery knowledge dimension.
[0091] Step S1518-3, for each node in the feature association network, evaluate the importance of the node in the feature association network based on the degree value, betweenness centrality and closeness centrality of the node, and generate a first feature importance evaluation result, wherein the degree value is the number of edges connected to the node, the betweenness centrality represents the frequency of a node appearing on the shortest path between other nodes, and the closeness centrality reflects the average distance from a node to other nodes.
[0092] Step S1518-4, obtaining a preset disaster recovery strategy, and adjusting the first feature importance evaluation result according to the disaster recovery strategy to generate an adjusted feature importance evaluation result.
[0093] Step S1518-5, determine a benchmark value, which is the average value of all feature importances or a basic value set according to a preset rule.
[0094] Step S1518-6, for each node, a calculation rule is constructed according to the relationship between the adjusted importance assessment result of the node and the benchmark value, and different dimension weights are set for the backup frequency knowledge dimension, the backup storage location knowledge dimension, and the data recovery strategy knowledge dimension, respectively, and a calculation framework is constructed according to the calculation rules and dimension weights.
[0095] Step S1518-7, calling the calculation framework to calculate the feature focusing factor of each node in the disaster recovery knowledge dimension based on the adjusted feature importance evaluation result and the dimension weight of the disaster recovery knowledge dimension, and summarizing the feature focusing factor of each node in the disaster recovery knowledge dimension to determine the feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension.
[0096] In this embodiment, a large-scale construction project cost data set is taken as an example. This data set contains cost information of many construction projects, such as the purchase cost of various types of building materials, labor costs of different types of work, equipment rental and purchase costs, and various management costs. In this data set, noise data may be manifested as abnormal fluctuations of individual data points due to accidental errors in data entry, such as the price of a certain type of building material being mistakenly entered as a value that is significantly deviated from the normal market price range; erroneous data may be due to incorrect classification or calculation of certain costs, such as a cost that should belong to equipment purchase costs being mistakenly included in the construction material purchase cost; incomplete data may be that some projects only record part of the cost information and lack key cost data, such as a construction project that only records the cost of the infrastructure part, but does not record the cost of the decoration part. Through data cleaning operations, specific algorithms and rules are used to identify and correct these data problems. For example, for data points that obviously deviate from the normal price range, they are corrected to reasonable values by comparing them with the market average price or price data of similar items. For misclassified data, they are reclassified according to the nature and purpose of the expenses. For incomplete data, they are supplemented if possible, otherwise the data of the item is specially marked or appropriately processed in subsequent analysis.
[0097] After the construction project cost data set that has undergone data cleaning operations is standardized, the standardized construction project cost data set is encoded to generate the data encoding results of the construction project cost data set. Standardization is intended to unify data of different types and units into a standard scale for subsequent calculation and analysis. For example, the purchase cost of building materials may be per ton or per cubic meter, while the labor cost may be per person per day. Through standardization, these different units of data are converted into a unified numerical representation. When encoding the data, according to the pre-set encoding rules, the standardized construction project cost data set is converted into an encoding form that can be recognized and processed by the computer. For example, encoding methods such as One-Hot Encoding can be used to convert different categorical variables (such as the type of building materials, the type of work, etc.) into binary vector forms, and corresponding encoding conversions are also performed for numerical variables (such as the amount of fees, etc.), and finally the data encoding results of the construction project cost data set are generated.
[0098] The data encoding result is input into the initial feature extraction layer of the disaster recovery backup model, and the data encoding result is linearly combined through multiple neurons in the initial feature extraction layer to obtain an intermediate feature extraction result, and the intermediate feature extraction result is processed by activation function to generate an initial feature set. Assume that the initial feature extraction layer of the disaster recovery backup model has several neurons, and each neuron has different weights for different parts of the input data encoding result. Take the building material cost data as an example. In the data encoding result, the cost data of different building materials correspond to different encoding values. These encoding values are input into the neurons of the initial feature extraction layer, and the neurons perform linear combination operations on these input values according to their weights. For example, for the encoding value corresponding to the cost of a certain building material, the weight of neuron A is 0.3, and the weight of neuron B is 0.2. Then the linear combination result is the value obtained by multiplying the encoding value by the corresponding weight and adding them together. After performing such a linear combination operation on all the encoding values, the intermediate feature extraction result is obtained. Then, the intermediate feature extraction result is processed by activation function. The activation function can use ReLU (Rectified Linear Unit) function, etc. The intermediate feature extraction result is converted into the initial feature set through the nonlinear transformation of the activation function. This initial feature set contains the feature information of the construction project cost data set after preliminary processing. These feature information have been extracted and converted to a certain extent, providing a basis for subsequent operations.
[0099] The feature crossover method is used to combine the features in the initial feature set to generate a combined feature sequence, and the combined feature sequence is screened based on the feature importance evaluation index to obtain the screened combined feature sequence, where the importance evaluation index is obtained by calculating the correlation between each feature and the disaster recovery backup target. The initial feature set includes various features related to construction costs, such as different construction material costs, labor costs of different types of work, etc. The feature crossover method is used to combine these features, for example, cross-combining the construction material cost feature with the labor cost feature to generate new combined features, such as "the correlation feature between a certain construction material cost and a specific type of labor cost", etc., so as to obtain a combined feature sequence. Then, the correlation between each feature and the disaster recovery backup target is calculated to obtain the feature importance evaluation index. Taking the disaster recovery backup target of ensuring the recoverability and integrity of construction project cost data in the event of a disaster as an example, the features with a high correlation with this target may be those related to data that have a greater impact on the project cost and are easily lost or damaged in the event of a disaster, such as the supply stability of key construction materials and the impact of cost fluctuations on the project cost and the recoverability of these data in the event of a disaster. According to the correlation calculation results, the combined feature sequence is screened, and those combined features with low correlation with the disaster recovery backup target are removed to obtain the screened combined feature sequence.
[0100] The filtered combined feature sequence is input into the deep feature mining layer in the disaster recovery backup model. The deep feature mining layer consists of multiple hidden layers. The filtered combined feature sequence is input into the first layer of the deep feature mining layer, and the neurons in the first layer perform a linear combination operation on the filtered combined feature sequence. After the linear combination operation result is processed by the activation function, the output result is passed to the next hidden layer to repeat the above linear combination and activation function processing process, and the output result of the last hidden layer is used as the deep feature set. Assume that the deep feature mining layer has three hidden layers, and the filtered combined feature sequence is input into the first hidden layer, which has several neurons. For example, for a certain combined feature in the combined feature sequence, the neurons in the first hidden layer perform a linear combination operation on the combined feature according to its weight, and then process it through the activation function (such as the tanh function) to obtain the output result and pass it to the second hidden layer. The neurons in the second hidden layer also perform a linear combination and activation function processing on the input result, and then pass the result to the third hidden layer. After the same operation, the third hidden layer uses its output result as the deep feature set. This deep feature set contains the feature information of the construction project cost data set after deep mining and conversion. This feature information is more abstract and advanced, and can better reflect the inherent structure of the data and information related to disaster recovery and backup.
[0101] The deep features in the deep feature set are semantically analyzed, and the deep features with similar semantics are grouped according to the semantic analysis results. Then, the grouped deep features are fused with knowledge to generate feature knowledge subsets. All feature knowledge subsets are combined to generate an integrated feature knowledge set. In the deep feature set, different deep features have different semantic meanings. For example, some deep features may be semantically related to the cost fluctuation of building materials in different time periods, and other deep features may be semantically related to the labor costs of different types of work in different construction stages. Through semantic analysis, deep features with similar semantics are identified, such as deep features related to the cost fluctuation of building materials, which are grouped together. Then, the grouped deep features are fused with knowledge to integrate the deep features in the same group, such as by weighted average or logical operation, to fuse them into a feature knowledge subset. After performing such operations on all groups, all feature knowledge subsets obtained are combined to generate an integrated feature knowledge set. This integrated feature knowledge set integrates various feature information in the deep feature set, and is more refined and meaningful through semantic analysis and knowledge fusion.
[0102] Determine the dimension of characteristic knowledge distribution, count the integrated characteristic knowledge set according to the determined dimension, and construct the structure of characteristic knowledge distribution data using matrix data structure or vector data structure according to the statistical results, wherein if a matrix structure is used, rows represent different dimensions, and columns can represent specific attributes of characteristic knowledge under each dimension. For example, the determined dimensions may include building material cost dimension, labor cost dimension, equipment cost dimension, etc. For the integrated characteristic knowledge set, statistics are performed according to these dimensions, and the quantity, proportion, and other information of characteristic knowledge under each dimension are counted. If a matrix data structure is used to construct the structure of characteristic knowledge distribution data, the building material cost dimension, labor cost dimension, and equipment cost dimension are used as rows, and for each row, the columns can represent different attributes, such as the fluctuation range of cost, the composition ratio of cost, etc. Through such statistics and structure construction, characteristic knowledge distribution data is generated, which can clearly show the distribution of characteristic knowledge of the construction project cost data set under different dimensions.
[0103] In determining the feature focusing factor of the feature knowledge distribution data and extracting the focused knowledge distribution data of the feature knowledge distribution data based on the feature focusing factor, first determine the feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension. Analyze the knowledge content related to the disaster recovery knowledge dimension classification in the feature knowledge distribution data. For example, in the construction project cost data, the content related to the disaster recovery knowledge dimension may include knowledge content in terms of data backup frequency, backup storage location, data recovery strategy, etc. Take each feature in the knowledge content classified under the disaster recovery knowledge dimension as a node to build a feature association network. For example, for features such as the backup frequency of building material cost and the backup frequency of labor cost under the backup frequency knowledge dimension, analyze the direct association relationship between them (such as the backup frequency of building material cost and the backup frequency of labor cost may be directly associated due to the overall budget control of the project) and the indirect association relationship (such as the indirect association through the risk assessment and management strategy of the project), and connect the relevant feature nodes with connecting edges based on these relationships to generate a feature association network. This feature association network can reflect the mutual relationship between each feature in the knowledge content classified under the disaster recovery knowledge dimension.
[0104] For each node in the feature association network, the importance of the node in the feature association network is evaluated based on the node's degree value, betweenness centrality, and proximity centrality, and the first feature importance evaluation result is generated. For example, for a node related to the frequency of backup of building material costs, its degree value represents the number of edges connected to this node. If there are 3 edges connected to it, the degree value is 3; betweenness centrality represents the frequency of this node appearing on the shortest path between other nodes. If this node appears on the shortest path between many other nodes, then the betweenness centrality is high; proximity centrality reflects the average distance from this node to other nodes. If the average distance from this node to other nodes is short, then the proximity centrality is high. Based on these indicators, the importance of the node in the feature association network is evaluated to obtain the first feature importance evaluation result. Obtain a pre-set disaster recovery strategy, such as using high-frequency backup and safe storage location strategies for important cost data. According to this disaster recovery strategy, adjust the first feature importance evaluation result to generate an adjusted feature importance evaluation result.
[0105] Determine a benchmark value, which can be the average value of all feature importances or a basic value set according to a preset rule. For example, if the average value of all feature importances is 0.5, then this 0.5 is used as the benchmark value. For each node, a calculation rule is constructed based on the relationship between the adjusted importance evaluation result of the node and the benchmark value, and different dimension weights are set for the backup frequency knowledge dimension, the backup storage location knowledge dimension, and the data recovery strategy knowledge dimension, such as the backup frequency knowledge dimension weight is 0.4, the backup storage location knowledge dimension weight is 0.3, and the data recovery strategy knowledge dimension weight is 0.3. According to the calculation rules and dimension weights, a calculation framework is constructed. Call this calculation framework based on the adjusted feature importance evaluation results and the dimension weight of the disaster recovery knowledge dimension to calculate the feature focus factor of each node under the disaster recovery knowledge dimension, and summarize the feature focus factors of each node under the disaster recovery knowledge dimension to determine the feature focus factor of the feature knowledge distribution data in the disaster recovery knowledge dimension.
[0106] The focused knowledge distribution data is determined based on the feature knowledge distribution data and the feature focusing factor of the disaster recovery knowledge dimension. First, the data structure of the feature knowledge distribution data is parsed. Based on the data structure of the feature knowledge distribution data, the feature elements associated with the disaster recovery knowledge dimension are searched. For example, the elements in the rows and columns related to the backup frequency, backup storage location, and data recovery strategy are searched in the matrix structure data. All feature elements related to the disaster recovery knowledge dimension are extracted by searching and matching in the parsed data structure to generate a feature element set. The feature focusing factor of the disaster recovery knowledge dimension is used to assign weights to each feature element in the feature element set to generate a weight sequence of the feature element set. The feature focusing factor is used to measure the index of the importance of each feature element under the disaster recovery knowledge dimension. According to the pre-set weight threshold or weight ratio, according to the weight sequence of the feature element set, the feature element set is screened to generate an optimized feature element subset. The data structure is reconstructed based on the screened feature element subset to generate the focused knowledge distribution data. This focused knowledge distribution data focuses more on the content related to the disaster recovery knowledge dimension, and can provide a more targeted data basis for subsequent disaster recovery backup prediction and other operations.
[0107] In a possible implementation, step S1519 includes:
[0108] Step S1519-1, parse the data structure of the characteristic knowledge distribution data, find the characteristic elements associated with the disaster recovery knowledge dimension based on the data structure of the characteristic knowledge distribution data, extract all the characteristic elements related to the disaster recovery knowledge dimension by searching and matching in the parsed data structure, and generate a characteristic element set. The disaster recovery knowledge dimension includes the backup frequency knowledge dimension, the backup storage location knowledge dimension, and the data recovery strategy knowledge dimension.
[0109] Step S1519-2, using the feature focusing factor of the disaster recovery knowledge dimension to assign weight to each feature element in the feature element set, to generate a weight sequence of the feature element set, wherein the feature focusing factor is an indicator for measuring the importance of each feature element in the disaster recovery knowledge dimension.
[0110] Step S1519-3, based on a preset weight threshold or weight ratio and according to the weight sequence of the feature element set, the feature element set is screened to generate an optimized feature element subset.
[0111] Step S1519-4, reconstructing the data structure based on the filtered feature element subset to generate the focused knowledge distribution data.
[0112] In this embodiment, the characteristic knowledge distribution data of a construction project cost data set is taken as an example, assuming that the data structure is represented in matrix form, the rows represent different cost components, such as building materials, labor costs, equipment rental, etc., and the columns represent various attributes under each cost component, such as cost fluctuation range, seasonal impact, etc. For the disaster recovery knowledge dimension, the backup frequency knowledge dimension may be related to the update frequency of cost data in different time periods, the backup storage location knowledge dimension may be related to the server location or storage medium where the cost data is stored, and the data recovery strategy knowledge dimension may be related to the recovery process and cost in the event of data loss or damage.
[0113] When analyzing the characteristic knowledge distribution data of this matrix structure, from the perspective of rows, for the row of building materials, it may be found that some attributes are related to the backup frequency knowledge dimension. For example, the supply stability of a special building material is low, which leads to the need for high-frequency backup of its cost data. Then this attribute related to the backup frequency is the characteristic element associated with the disaster recovery knowledge dimension. From the perspective of columns, in the column of cost fluctuation range, it may be found that the cost fluctuation of labor costs is related to the data recovery strategy knowledge dimension, because the fluctuation of labor costs may affect the budget adjustment during data recovery, which is also a characteristic element related to the disaster recovery knowledge dimension. Through careful search and matching of the entire matrix structure, all characteristic elements related to the backup frequency knowledge dimension, the backup storage location knowledge dimension, and the data recovery strategy knowledge dimension are extracted to form a set of characteristic elements.
[0114] Next, the feature focus factor of the disaster recovery knowledge dimension is used to assign weights to each feature element in the feature element set, and a weight sequence of the feature element set is generated. The feature focus factor is used to measure the index of the importance of each feature element in the disaster recovery knowledge dimension. Continuing with the above-mentioned construction project cost data set as an example, it is assumed that the feature focus factors of each feature element in the disaster recovery knowledge dimension have been obtained in the previous calculation. For a feature element in the feature element set that is related to the backup frequency of building materials, its feature focus factor may be high, which means that the feature element is more important in the disaster recovery knowledge dimension. According to this feature focus factor, a relatively high weight is assigned to this feature element, such as 0.8. For another feature element related to the labor cost data recovery strategy, if its feature focus factor is low, it may be assigned a weight of 0.3. In this way, each feature element in the feature element set is weighted according to its corresponding feature focus factor, thereby generating a weight sequence of the feature element set.
[0115] Then, according to the preset weight threshold or weight ratio, according to the weight sequence of the feature element set, the feature element set is screened to generate an optimized feature element subset. The preset weight threshold or weight ratio is determined according to the actual disaster recovery requirements and experience. For example, the weight threshold is set to 0.5. For each feature element in the feature element set, if its weight is greater than or equal to 0.5, it will be retained in the optimized feature element subset. Feature elements with a weight less than 0.5 will be discarded. Or, using the weight ratio method, assuming that the first 60% of feature elements with higher weights are selected, then after sorting the feature element set according to the weight, the first 60% of feature elements are selected to form the optimized feature element subset. Taking the feature element related to the backup frequency of building materials as an example, if its weight is 0.8, which is greater than the weight threshold of 0.5, it will be retained in the optimized feature element subset; and if a feature element related to the equipment rental storage location has a weight of 0.4, which is less than the weight threshold of 0.5, it will be discarded.
[0116] Finally, the data structure is reconstructed based on the filtered feature element subset to generate focused knowledge distribution data. Since the filtered feature element subset only contains feature elements that are highly relevant to the disaster recovery knowledge dimension and have high importance, reconstructing the data structure based on these feature elements can more accurately reflect the knowledge distribution related to disaster recovery. If the previous feature knowledge distribution data adopts a matrix structure, then when reconstructing the focused knowledge distribution data, only the rows and columns related to the filtered feature element subset may be retained. For example, in the previous matrix, if the feature elements related to the backup frequency of building materials are retained, then in the new focused knowledge distribution data structure, the columns related to the backup frequency in the row of building materials, as well as other rows and columns related to the retained feature elements, will be retained. Through such reconstruction, the generated focused knowledge distribution data is more focused on the content related to the disaster recovery knowledge dimension, which can provide a more targeted data basis for subsequent operations such as disaster recovery backup prediction, and help improve the accuracy and efficiency of disaster recovery backup processing of construction project cost data.
[0117] Figure 2 The hardware structure of the disaster recovery management system 100 for implementing the above-mentioned construction project cost data disaster recovery method provided by the embodiment of the present invention is shown as follows: Figure 2 As shown, the disaster recovery management system 100 may include a processor 110 , a machine-readable storage medium 120 , a bus 130 , and a communication unit 140 .
[0118] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data obtained from an external terminal. In some embodiments, the machine-readable storage medium 120 may store data and / or instructions that the disaster recovery management system 100 uses to execute or use to complete the exemplary method described in the present invention.
[0119] During the specific implementation process, one or more processors 110 execute computer executable instructions stored in the machine-readable storage medium 120, so that the processor 110 can execute the construction project cost data disaster recovery method of the above method embodiment. The processor 110, the machine-readable storage medium 120 and the communication unit 140 are connected through the bus 130, and the processor 110 can be used to control the sending and receiving actions of the communication unit 140.
[0120] The specific implementation process of the processor 110 can refer to the various method embodiments executed by the above-mentioned disaster recovery management system 100. The implementation principles and technical effects are similar, and this embodiment will not be repeated here.
[0121] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer executable instructions are preset. When a processor executes the computer executable instructions, the above-mentioned construction project cost data disaster recovery method is implemented.
[0122] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, various features are sometimes combined into one embodiment, drawing or description thereof.
Claims
1. A construction project cost data disaster recovery method, characterized in that: The method comprises: Acquire a sample construction project cost data sequence and initialize a neural network; the sample construction project cost data sequence includes a plurality of candidate sample construction project cost data; each candidate sample construction project cost data carries disaster recovery backup allocation space data annotation data; Perform disaster recovery backup prediction on each candidate sample construction project cost data according to the initialized neural network, and generate training disaster recovery backup allocation space data for each candidate sample construction project cost data; For each of the candidate sample construction project cost data, determine the global network learning error according to the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data, and extract valid sample construction project cost data from the sample construction project cost data sequence according to the global network learning error of each of the candidate sample construction project cost data; Performing network parameter learning on the initialized neural network based on the extracted valid sample construction project cost data, and outputting a trained disaster recovery backup model; A construction project cost data set to be processed for data disaster recovery is obtained, and disaster recovery backup processing is performed on the construction project cost data set according to the trained disaster recovery backup model and the predicted disaster recovery backup allocation space data.
2. The construction project cost data disaster recovery method according to claim 1, characterized in that: The initialized neural network includes two or more; The method of performing disaster recovery backup prediction on each candidate sample construction project cost data according to the initialized neural network to generate training disaster recovery backup allocation space data for each candidate sample construction project cost data includes: Loading each candidate sample construction project cost data into a plurality of initialized neural networks respectively, generating training disaster recovery backup allocation space data corresponding to each initialized neural network for each candidate sample construction project cost data; For each candidate sample construction project cost data, determining the global network learning error based on the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data includes: Determining a first network learning error corresponding to each of the initialized neural networks according to the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data corresponding to each of the initialized neural networks; The global network learning error of the candidate sample construction project cost data is determined based on the first network learning error corresponding to each of the initialized neural networks.
3. The construction project cost data disaster recovery method according to claim 2, characterized in that: The method of learning network parameters of the initialized neural network based on the extracted valid sample construction project cost data and outputting the trained disaster recovery backup model includes: Performing network parameter learning on the multiple initialized neural networks respectively according to the valid sample construction project cost data to generate multiple initialized neural networks; By performing performance verification on the multiple initialized neural networks, the trained disaster recovery backup model is determined from the multiple initialized neural networks.
4. The construction project cost data disaster recovery method according to claim 2, characterized in that: Determining the global network learning error of the candidate sample construction project cost data according to the first network learning error corresponding to each of the initialized neural networks includes: Determine the relative learning error between each two of the training disaster recovery backup allocation space data; The sum of each of the first network learning errors and the relative learning error is determined as the global network learning error.
5. The construction project cost data disaster recovery method according to claim 1, characterized in that: The method of learning network parameters of the initialized neural network based on the extracted valid sample construction project cost data and outputting the trained disaster recovery backup model includes: Iteratively learning network parameters of the initialized neural network according to the extracted valid sample construction project cost data to generate a first temporary neural network, and if the first temporary neural network meets the network convergence requirements, generating the trained disaster recovery backup model according to the first temporary neural network; If the first temporary neural network does not meet the network convergence requirement, the method further includes: The first temporary neural network after learning the network parameters is used as an iterative neural network, and the following steps are iteratively performed until the generated second temporary neural network meets the network convergence requirements, and the trained disaster recovery backup model is generated according to the second temporary neural network that meets the network convergence requirements: Obtain iterative sample construction project cost data sequence; Performing disaster recovery backup prediction on each candidate sample construction project cost data in the iterative sample construction project cost data sequence according to the initialized neural network, and generating training disaster recovery backup allocation space data for each candidate sample construction project cost data; For each candidate sample construction project cost data, determining a global network learning error based on the training disaster recovery backup allocation space data and the disaster recovery backup allocation space data labeling data; Extracting valid sample construction project cost data from the sample construction project cost data sequence according to the global network learning error of each candidate sample construction project cost data; Iterative network parameter learning is performed on the initialized neural network according to the extracted valid sample construction project cost data to generate a second temporary neural network. If the second temporary neural network does not meet the network convergence requirements, the second temporary neural network is used as an iterative neural network.
6. The construction project cost data disaster recovery method according to claim 2, characterized in that: The extracting of valid sample construction project cost data from the sample construction project cost data sequence based on the global network learning error of each candidate sample construction project cost data includes at least one of the following: Marking the sample construction project cost data with a set proportion and the smallest global network learning error as the valid sample construction project cost data; The sample construction project cost data whose global network learning error is less than the set error is marked as the valid sample construction project cost data.
7. The construction project cost data disaster recovery method according to claim 1, characterized in that: The performing disaster recovery backup processing on the construction project cost data set according to the trained disaster recovery backup model and the predicted disaster recovery backup allocation space data includes: Utilizing the disaster recovery backup model, extracting characteristic knowledge distribution data of the construction project cost data set, determining a characteristic focusing factor of the characteristic knowledge distribution data, and extracting focused knowledge distribution data of the characteristic knowledge distribution data according to the characteristic focusing factor; Performing disaster recovery backup prediction on the focused knowledge distribution data to generate confidence levels of the construction project cost data set corresponding to a plurality of reference disaster recovery backup allocation space data; The predicted disaster recovery backup allocation space data of the construction project cost data set is determined according to the confidence levels of the plurality of reference disaster recovery backup allocation space data corresponding to the construction project cost data set.
8. The construction project cost data disaster recovery method according to claim 7, characterized in that: The step of extracting characteristic knowledge distribution data of the construction project cost data set by using the disaster recovery backup model includes: Performing a data cleaning operation on the construction project cost data set to remove noise data, erroneous data and incomplete data in the construction project cost data set; After the construction project cost data set that has undergone the data cleaning operation is standardized, data encoding is performed on the standardized construction project cost data set to generate a data encoding result of the construction project cost data set; Inputting the data encoding result into the initial feature extraction layer of the disaster recovery backup model, performing a linear combination operation on the data encoding result through multiple neurons in the initial feature extraction layer to obtain an intermediate feature extraction result, and performing activation function processing on the intermediate feature extraction result to generate an initial feature set; Using a feature crossover method, combining the features in the initial feature set to generate a combined feature sequence, and screening the combined feature sequence based on a feature importance evaluation index to obtain a screened combined feature sequence, wherein the importance evaluation index is obtained by calculating the correlation between each feature and a disaster recovery backup target; The screened combined feature sequence is input into the deep feature mining layer in the disaster recovery backup model, the deep feature mining layer is composed of multiple hidden layers, the screened combined feature sequence is input into the first layer of the deep feature mining layer, and the screened combined feature sequence is linearly combined through the neurons of the first layer, and the linear combination operation result is activated, and the output result is passed to the next hidden layer to repeat the above linear combination and activation function processing process, and the output result of the last hidden layer is used as the deep feature set; Performing semantic analysis on the deep features in the deep feature set, grouping the deep features with similar semantics according to the semantic analysis results, performing knowledge fusion operation on the grouped deep features to generate feature knowledge subsets, and combining all feature knowledge subsets to generate an integrated feature knowledge set; Determine the dimension of feature knowledge distribution, perform statistics on the integrated feature knowledge set according to the determined dimension, and construct the structure of feature knowledge distribution data using a matrix data structure or a vector data structure based on the statistical results. If a matrix structure is used, rows represent different dimensions, and columns can represent specific attributes of feature knowledge under each dimension. The step of determining a feature focusing factor of the feature knowledge distribution data and extracting focused knowledge distribution data of the feature knowledge distribution data according to the feature focusing factor comprises: Determine a feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension; Determining the focused knowledge distribution data according to the characteristic knowledge distribution data and the characteristic focusing factor of the disaster recovery knowledge dimension; The step of determining the feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension includes: Analyzing the knowledge content related to the disaster recovery knowledge dimension classification in the characteristic knowledge distribution data; Taking each feature in the knowledge content classified under the disaster recovery knowledge dimension as a node, a feature association network is constructed, wherein the direct association relationship and the indirect association relationship between each feature are specifically analyzed, and according to the direct and indirect association relationship between the features, the related feature nodes are connected with connecting edges to generate the feature association network, and the feature association network is used to reflect the mutual relationship between each feature in the knowledge content classified under the disaster recovery knowledge dimension; For each node in the feature association network, based on the degree value, betweenness centrality and closeness centrality of the node, evaluating the importance of the node in the feature association network, and generating a first feature importance evaluation result, the degree value is the number of edges connected to the node, the betweenness centrality represents the frequency of a node appearing on the shortest path between other nodes, and the closeness centrality reflects the average distance from a node to other nodes; Obtaining a preset disaster recovery strategy, and adjusting the first feature importance evaluation result according to the disaster recovery strategy to generate an adjusted feature importance evaluation result; Determine a benchmark value, which is an average value of the importance of all features or a basic value set according to a preset rule; For each node, a calculation rule is constructed according to the relationship between the adjusted importance evaluation result of the node and the benchmark value, and different dimension weights are set for the backup frequency knowledge dimension, the backup storage location knowledge dimension, and the data recovery strategy knowledge dimension, and a calculation framework is constructed according to the calculation rule and the dimension weights; The calculation framework is called to calculate the feature focusing factor of each node in the disaster recovery knowledge dimension based on the adjusted feature importance evaluation result and the dimension weight of the disaster recovery knowledge dimension, and the feature focusing factor of each node in the disaster recovery knowledge dimension is summarized to determine the feature focusing factor of the feature knowledge distribution data in the disaster recovery knowledge dimension.
9. The construction project cost data disaster recovery method according to claim 8, characterized in that: The determining the focused knowledge distribution data according to the feature knowledge distribution data and the feature focusing factor of the disaster recovery knowledge dimension includes: Parsing the data structure of the characteristic knowledge distribution data, searching for characteristic elements associated with the disaster recovery knowledge dimension based on the data structure of the characteristic knowledge distribution data, extracting all characteristic elements related to the disaster recovery knowledge dimension by searching and matching in the parsed data structure, and generating a characteristic element set; the disaster recovery knowledge dimension includes a backup frequency knowledge dimension, a backup storage location knowledge dimension, and a data recovery strategy knowledge dimension; Using the feature focusing factor of the disaster recovery knowledge dimension to assign weights to each feature element in the feature element set, generating a weight sequence of the feature element set, wherein the feature focusing factor is used to measure an indicator of the importance of each feature element in the disaster recovery knowledge dimension; According to a preset weight threshold or weight ratio, and according to the weight sequence of the feature element set, the feature element set is screened to generate an optimized feature element subset; The data structure is reconstructed based on the filtered feature element subset to generate the focused knowledge distribution data.
10. A disaster recovery management system, characterized in that: The disaster recovery management system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the construction project cost data disaster recovery method described in any one of claims 1-9 above.
Citation Information
Patent Citations
HSS (home subscriber server) data disaster tolerant algorithm based on artificial intelligence and neural network
CN102724070A
Bridge data intelligent disaster recovery backup system and method
CN117149522A
System and method for predicting power plant operational parameters utilizing artificial neural network deep learning methodologies
US20170091615A1