Engineering cost data analysis method and system based on big data analysis and cloud computing

Through big data analysis and cloud computing technology, the engineering profile feature extraction algorithm and word vector technology are used to match the approximate historical project data for engineering cost analysis, solving the accuracy problems of existing methods in nonlinear and multivariable situations, and achieving efficient and accurate abnormal identification and management of engineering cost data.

CN120579772APending Publication Date: 2025-09-02ZHEJIANG HUAPU ENG MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510734264.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When faced with nonlinear relationships or multivariate complex situations, the analysis accuracy of existing engineering cost data declines, making it difficult to find potential correlations between data, resulting in large deviations from the actual situation. Traditional methods also have high requirements for data quality, making it difficult to accurately capture the nonlinear relationships and interactions in engineering cost.

Method used

Using a method based on big data analysis and cloud computing, the cloud server matches the approximate historical project data, uses the engineering profile feature extraction algorithm and word vector technology to process the engineering data, and combines the historical project database for comparison and analysis to identify abnormalities in the engineering cost data.

Benefits of technology

The abnormal analysis of current engineering cost data is realized, the analysis accuracy and efficiency is improved, the manual verification time cost is reduced, more reliable reference basis and risk warning are provided, and decision-making support capabilities for engineering cost management are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579772A_ABST
    Figure CN120579772A_ABST
Patent Text Reader

Abstract

The invention provides a project cost data analysis method and system based on big data analysis and cloud computing. A cost accounting process is optimized through big data analysis and cloud computing. The method comprises the following steps: firstly, extracting project general situation data from current project cost data, generating feature data by applying a preset algorithm, and sending the feature data to a cloud server; and then the cloud server matches an approximate historical engineering project in a historical engineering project database, determines corresponding use amount index typical feature data, and returns the use amount index typical feature data to be compared and analyzed with the current engineering use amount index data. Through the process, the abnormal condition in the current project cost data can be automatically identified, a traditional manual verification mode is effectively replaced, the time cost is remarkably reduced, and the resource configuration efficiency is improved. By combining a data processing algorithm with historical project data, a scientific basis is provided for project cost decision making, and accurate judgment of cost data abnormity is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of engineering cost analysis, and in particular to an engineering cost data analysis method and system based on big data analysis and cloud computing. Background Art

[0002] With the continuous development of the construction industry, the methods for analyzing construction cost data are becoming increasingly diverse, but each existing analysis method has its inherent limitations.

[0003] Traditional statistical analysis methods, due to their ease of use and intuitive understanding, have held a prominent position in the field of engineering cost data analysis. By calculating statistical indicators such as mean and standard deviation, these methods can quickly capture the overall characteristics of the data. However, this approach relies heavily on the quality and completeness of historical data, significantly reducing analytical accuracy when faced with nonlinear relationships or complex multivariate scenarios. Furthermore, statistical analysis methods struggle to identify potential correlations between data, easily overlooking the underlying causes of cost fluctuations and resulting in significant deviations between predicted results and actual conditions.

[0004] Parametric estimation, a common analytical method used in the early stages of a project, can quickly generate cost estimates based on limited design information, facilitating project feasibility assessments for decision makers. However, a drawback of this method is that its accuracy is limited by the rationality of parameter selection. When a project is unique or innovative, traditional parameters may not accurately reflect actual costs.

[0005] By establishing a mathematical model between construction costs and influencing factors, regression analysis can quantitatively analyze the impact of various factors on construction costs, providing a scientific basis for cost forecasting. However, this method has extremely high requirements for data quality. When the input data contains outliers or the sample size is insufficient, the model accuracy will be significantly reduced. To make matters more difficult, regression models are often based on linear assumptions, making it difficult to accurately capture the nonlinear relationships in construction costs. Especially in complex engineering projects, the interactions between cost influencing factors cannot be expressed by simple regression models, resulting in distorted analysis results. Summary of the Invention

[0006] This application provides a method for analyzing engineering cost data based on big data analysis and cloud computing, comprising the following steps: A1, obtains the current project overview data from the preset current project cost data; A2, generating corresponding current project profile feature data according to the current project profile data using a preset project profile feature extraction algorithm; A3, sending the current project profile feature data to the preset cloud server; A4, matching the current project profile feature data with a preset historical project database on the cloud server to determine corresponding approximate historical project cost data; A5, determining, at the cloud server, corresponding typical characteristic data of approximate engineering quantity indicators based on the historical engineering quantity indicator data of each approximate historical engineering project cost data; A6, receiving typical characteristic data of approximate engineering usage indicators from the cloud server; A7, compare and analyze the current engineering quantity indicator data in the current engineering cost data based on the typical characteristic data of the approximate engineering quantity indicator.

[0007] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can match the engineering cost data of similar historical projects through the cloud server, realize the comparative analysis of the current project and similar historical projects, and thus realize the abnormal analysis of the current engineering cost data. The data is processed through the engineering profile feature extraction algorithm, and the historical project database is combined to match similar projects. The typical characteristics of the usage indicators are analyzed to provide data support for cost data analysis, which can determine whether there are abnormalities in the current engineering cost data, save the time cost of manual verification, and optimize resource allocation.

[0008] Optionally, the engineering profile feature extraction algorithm includes the following steps: B1, obtain the name of each overview subject and the corresponding overview subject key value in the current project overview data; B2, generating corresponding overview subject name word vectors based on the names of each overview subject using a preset word vector algorithm; B3, generating a key-value pair of each overview subject name word vector according to the overview subject name word vector and the overview subject key value combination corresponding to each overview subject name; B4, generates the current project overview feature data based on the key-value pairs of the word vectors of each overview subject name.

[0009] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can realize the semantic expression and characterization processing of engineering data by converting the names of various profile subjects into word vectors and combining them with corresponding key values. It can improve the accuracy of matching similar engineering projects and reduce the errors in subject matching in data processing. At the same time, it provides a standardized data structure for subsequent comparative analysis, enhances the system's unified recognition ability of different expressions, and makes engineering cost analysis more accurate and efficient.

[0010] Optionally, step A4 includes the following steps: A401, sequentially obtaining historical engineering project overview feature data corresponding to the cost data of each historical engineering project from the historical engineering project database of the cloud server and combining them to generate a historical engineering overview feature data set; A402, sequentially obtain the word vector key-value pairs of each overview subject name from the current project overview feature data; A403, sequentially matching the historical project overview feature data according to the overview subject name word vector key-value pairs; A404: If the historical engineering overview feature data does not contain a historical engineering overview subject name word vector that matches the key-value pair of the overview subject name word vector, then the historical engineering overview feature data is removed from the historical engineering overview feature data set; A405 defines the remaining historical engineering overview feature data in the historical engineering overview feature data set as approximate historical engineering project cost data.

[0011] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can obtain the historical engineering profile feature data set from the cloud server database, and then use the subject name word vector key-value pairs of the current project to match one by one, intelligently eliminate historical data that does not meet the conditions, and finally retain similar historical engineering data, thereby improving the matching accuracy, reducing unnecessary data interference, and ensuring that subsequent analysis is based on the most relevant historical data, thereby providing a more reliable reference basis for engineering cost estimation and improving the accuracy of the analysis results.

[0012] Optionally, step A403 includes the following steps: A4031, sequentially obtains the word vector key-value pairs of each historical engineering overview subject name from the historical engineering overview feature data; A4032, based on the overview subject name word vector in the overview subject name word vector key-value pair and the historical engineering overview subject name word vector in each historical engineering overview subject name word vector key-value pair, calculates the corresponding cosine similarity and defines it as the subject name word vector similarity; A4033, if the subject name word vector similarity is greater than or equal to the preset subject name similarity threshold, then respectively obtain the overview subject key value corresponding to the overview subject name word vector and the historical engineering overview subject key value corresponding to the historical engineering overview subject name word vector; A4034, if the data types of the overview subject key value and the historical engineering overview subject key value are both text type, then the corresponding overview subject key value word vector and historical engineering overview subject key value word vector are generated respectively based on the overview subject key value and the historical engineering overview subject key value using the word vector algorithm; A4035 calculates the corresponding cosine similarity based on the overview subject key word vector and the historical engineering overview subject key word vector and defines it as the subject key word vector item similarity; A4036, if the subject key word vector item similarity is greater than or equal to the preset subject key similarity threshold, then the profile subject name word vector key-value pair is defined to match the historical project profile feature data.

[0013] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can perform initial screening by calculating the cosine similarity between the subject name word vectors, and then convert the text type subject key values ​​into word vectors and calculate their similarity. The double-layer screening mechanism can not only identify engineering overview elements that are expressed differently but are essentially the same or similar, but can also flexibly adjust the matching accuracy through preset thresholds, effectively avoiding the limitations of traditional character matching, improving the comparison accuracy between heterogeneous data, and providing a more accurate data basis for engineering cost analysis.

[0014] Optionally, step A5 includes the following steps: A501, generating corresponding historical engineering usage standardization index data using a preset data standardization algorithm based on each historical engineering usage index data; A502, generating a corresponding historical engineering quantity feature vector based on the standardized index data of each historical engineering quantity using a preset engineering quantity feature extraction algorithm; A503 calculates the corresponding central vector based on each historical engineering quantity characteristic vector and defines it as the typical engineering quantity characteristic vector; A504 calculates the corresponding Euclidean distance between the typical engineering quantity characteristic vector and each historical engineering quantity characteristic vector and defines it as the engineering quantity characteristic distance; A505 determines the minimum value among all engineering usage characteristic distances and defines its corresponding historical engineering usage characteristic vector as the typical characteristic data of the approximate engineering usage indicator.

[0015] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can process historical engineering usage indicators through data standardization, eliminate the influence of different units and magnitudes, and then use the feature extraction algorithm to convert complex engineering usage data into feature vectors, and calculate the central vector as a typical feature reference. The historical engineering cost data closest to the typical feature is found through Euclidean distance calculation, overcoming the limitations of traditional direct comparison, and can capture the essential characteristics of engineering usage from multiple dimensions, improve the accuracy and representativeness of feature extraction, provide more scientific data support for engineering cost analysis, and enhance the accuracy and reference value of the analysis results.

[0016] Optionally, step A7 includes the following steps: A701, generating corresponding current engineering quantity standardization indicator data through a data standardization algorithm based on the current engineering quantity indicator data; A702, generating a corresponding current engineering quantity feature vector based on the current engineering quantity standardized index data using an engineering quantity feature extraction algorithm; A703 calculates the corresponding cosine similarity based on the current engineering quantity feature vector and the typical feature data of the approximate engineering quantity index and defines it as the quantity index comparison similarity; A704: If the usage indicator comparison similarity is less than the preset comparison similarity threshold, it is determined that the current project cost data is abnormal.

[0017] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can convert complex data into feature vectors by standardizing and extracting features of current engineering quantity indicators, and then calculate the cosine similarity between the vectors corresponding to the typical feature data of the approximate engineering quantity indicators, and determine the degree of similarity through a preset threshold, thereby realizing the identification of potential anomalies in the engineering cost data, providing an objective data evaluation standard, and effectively avoiding the subjectivity and inconsistency of human judgment. Compared with traditional methods, it can better capture the essential differences between data, improve the accuracy and reliability of anomaly detection, provide timely risk warnings for engineering cost management, and help decision makers quickly discover and correct potential problems.

[0018] Optionally, the engineering cost data analysis method based on big data analysis and cloud computing further includes the following steps: A8: If the current project cost data is abnormal, multiple corresponding current project first-level subject usage index standardization data are generated based on the current project usage standardization index data and the preset first-level subject list; A9, receiving historical engineering quantity standardization indicator data corresponding to the typical characteristic data of the approximate engineering quantity indicator from the cloud server, and generating a plurality of corresponding typical engineering first-level subject quantity indicator standardization data according to the historical engineering quantity standardization indicator data and the first-level subject list; A10 generates a corresponding current project first-level subject usage feature vector based on the standardized data of the current project first-level subject usage indicators using a project usage feature extraction algorithm; A11, based on the standardized data of the consumption indicators of the first-level subjects of each typical project, generates the corresponding typical project first-level subject consumption feature vector through the engineering consumption feature extraction algorithm; A12, based on the first-level subject usage feature vector of each current project and the first-level subject usage feature vector of the corresponding typical project, calculates the corresponding cosine similarity and defines it as the first-level subject feature similarity; A13: If the first-level subject feature similarity is less than the preset subject similarity threshold, the corresponding subject name is determined based on the first-level subject list and defined as an abnormal usage subject.

[0019] By adopting the above technical solution, the engineering cost data analysis method based on big data analysis and cloud computing can accurately identify which specific first-level subjects have cost anomalies by dividing the engineering data into first-level subjects, extracting feature vectors respectively and performing cosine similarity calculation with the corresponding subjects of typical projects, thereby improving the meticulousness and pertinence of anomaly diagnosis, making the anomaly location clearer, providing a clear direction for subsequent rectification, and improving the pertinence and decision-making support capabilities of cost anomaly analysis.

[0020] This application also provides an engineering cost data analysis system based on big data analysis and cloud computing, including: Cloud computing module; Local computing module; The cloud computing module and the local computing module are communicatively connected; The cloud computing module includes a cloud storage module and a cloud processing module, and the cloud storage module and the cloud processing module are data-connected. The engineering cost data analysis system based on big data analysis and cloud computing further includes an engineering cost analysis strategy, including the following steps: C1, obtaining current project overview data from preset current project cost data through the local calculation module; C2, generating corresponding current engineering profile feature data by the local computing module according to the current engineering profile data using a preset engineering profile feature extraction algorithm; C3, sending the current project profile feature data to the cloud computing module through the local computing module; C4, the cloud computing module matches the current project profile feature data with a preset historical project database to determine corresponding approximate historical project cost data; C5, determining, in the cloud computing module, corresponding typical characteristic data of approximate engineering quantity indicators based on the historical engineering quantity indicator data of each approximate historical engineering project cost data; C6, receiving typical characteristic data of approximate engineering usage indicators from the cloud computing module to the local computing module; C7, comparing and analyzing the current engineering quantity indicator data in the current engineering cost data according to the typical characteristic data of the approximate engineering quantity indicator by the local calculation module.

[0021] By adopting the above technical solution, the engineering cost data analysis system based on big data analysis and cloud computing can match the engineering cost data of similar historical projects through the cloud server, realize comparative analysis of the current project with similar historical projects, and thus realize abnormal analysis of the current engineering cost data. It processes data through the engineering profile feature extraction algorithm, matches similar projects in combination with the historical project database, and analyzes the typical characteristics of usage indicators to provide data support for cost data analysis. It can determine whether there are abnormalities in the current engineering cost data, save the time cost of manual verification, and optimize resource allocation.

[0022] In summary, this application includes at least one of the following beneficial technical effects: 1. The cloud server can be used to match the engineering cost data of similar historical projects to achieve comparative analysis of the current project with similar historical projects, thereby realizing abnormal analysis of the current engineering cost data. Data processing is performed through the engineering profile feature extraction algorithm, combined with the historical project database to match similar projects, and analyze the typical characteristics of usage indicators to provide data support for cost data analysis. It can determine whether there are abnormalities in the current engineering cost data, save the time cost of manual verification, and optimize resource allocation.

[0023] 2. By converting the names of each overview subject into word vectors and combining them with corresponding key values, the semantic expression and feature processing of engineering data can be achieved, which can improve the accuracy of matching similar engineering projects and reduce the error of subject matching in data processing. At the same time, it provides a standardized data structure for subsequent comparative analysis, enhances the system's unified recognition ability of different expressions, and makes engineering cost analysis more accurate and efficient.

[0024] 3. By obtaining a historical engineering overview feature dataset from the cloud server database, and then using the subject name word vector key-value pairs of the current project to match them one by one, the historical data that does not meet the conditions can be intelligently eliminated, and similar historical engineering data can be retained. This improves the matching accuracy, reduces unnecessary data interference, and ensures that subsequent analysis is based on the most relevant historical data, thereby providing a more reliable reference basis for engineering cost estimation and improving the accuracy of the analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a process diagram of an engineering cost data analysis method based on big data analysis and cloud computing of the present invention.

[0026] Figure 2 It is a schematic diagram of the principle of an engineering cost data analysis system based on big data analysis and cloud computing of the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0028] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.

[0029] refer to Figure 1 The present invention provides a method for analyzing engineering cost data based on big data analysis and cloud computing, which is used to analyze whether there are abnormalities in the current project cost data by combining the cost data of historical projects, including the following steps: A1, obtains the current project overview data from the preset current project cost data; Current project cost data refers to the cost data of the project that needs to be analyzed, which is compiled by the staff in a certain established form or format, usually in the form of a table; The current project overview data is the overview data and information about the current project in the current project cost data, such as building type, number of buildings, number of floors, building height, building area and other information data.

[0030] A2, generating corresponding current project profile feature data according to the current project profile data using a preset project profile feature extraction algorithm; The engineering profile feature extraction algorithm is a pre-set feature extraction algorithm used to extract features from the current engineering profile data; The current project overview characteristic data is characteristic data corresponding to the current project overview data.

[0031] A3, sending the current project profile feature data to the preset cloud server; By sending the current project profile feature data to the cloud server, unnecessary leakage of the current project information can be avoided, especially before the actual construction of the project is carried out.

[0032] A4, matching the current project profile feature data with a preset historical project database on the cloud server to determine corresponding approximate historical project cost data; The historical engineering project database is a pre-set database that stores a large amount of engineering cost data of historical engineering projects; Approximate historical engineering project cost data is historical engineering project cost data that is identified as having the same or similar engineering type as the current engineering cost data through matching with current engineering profile feature data. The matching process is performed on the cloud server, which can fully utilize the computing power of the cloud server.

[0033] A5, determining, at the cloud server, corresponding typical characteristic data of approximate engineering quantity indicators based on the historical engineering quantity indicator data of each approximate historical engineering project cost data; The typical characteristic data of approximate engineering quantity indicators are characteristic data of representative similar projects, which are obtained by comprehensive calculation or feature extraction based on the cost data of various approximate historical engineering projects.

[0034] A6, receiving typical characteristic data of approximate engineering usage indicators from the cloud server; By only receiving the feature data from the cloud server, the leakage of engineering-related data is avoided while reducing the data transmission volume.

[0035] A7, comparing and analyzing the current engineering quantity indicator data in the current engineering cost data based on the typical characteristic data of the approximate engineering quantity indicator; The current engineering usage index data is the engineering usage index data in the current engineering cost data, which reflects the usage of materials, labor and machinery in the engineering project. Between similar engineering projects, the usage indicators usually have a certain degree of similarity and can be used to compare and analyze abnormal situations in the cost data. For example, feature extraction can be performed on the current engineering usage index data and feature comparison can be performed with typical feature data of similar engineering usage indicators to discover abnormalities.

[0036] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can match the engineering cost data of similar historical projects through the cloud server, realize the comparative analysis of the current project and similar historical projects, and thus realize the abnormal analysis of the current engineering cost data. The data is processed through the engineering profile feature extraction algorithm, and the historical project database is combined to match similar projects. The typical characteristics of the usage indicators are analyzed to provide data support for cost data analysis, which can determine whether there are abnormalities in the current engineering cost data, save the time cost of manual verification, and optimize resource allocation.

[0037] Furthermore, the engineering profile feature extraction algorithm includes the following steps: B1, obtain the name of each overview subject and the corresponding overview subject key value in the current project overview data; The general subject names are the names of the subjects in the current project general data, such as the number of buildings, number of floors, floor height, building height, number of households, building area, etc. The key value of the overview subject is the numerical value corresponding to each subject name. For example, the key value corresponding to the number of buildings is 5, the key value corresponding to the number of floors is 32, the key value corresponding to the floor height is 3.6, and so on.

[0038] B2, generating corresponding overview subject name word vectors based on the names of each overview subject using a preset word vector algorithm; The word vector algorithm is a pre-set algorithm used to convert the names of various general subject areas into corresponding vectors. The word vector algorithm can be an existing general word vector algorithm or an algorithm that is self-adjusted and set based on relevant vocabulary of construction projects. The overview subject name word vector is the word vector corresponding to the overview subject name.

[0039] B3, generating a key-value pair of each overview subject name word vector according to the overview subject name word vector and the overview subject key value combination corresponding to each overview subject name; The profile subject name word vector key-value pair is a data pair generated by combining the profile subject name word vector and the profile subject key value corresponding to the profile subject name.

[0040] B4, generates the current project overview feature data based on the key-value pairs of the word vectors of each overview subject name.

[0041] The current engineering overview feature data is a data collection of all overview subject name word vector key-value pairs.

[0042] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can realize the semantic expression and characterization processing of engineering data by converting the names of each profile subject into word vectors and combining them with corresponding key values, which can improve the accuracy of matching similar engineering projects and reduce the error of subject matching in data processing. At the same time, it provides a standardized data structure for subsequent comparative analysis, enhances the system's unified recognition ability of different expressions, and makes engineering cost analysis more accurate and efficient.

[0043] Furthermore, the step A4 includes the following steps: A401, sequentially obtaining historical engineering project overview feature data corresponding to the cost data of each historical engineering project from the historical engineering project database of the cloud server and combining them to generate a historical engineering overview feature data set; The historical engineering overview feature data is feature data of the engineering overview data of the historical engineering project cost data, which can be pre-selected in the cloud service and feature extracted from the engineering overview data of the historical engineering project cost data using an engineering overview feature extraction algorithm and stored in the historical engineering project database; The historical engineering profile feature dataset is a collection of the profile feature data of various historical engineering projects.

[0044] A402, sequentially obtain the word vector key-value pairs of each overview subject name from the current project overview feature data; Obtain the key-value pairs of the overview subject name word vectors one by one from the current project overview feature data for matching.

[0045] A403, sequentially matching the historical project overview feature data according to the overview subject name word vector key-value pairs; Try to match the overview feature data of each historical project through the overview subject name word vector key-value pairs.

[0046] A404: If the historical engineering overview feature data does not contain a historical engineering overview subject name word vector that matches the key-value pair of the overview subject name word vector, then the historical engineering overview feature data is removed from the historical engineering overview feature data set; The word vector of the subject name of the historical engineering overview is the word vector corresponding to the name of each subject name of the historical engineering overview in the historical engineering overview feature data; If the overview subject name word vector in the overview subject name word vector key-value pair fails to match the same or similar historical engineering overview subject name word vector in the historical engineering overview feature data, the historical engineering overview feature data is removed from the historical engineering overview feature data set.

[0047] A405, defining the remaining historical engineering overview feature data in the historical engineering overview feature data set as approximate historical engineering project cost data; After removing all unmatched historical engineering overview feature data sets, the remaining historical engineering overview feature data are historical engineering project cost data that are similar to the current engineering cost data.

[0048] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can obtain the historical engineering profile feature data set from the cloud server database, and then use the subject name word vector key-value pairs of the current project to match one by one, intelligently eliminate historical data that does not meet the conditions, and finally retain similar historical engineering data, thereby improving the matching accuracy, reducing unnecessary data interference, and ensuring that subsequent analysis is based on the most relevant historical data, thereby providing a more reliable reference basis for engineering cost estimation and improving the accuracy of the analysis results.

[0049] Furthermore, the step A403 includes the following steps: A4031, sequentially obtains the word vector key-value pairs of each historical engineering overview subject name from the historical engineering overview feature data; Obtain the historical engineering overview subject name word vector key-value pairs one by one from the historical engineering overview feature data; The word vector key-value pairs of the subject names of the historical engineering overview are the word vector key-value pairs of the subject names of the historical engineering overview corresponding to the historical engineering overview in the historical engineering overview feature data.

[0050] A4032, based on the overview subject name word vector in the overview subject name word vector key-value pair and the historical engineering overview subject name word vector in each historical engineering overview subject name word vector key-value pair, calculates the corresponding cosine similarity and defines it as the subject name word vector similarity; The subject name word vector similarity is the cosine similarity between the overview subject name word vector and the historical engineering overview subject name word vector.

[0051] A4033, if the subject name word vector similarity is greater than or equal to the preset subject name similarity threshold, then respectively obtain the overview subject key value corresponding to the overview subject name word vector and the historical engineering overview subject key value corresponding to the historical engineering overview subject name word vector; The subject name similarity threshold is a pre-set reference value used to determine the degree of similarity between subject name word vectors. It can be set manually based on experience or statistical data. The key value of the historical engineering overview subject is the key value corresponding to the historical engineering overview subject name word vector in the historical engineering overview subject name word vector key-value pair.

[0052] A4034, if the data types of the overview subject key value and the historical engineering overview subject key value are both text type, then the corresponding overview subject key value word vector and historical engineering overview subject key value word vector are generated respectively based on the overview subject key value and the historical engineering overview subject key value using the word vector algorithm; The profile subject key word vector is the word vector corresponding to the profile subject key value of the text type; The key-value word vector of the historical engineering overview subject is the word vector corresponding to the key-value of the historical engineering overview subject of the text type.

[0053] A4035 calculates the corresponding cosine similarity based on the overview subject key word vector and the historical engineering overview subject key word vector and defines it as the subject key word vector item similarity; The similarity of subject key word vector items is the cosine similarity between the overview subject key word vector and the historical engineering overview subject key word vector.

[0054] A4036, if the subject key word vector item similarity is greater than or equal to the preset subject key similarity threshold, then the profile subject name word vector key-value pair is defined to match the historical project profile feature data; The subject key value similarity threshold is a pre-set reference value used to determine the similarity between subject key value word vector items; When the similarity of the subject key word vector item is not less than the subject key similarity threshold, it means that there is a high degree of correlation between the overview subject key word vector and the historical engineering overview subject key word vector, and it can be determined that the overview subject name word vector key-value pair matches the historical engineering overview feature data.

[0055] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can perform initial screening by calculating the cosine similarity between the subject name word vectors, and then convert the text type subject key values ​​into word vectors and calculate their similarity. The double-layer screening mechanism can not only identify engineering overview elements that are expressed differently but are essentially the same or similar, but can also flexibly adjust the matching accuracy through preset thresholds, effectively avoiding the limitations of traditional character matching, improving the comparison accuracy between heterogeneous data, and providing a more accurate data basis for engineering cost analysis.

[0056] Furthermore, the step A5 includes the following steps: A501, generating corresponding historical engineering usage standardization index data using a preset data standardization algorithm based on each historical engineering usage index data; The data standardization algorithm is a pre-set algorithm used to standardize data, including, for example, data cleaning, data completion, and normalization. The historical engineering quantity standardized indicator data is the historical engineering quantity indicator data that has been standardized.

[0057] A502, generating a corresponding historical engineering quantity feature vector based on the standardized index data of each historical engineering quantity using a preset engineering quantity feature extraction algorithm; The engineering quantity feature extraction algorithm is a pre-set feature extraction algorithm used to extract feature data from the historical engineering quantity standardized indicator data; The historical engineering usage characteristic vector is the characteristic data corresponding to the historical engineering usage standardized indicator data.

[0058] A503 calculates the corresponding central vector based on each historical engineering quantity characteristic vector and defines it as the typical engineering quantity characteristic vector; The typical engineering quantity characteristic vector is the central vector corresponding to each historical engineering quantity characteristic vector, which can be determined by calculating the vector mean or clustering algorithm.

[0059] A504 calculates the corresponding Euclidean distance between the typical engineering quantity characteristic vector and each historical engineering quantity characteristic vector and defines it as the engineering quantity characteristic distance; The engineering quantity characteristic distance is the Euclidean distance between the characteristic vectors of each historical engineering quantity and the characteristic vector of the typical engineering quantity.

[0060] A505, determining the minimum value among all engineering quantity characteristic distances and defining the corresponding historical engineering quantity characteristic vector as the typical characteristic data of the approximate engineering quantity indicator; The historical engineering usage characteristic vector closest to the typical engineering usage characteristic vector is determined by finding the engineering usage characteristic distance with the smallest value, which is the typical characteristic data of the approximate engineering usage indicator.

[0061] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can process historical engineering usage indicators through data standardization, eliminate the influence of different units and magnitudes, and then use the feature extraction algorithm to convert complex engineering usage data into feature vectors, and calculate the center vector as a typical feature reference. The historical engineering cost data closest to the typical feature is found through Euclidean distance calculation, which overcomes the limitations of traditional direct comparison, can capture the essential characteristics of engineering usage from multiple dimensions, improve the accuracy and representativeness of feature extraction, provide more scientific data support for engineering cost analysis, and enhance the accuracy and reference value of the analysis results.

[0062] Furthermore, the step A7 includes the following steps: A701, generating corresponding current engineering quantity standardization indicator data through a data standardization algorithm based on the current engineering quantity indicator data; The current engineering quantity standardization indicator data is the standardized data corresponding to the current engineering quantity indicator data.

[0063] A702, generating a corresponding current engineering quantity feature vector based on the current engineering quantity standardized index data using an engineering quantity feature extraction algorithm; The current engineering usage characteristic vector is characteristic data corresponding to the current engineering usage standardized index data.

[0064] A703 calculates the corresponding cosine similarity based on the current engineering usage characteristic vector and the vector corresponding to the typical characteristic data of the approximate engineering usage indicator and defines it as the usage indicator comparison similarity; The usage indicator comparison similarity is the cosine similarity between the current engineering usage feature vector and the vector corresponding to the typical feature data of the approximate engineering usage indicator.

[0065] A704: If the usage indicator comparison similarity is less than the preset comparison similarity threshold, it is determined that the current project cost data is abnormal.

[0066] The comparison similarity threshold is a pre-set reference value used to determine the size of the usage index comparison similarity. If the usage index comparison similarity is too small, it means that there is a large deviation between the current project cost data and the typical approximate project cost data, and there may be anomalies.

[0067] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can standardize and extract features of current engineering quantity indicators, convert complex data into feature vectors, and then calculate the cosine similarity between the vectors corresponding to the typical feature data of the approximate engineering quantity indicators, and determine the degree of similarity through a preset threshold, thereby realizing the identification of potential anomalies in the engineering cost data, providing an objective data evaluation standard, and effectively avoiding the subjectivity and inconsistency of human judgment. Compared with traditional methods, it can better capture the essential differences between data, improve the accuracy and reliability of anomaly detection, provide timely risk warnings for engineering cost management, and help decision makers quickly discover and correct potential problems.

[0068] Furthermore, the engineering cost data analysis method based on big data analysis and cloud computing further includes the following steps: A8: If the current project cost data is abnormal, multiple corresponding current project first-level subject usage index standardization data are generated based on the current project usage standardization index data and the preset first-level subject list; The first-level subject list is a pre-set list used to identify the first-level subjects in the engineering quantity indicator data. The first-level subjects can be assigned a certain code number for easy identification and search. For example, the first-level subject of a certain code number includes the quantity indicators of sub-subjects such as earthwork engineering, foundation pit support, slope, underground continuous front and foundation treatment. According to the first-level subject list, the current engineering quantity standardization indicator data can be divided into the standardized data corresponding to each first-level subject; The standardized data of the current engineering first-level subject usage indicators are the standardized data corresponding to each first-level subject in the current engineering usage standardized indicator data.

[0069] A9, receiving historical engineering quantity standardization indicator data corresponding to the typical characteristic data of the approximate engineering quantity indicator from the cloud server, and generating a plurality of corresponding typical engineering first-level subject quantity indicator standardization data according to the historical engineering quantity standardization indicator data and the first-level subject list; The standardized data of the usage indicators of the first-level subjects of typical projects are the standardized data corresponding to each first-level subject in the historical engineering usage standardized indicator data corresponding to the typical characteristic data of the approximate engineering usage indicators.

[0070] A10 generates a corresponding current project first-level subject usage feature vector based on the standardized data of the current project first-level subject usage indicators using a project usage feature extraction algorithm; The current engineering first-level subject usage characteristic vector is the characteristic vector corresponding to the standardized data of the current engineering first-level subject usage indicator.

[0071] A11, based on the standardized data of the consumption indicators of the first-level subjects of each typical project, generates the corresponding typical project first-level subject consumption feature vector through the engineering consumption feature extraction algorithm; The typical engineering first-level subject usage characteristic vector is the characteristic vector corresponding to the standardized data of the usage indicators of each typical engineering first-level subject.

[0072] A12, based on the first-level subject usage feature vector of each current project and the first-level subject usage feature vector of the corresponding typical project, calculates the corresponding cosine similarity and defines it as the first-level subject feature similarity; The typical engineering first-level subject usage characteristic vector corresponding to the current engineering first-level subject usage characteristic vector can be determined according to the first-level subject code sequence number; The first-level subject feature similarity is the cosine similarity between the first-level subject usage feature vector of the current project and the first-level subject usage feature vector of the corresponding typical project.

[0073] A13: If the first-level subject feature similarity is less than the preset subject similarity threshold, the corresponding subject name is determined based on the first-level subject list and defined as an abnormal usage subject. The subject similarity threshold is a pre-set reference value used to determine the similarity of the first-level subject features; If the similarity of the first-level subject characteristics is too small, it means that there is a large deviation between the usage indicators of the first-level subject in the current engineering usage indicator data and the usage indicators of the first-level subject in the typical historical engineering usage indicator data corresponding to the typical characteristic data of the approximate engineering usage indicators, and there may be anomalies.

[0074] Through the above steps, the engineering cost data analysis method based on big data analysis and cloud computing can accurately identify which specific first-level subjects have cost anomalies by dividing the engineering data into first-level subjects, extracting feature vectors respectively and performing cosine similarity calculation with the corresponding subjects of typical projects, thereby improving the meticulousness and pertinence of anomaly diagnosis, making the anomaly location clearer, providing a clear direction for subsequent rectification, and improving the pertinence and decision-making support capabilities of cost anomaly analysis.

[0075] refer to Figure 2 , this application also provides an engineering cost data analysis system based on big data analysis and cloud computing, including: Cloud computing module 10; local computing module 20; The cloud computing module 10 and the local computing module 20 are communicatively connected; The cloud computing module 10 includes a cloud storage module 11 and a cloud processing module 12, and the cloud storage module 11 and the cloud processing module 12 are data-connected. The cloud computing module 10 is mainly used to be deployed on a cloud server to provide sufficient computing power for data processing and analysis; The cloud storage module 11 is mainly used to be deployed on cloud services to store a large amount of historical project cost data; The cloud processing module 12 is mainly used to be deployed on the cloud service to perform data analysis and processing on a large amount of historical engineering cost data; The local computing module 20 is mainly used to perform partial calculations locally, and to exchange data and perform collaborative processing with the cloud computing module 10 .

[0076] The engineering cost data analysis system based on big data analysis and cloud computing further includes an engineering cost analysis strategy, including the following steps: C1, obtaining current project overview data from preset current project cost data through the local calculation module 20; C2, generating corresponding current engineering profile feature data by the local computing module 20 according to the current engineering profile data using a preset engineering profile feature extraction algorithm; C3, sending the current project profile feature data to the cloud computing module 10 through the local computing module 20; C4, the cloud computing module 10 matches the current project profile feature data with a preset historical project database to determine corresponding approximate historical project cost data; C5, determining corresponding typical characteristic data of approximate engineering quantity indicators based on the historical engineering quantity indicator data of each approximate historical engineering project cost data in the cloud computing module 10; C6, receiving typical characteristic data of approximate engineering usage indicators from the cloud computing module 10 to the local computing module 20; C7, comparing and analyzing the current engineering quantity indicator data in the current engineering cost data according to the typical characteristic data of the approximate engineering quantity indicator by the local calculation module 20.

[0077] Through the above technical solution, the engineering cost data analysis system based on big data analysis and cloud computing can match the engineering cost data of similar historical projects through the cloud server, realize the comparative analysis of the current project and similar historical projects, and thus realize the abnormal analysis of the current engineering cost data. It processes the data through the engineering profile feature extraction algorithm, matches similar projects in combination with the historical project database, and analyzes the typical characteristics of the usage indicators to provide data support for cost data analysis. It can determine whether there are abnormalities in the current engineering cost data, save the time cost of manual verification, and optimize resource allocation.

[0078] The above are all preferred embodiments of the present application and are not intended to limit the scope of protection of this application. Unless otherwise specified, any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features. In other words, unless otherwise specified, each feature is merely an example of a series of equivalent or similar features.

Claims

1. A method for analyzing engineering cost data based on big data analysis and cloud computing, characterized in that: The following steps are involved: A1, obtains the current project overview data from the preset current project cost data; A2, generating corresponding current project profile feature data according to the current project profile data using a preset project profile feature extraction algorithm; A3, sending the current project profile feature data to the preset cloud server; A4, matching the current project profile feature data with a preset historical project database on the cloud server to determine corresponding approximate historical project cost data; A5, determining, at the cloud server, corresponding typical characteristic data of approximate engineering quantity indicators based on the historical engineering quantity indicator data of each approximate historical engineering project cost data; A6, receiving typical characteristic data of approximate engineering usage indicators from the cloud server; A7, compare and analyze the current engineering quantity indicator data in the current engineering cost data based on the typical characteristic data of the approximate engineering quantity indicator.

2. The method for analyzing engineering cost data based on big data analysis and cloud computing according to claim 1, characterized in that: The engineering profile feature extraction algorithm comprises the following steps: B1, obtain the name of each overview subject and the corresponding overview subject key value in the current project overview data; B2, generating corresponding overview subject name word vectors based on the names of each overview subject using a preset word vector algorithm; B3, generating a key-value pair of each overview subject name word vector according to the overview subject name word vector and the overview subject key value combination corresponding to each overview subject name; B4, generates the current project overview feature data based on the key-value pairs of the word vectors of each overview subject name.

3. The method for analyzing engineering cost data based on big data analysis and cloud computing according to claim 2 is characterized in that: Step A4 includes the following steps: A401, sequentially obtaining historical engineering project overview feature data corresponding to the cost data of each historical engineering project from the historical engineering project database of the cloud server and combining them to generate a historical engineering overview feature data set; A402, sequentially obtain the word vector key-value pairs of each overview subject name from the current project overview feature data; A403, sequentially matching the historical project overview feature data according to the overview subject name word vector key-value pairs; A404: If the historical engineering overview feature data does not contain a historical engineering overview subject name word vector that matches the key-value pair of the overview subject name word vector, then the historical engineering overview feature data is removed from the historical engineering overview feature data set; A405 defines the remaining historical engineering overview feature data in the historical engineering overview feature data set as approximate historical engineering project cost data.

4. The method for analyzing engineering cost data based on big data analysis and cloud computing according to claim 3 is characterized in that: Step A403 includes the following steps: A4031, sequentially obtains the word vector key-value pairs of each historical engineering overview subject name from the historical engineering overview feature data; A4032, based on the overview subject name word vector in the overview subject name word vector key-value pair and the historical engineering overview subject name word vector in each historical engineering overview subject name word vector key-value pair, calculates the corresponding cosine similarity and defines it as the subject name word vector similarity; A4033, if the subject name word vector similarity is greater than or equal to the preset subject name similarity threshold, then respectively obtain the overview subject key value corresponding to the overview subject name word vector and the historical engineering overview subject key value corresponding to the historical engineering overview subject name word vector; A4034, if the data types of the overview subject key value and the historical engineering overview subject key value are both text type, then the corresponding overview subject key value word vector and historical engineering overview subject key value word vector are generated respectively based on the overview subject key value and the historical engineering overview subject key value using the word vector algorithm; A4035 calculates the corresponding cosine similarity based on the overview subject key word vector and the historical engineering overview subject key word vector and defines it as the subject key word vector item similarity; A4036, if the subject key word vector item similarity is greater than or equal to the preset subject key similarity threshold, then the profile subject name word vector key-value pair is defined to match the historical project profile feature data.

5. The method for analyzing engineering cost data based on big data analysis and cloud computing according to claim 4 is characterized in that: Step A5 includes the following steps: A501, generating corresponding historical engineering usage standardization index data using a preset data standardization algorithm based on each historical engineering usage index data; A502, generating a corresponding historical engineering quantity feature vector based on the standardized index data of each historical engineering quantity using a preset engineering quantity feature extraction algorithm; A503 calculates the corresponding central vector based on each historical engineering quantity characteristic vector and defines it as the typical engineering quantity characteristic vector; A504 calculates the corresponding Euclidean distance between the typical engineering quantity characteristic vector and each historical engineering quantity characteristic vector and defines it as the engineering quantity characteristic distance; A505 determines the minimum value among all engineering usage characteristic distances and defines its corresponding historical engineering usage characteristic vector as the typical characteristic data of the approximate engineering usage indicator.

6. The method for analyzing engineering cost data based on big data analysis and cloud computing according to claim 5 is characterized in that: Step A7 includes the following steps: A701, generating corresponding current engineering quantity standardization indicator data through a data standardization algorithm based on the current engineering quantity indicator data; A702, generating a corresponding current engineering quantity feature vector based on the current engineering quantity standardized index data using an engineering quantity feature extraction algorithm; A703 calculates the corresponding cosine similarity based on the current engineering usage characteristic vector and the vector corresponding to the typical characteristic data of the approximate engineering usage indicator and defines it as the usage indicator comparison similarity; A704: If the usage indicator comparison similarity is less than the preset comparison similarity threshold, it is determined that the current project cost data is abnormal.

7. The method for analyzing engineering cost data based on big data analysis and cloud computing according to claim 6, characterized in that: Further comprising the steps of: A8: If the current project cost data is abnormal, multiple corresponding current project first-level subject usage index standardization data are generated based on the current project usage standardization index data and the preset first-level subject list; A9, receiving historical engineering quantity standardization indicator data corresponding to the typical characteristic data of the approximate engineering quantity indicator from the cloud server, and generating a plurality of corresponding typical engineering first-level subject quantity indicator standardization data according to the historical engineering quantity standardization indicator data and the first-level subject list; A10 generates a corresponding current project first-level subject usage feature vector based on the standardized data of the current project first-level subject usage indicators using a project usage feature extraction algorithm; A11, based on the standardized data of the consumption indicators of the first-level subjects of each typical project, generates the corresponding typical project first-level subject consumption feature vector through the engineering consumption feature extraction algorithm; A12, based on the first-level subject usage feature vector of each current project and the first-level subject usage feature vector of the corresponding typical project, calculates the corresponding cosine similarity and defines it as the first-level subject feature similarity; A13: If the first-level subject feature similarity is less than the preset subject similarity threshold, the corresponding subject name is determined based on the first-level subject list and defined as an abnormal usage subject.

8. An engineering cost data analysis system based on big data analysis and cloud computing, characterized in that: include: Cloud computing module; Local computing module; The cloud computing module and the local computing module are communicatively connected; The cloud computing module includes a cloud storage module and a cloud processing module, and the cloud storage module and the cloud processing module are data-connected. The engineering cost data analysis system based on big data analysis and cloud computing further includes an engineering cost analysis strategy, including the following steps: C1, obtaining current project overview data from preset current project cost data through the local calculation module; C2, generating corresponding current engineering profile feature data by the local computing module according to the current engineering profile data using a preset engineering profile feature extraction algorithm; C3, sending the current project profile feature data to the cloud computing module through the local computing module; C4, the cloud computing module matches the current project profile feature data with a preset historical project database to determine corresponding approximate historical project cost data; C5, determining, in the cloud computing module, corresponding typical characteristic data of approximate engineering quantity indicators based on the historical engineering quantity indicator data of each approximate historical engineering project cost data; C6, receiving typical characteristic data of approximate engineering usage indicators from the cloud computing module to the local computing module; C7, comparing and analyzing the current engineering quantity indicator data in the current engineering cost data according to the typical characteristic data of the approximate engineering quantity indicator by the local calculation module.

Citation Information

Patent Citations

  • Geological disaster prevention and control project cost management system and method thereof

    CN116977001A

  • Power grid project cost auxiliary analysis method and system

    CN117057835A

  • Engineering cost estimation method and device based on artificial intelligence

    CN117314179A