Engineering cost information optimization integration method based on Internet big data service
Through the optimization and integration method of engineering cost information based on Internet big data services, the problem of inefficient engineering cost information management in the existing technology is solved, efficient data collection, in-depth analysis, accurate screening and intuitive presentation are realized, and the security and reliability of data are ensured.
Patent Information
- Application Number
- CN202510343701.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-22
- Publication Date
- 2025-06-27
AI Technical Summary
The existing engineering cost information management methods are inefficient and prone to errors, making it difficult to achieve efficient data collection, in-depth analysis, precise screening, intuitive presentation and security management.
The engineering cost information optimization integration method based on Internet big data services is adopted, data is collected through intelligent crawler technology, data integration is carried out based on semantic analysis, data mining algorithm is used to analyze association relationships, build prediction models, and information is presented through visualization technology to ensure the security and reliability of the data.
It realizes efficient collection and integration of engineering cost information, improves the accuracy of data analysis and prediction, improves the efficiency and readability of information screening and presentation, and ensures the security and reliability of data.
Smart Images

Figure CN120216749A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of project cost, and particularly to an optimized integration method for project cost information based on Internet big data services. Background Art
[0002] In the current field of construction projects, the accuracy and efficiency of project cost management are crucial for the successful implementation of projects. With the rapid development of Internet technology and the advent of the big data era, a vast amount of project cost information has been continuously generated and accumulated on the Internet. However, this information faces many problems and urgently requires an effective optimized integration method.
[0003] Traditional acquisition of project cost information mainly relies on manual collection and collation, which is inefficient and error-prone. Project cost engineers need to spend a lot of time collecting information from various scattered channels, such as building materials market research reports, data from engineering consulting companies, cost index documents issued by the government, etc. Not only is the data collection cycle long, but it is also difficult to ensure the integrity and timeliness of the data. The data formats and standards from different sources are often inconsistent, and a large amount of manual checking and conversion work is required during the integration process, further increasing the workload and the probability of errors.
[0004] In terms of data analysis, traditional methods are mostly based on simple statistical analysis and empirical judgment, and it is difficult to discover the deep-seated correlations and laws behind the data. For example, it is difficult to accurately analyze the complex relationships between the price fluctuations of building materials and many factors, such as the macroeconomic situation, regional supply and demand relationships, seasonal changes, project types, etc., thus unable to provide strong support for the accurate prediction of project costs. This results in situations where the project budget is often overrun or too low, affecting the project quality during the project budget preparation process.
[0005] There are also dilemmas in information screening and application. Facing a vast amount of data, there is a lack of an effective screening mechanism, and it is difficult to quickly and accurately extract useful information according to the unique needs of specific projects. Moreover, the traditional project cost report forms are single, mostly in text and simple tables, lacking intuitiveness and interactivity, which is not conducive to the quick understanding and decision-making of all parties involved in the project.
[0006] In addition, the issues of data security and privacy protection are becoming increasingly prominent. In the Internet environment, project cost information involves the commercial secrets of many enterprises and personal privacy. Without effective encryption and management measures, it is very easy to suffer from data leakage and malicious attacks, bringing huge losses to enterprises and individuals.
[0007] In summary, the existing methods for engineering cost information management can no longer meet the requirements of the times. Therefore, it is of great practical significance to develop an optimized integration method for engineering cost information based on Internet big data services, which can achieve efficient collection, in-depth analysis, precise screening, intuitive presentation, and secure management of engineering cost information, improve the overall level of engineering cost management, and promote the healthy and sustainable development of the construction engineering industry. Summary of the Invention
[0008] An optimized integration method for engineering cost information based on Internet big data services proposed by the present invention is used to solve the problems mentioned in the above-mentioned prior art.
[0009] To achieve the above object, the present invention adopts the following technical solutions: An optimized integration method for engineering cost information based on Internet big data services includes the following steps:
[0010] S1. Data collection and integration step: Collect engineering cost-related data from multiple Internet data sources; use intelligent web crawler technology to accurately locate and collect data according to preset keywords and data characteristics;
[0011] Clean and integrate the collected data, remove duplicate data, error data, and abnormal data, and associate and integrate data from different sources through a data matching algorithm based on semantic analysis. Set the data semantic similarity threshold as M t , when the data semantic similarity is greater than M t , M t =8, perform data fusion to form a standardized basic engineering cost information database;
[0012] S2. Data analysis and mining step: Analyze the integrated engineering cost data using data mining algorithms, and use the Apriori association rule algorithm to mine the multi-dimensional association relationships between building material prices, engineering geographical locations, seasonal factors, and engineering types. Set the minimum support as S min , S min =0.2, the minimum confidence as C min , C min =0.6, and mine strongly associated data item sets;
[0013] Build an engineering cost prediction model based on big data analysis, and use a neural network integrated with a particle swarm optimization algorithm to build the model. The neurons in the input layer of the neural network correspond to engineering feature parameters, and the output layer is the predicted engineering cost value. The particle swarm optimization algorithm is used to optimize the weights and thresholds of the neural network. Set the particle swarm size as N, N = 50, the inertia weight w linearly decreases from 0.9 to 0.4, and the learning factors c1 = c2 = 2;
[0014] S3. Information Optimization and Screening Step: Optimize and screen the project cost information according to the user's needs and specific project requirements. The user can set screening conditions, including the project location, project budget range, project construction period, and project environmental protection requirements. The system screens out the information that meets the requirements from the basic information database according to the set conditions;
[0015] Rank the screened information by priority. Considering comprehensively the timeliness, reliability, authority of data sources, and the matching degree between data and user needs, adopt a dynamic weighted ranking algorithm. Set the timeliness weight as W t and the reliability weight as W r and the authority weight as W a and the matching degree weight as W m , where W t = 0.25, W r = 0.3, W a = 0.2, W m = 0.25;
[0016] S4. Integration and Visual Presentation Step: Integrate the optimized and screened project cost information to form a complete project cost information report;
[0017] Adopt visualization technology to present the project cost information in an intuitive chart form, including bar charts showing the proportion of costs in different parts, line charts reflecting the trend of material prices, and heat maps presenting cost differences in different regions. When the user hovers the mouse over a data point, the detailed information of the data point and the associated data information can be automatically displayed;
[0018] Data Update and Verification Step: Establish a real-time data update mechanism, establish a long connection with the data source. When there is data update in the data source, immediately obtain and process it. Adopt data fingerprint technology to verify the integrity of the updated data. Set the data fingerprint function as F(d), where d is the data block. By comparing the fingerprints of the data before and after the update, if the fingerprints match, perform the data update operation;
[0019] Regularly verify the data in the information database. Adopt the cross-validation method, divide the data into multiple groups, and each group of data is used as the validation set and the training set respectively. By comparing the model prediction results with the actual data, calculate the deviation rate of the data P is the predicted value, A is the actual value. When the deviation rate exceeds the set threshold of 5%, re-collect and correct the data;
[0020] S5. Multi-source Data Fusion and Calibration Step: For data from different sources with differences, adopt the Bayesian estimation method for fusion and calibration. Let the data from different sources be D1, D2,..., D n , and their prior probabilities be P(D1), P(D2),..., P(D n), the likelihood functions are L(D1), L(D2), …, L(D n ), the fused data D f Calculated according to the formula
[0021] During the fusion process, the credibility of the data is evaluated, and the credibility weight is determined according to the historical accuracy of the data source and the data update frequency factor. The data with low credibility is marked and further verified;
[0022] S6. Model self - learning and optimization step: During the operation of the project cost prediction model, new data samples are automatically collected, and the incremental learning algorithm is used to update and optimize the model in real - time. Let the new data sample set be S n , and the model adjusts the neuron connection weight W n according to the data in S ij , and the adjustment formula is where η is the learning rate, δ i is the error term, x j is the input feature, α is the momentum factor, is the previous weight adjustment amount;
[0023] The performance of the model is evaluated, and the confusion matrix and F1 - value are used as evaluation indicators. When the F1 - value is lower than the set threshold of 0.85, the model reconstruction mechanism is triggered, and the algorithm is re - selected or the model structure is adjusted.
[0024] Furthermore, in the data collection and integration step, for the collection of building material price data, a multi - channel collection method is adopted, including data docking with major building material e - commerce platforms, scraping data from government price department websites, and collecting official quotes from building material production enterprises. The blockchain technology is used to build a data traceability system, and the fusion weight of data from different channels is determined according to the reliability of the data source.
[0025] Furthermore, the project cost prediction model constructed in the data analysis and mining step adopts a combination of the long short - term memory network LSTM and the random forest algorithm in deep learning algorithms. LSTM is used to learn the time - series characteristics of project cost data, and the random forest algorithm is used for classification and regression analysis of the data. Through the feature importance evaluation algorithm, the key features that have a greater impact on the project cost are screened out.
[0026] Furthermore, in the information optimization and screening step, when the screening conditions set by the user are fuzzy or incomplete, the system adopts an intelligent recommendation algorithm to automatically supplement and improve the screening conditions according to the user's historical usage records and partial feature information of the current project, using the fuzzy matching algorithm and the case - based reasoning method.
[0027] Furthermore, in the integration and visualization presentation step, add data source annotation and data update timestamp to the project cost information report, with the detailed degree of the annotation information not less than 95%. Set an interactive function in the visualization chart, where users can obtain detailed information of data points in the chart through mouse click operations, and export the visualization chart as a high-definition picture or vector graph format according to user needs.
[0028] Furthermore, it also includes the data security and privacy protection step. Use an encryption algorithm to encrypt and store and transmit the collected project cost data. The encryption algorithm uses the AES encryption algorithm with a key length of 256 bits. Anonymize the user's personal information and operation records. At the same time, establish a data access permission management system to allocate different access permissions according to user roles and operation needs.
[0029] Furthermore, establish a data quality monitoring and evaluation mechanism to evaluate the data quality in the project cost information database. Use the principal component analysis PCA algorithm to perform dimensionality reduction processing on the data, extract features, and then calculate the information entropy of the data p i as the probability distribution of data features to evaluate the uncertainty of the data. When the data quality is lower than the set threshold, automatically start the data update and correction program.
[0030] Furthermore, in the data analysis and mining step, establish an outlier detection model for project cost data, combining a statistics-based method and a clustering analysis method. Set the outlier judgment threshold as O t . When the data deviates from the normal range by more than O t , it is determined as an outlier. Use the isolation forest algorithm to analyze the outlier to determine the isolation degree of the outlier E(h(x)) is the average path length of the data point x in the isolation forest, and C(h(x)) is the average path length of the data point x in the completely random tree. Mark and analyze the outlier to analyze the cause of the outlier.
[0031] Compared with the existing technology, the beneficial effects of the present invention are:
[0032] First, in terms of data processing efficiency and accuracy, the intelligent crawler and the precise matching algorithm achieve high-speed collection and efficient integration of massive data, with the data accuracy exceeding 95%, greatly saving labor and time costs, and reducing errors caused by manual operations. For example, compared with the traditional manual collection method, the data collection speed can be increased by dozens of times, and it can quickly associate data from different sources to form a comprehensive information database.
[0033] Secondly, powerful data analysis and mining capabilities are highlighted. By using advanced algorithms to mine multi-dimensional correlation relationships and constructing a high-precision prediction model, the prediction error is controlled within ±8%, which can provide a scientific and reliable reference basis for project cost, effectively avoid budget overrun or shortage, and improve the accuracy and scientific nature of project cost control.
[0034] Furthermore, the information optimization screening and presentation are excellent. According to the dynamic weighted ranking of multiple factors, the correlation of the screening results exceeds 92%, and the required information is accurately pushed to users. The visual charts combined with the intelligent prompt function not only intuitively display the cost composition and trends, with the interactive response time not exceeding 0.3 seconds, but also facilitate users to deeply explore the data details, greatly improving the readability and usability of the information and facilitating the rapid decision-making of all parties involved in the project.
[0035] In terms of data management, the real-time update and verification mechanism and the multi-source data fusion and calibration ensure that the data is timely, accurate and consistent, and the data update and correction efficiency exceeds 90%. The data quality monitoring and evaluation mechanism further guarantees the data reliability and provides a solid data foundation for project cost management.
[0036] From the security perspective, encrypted storage and transmission, anonymization processing and strict access right management minimize the risk of data leakage. The identifiability of anonymized information is less than 0.5%, effectively protecting the enterprise's business secrets and personal privacy and enhancing users' trust in the system. Brief Description of the Drawings
[0037] Figure 1 It is a schematic block diagram of an optimized integration method for project cost information based on Internet big data services proposed by the present invention. Detailed Embodiment
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0040] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, and it can be the internal connection of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0041] Reference Figure 1 :A method for optimizing and integrating engineering cost information based on Internet big data services, comprising the following steps:
[0042] Data collection and integration steps: Collect project cost-related data from multiple Internet data sources, including building material price data, project equipment rental price data, labor cost data, project cost index data in different regions, project design parameter data, and project construction process data. Using intelligent crawler technology, data can be accurately located and collected based on preset keywords and data features. The data collection speed is not less than 1,000 pieces / second, and the accuracy of the collected data is not less than 95%.
[0043] Clean and integrate the collected data to remove duplicate data, erroneous data, and abnormal data. Use a data matching algorithm based on semantic analysis to associate and integrate data from different sources. Set the data semantic similarity threshold as M. t , when the data semantic similarity is greater than M t , M t =8, data fusion is performed to form a standardized basic information database of engineering cost, and the data in the database covers at least 80 different engineering cost information categories.
[0044] Data analysis and mining steps: Use data mining algorithms to conduct in-depth analysis of the integrated project cost data to mine the inherent associations and potential rules between the data. Use the improved Apriori association rule algorithm to mine the multi-dimensional association between the price of building materials and the geographical location of the project, seasonal factors, and the type of project. Set the minimum support as S min , S min =0.2, the minimum confidence level is C min , C min= 0.6, and strongly correlated data item sets are mined out.
[0045] Based on big data analysis, a project cost prediction model is constructed. The amount of model training data is not less than 1.5 million, covering the cost information of projects of different types and scales. A neural network integrated with a particle swarm optimization algorithm is used to construct the model. The input layer neurons of the neural network correspond to project feature parameters, and the output layer is the predicted project cost value. The particle swarm optimization algorithm is used to optimize the weights and thresholds of the neural network. Let the size of the particle swarm be N, N = 50, the inertial weight w linearly decreases from 0.9 to 0.4, and the learning factors c1 = c2 = 2. Through continuous iterative adjustment, the prediction error of the model is controlled within ±8%.
[0046] Steps of information optimization and screening: Optimize and screen the project cost information according to user requirements and specific project requirements. Users can set screening conditions, such as project location, project budget range, project construction period, project environmental protection requirements, etc. The system screens out the information that meets the requirements from the basic information database according to the set conditions, and the relevance of the screening results is not less than 92%.
[0047] Rank the screened information by priority, comprehensively considering factors such as the timeliness, reliability, authority of data sources, and the degree of match between the data and user requirements. Use a dynamic weighted ranking algorithm. Let the timeliness weight be W t 、the reliability weight be W r 、the authority weight be W a 、and the match degree weight be W m , where W t = 0.25, W r = 0.3, W a = 0.2, W m = 0.25, providing users with accurate and orderly reference for project cost information.
[0048] Steps of integration and visual presentation: Integrate the optimized and screened project cost information to form a complete project cost information report. The report content includes project overview, details of various expenses, cost analysis results, risk assessment, cost control suggestions, etc. The report format is standardized and can be customized according to user requirements. Multiple formats such as PDF and Excel can be generated, and the time to generate the report does not exceed 5 minutes.
[0049] Visualization technology is adopted to present project cost information in an intuitive chart form. For example, bar charts show the proportion of costs in different parts, line charts reflect the trend of material prices, and heat maps present cost differences in different regions, etc. The clarity of the visualization charts is not less than 1080p, and an intelligent hint function is set in the visualization charts. When the user's mouse hovers over a data point, the detailed information of the data point and the associated data information can be automatically displayed, and the interactive response time does not exceed 0.3 seconds, facilitating the user to quickly understand and analyze project cost information and improving the readability and usability of the information.
[0050] Data update and verification steps: Establish a real-time data update mechanism, establish a long connection with the data source. When there is data update in the data source, immediately obtain and process it. Adopt data fingerprint technology to verify the integrity of the updated data. Let the data fingerprint function be F(d), where d is the data block. By comparing the fingerprints of the data before and after the update, if the fingerprints match, then perform the data update operation to ensure that the data is updated to the project cost basic information database in a timely and accurate manner.
[0051] Regularly verify the data in the information database. Adopt the cross-validation method, divide the data into multiple groups, and each group of data is used as the validation set and the training set respectively. By comparing the model prediction results with the actual data, calculate the deviation rate of the data. P is the predicted value, A is the actual value. When the deviation rate exceeds the set threshold of 5%, re-collect and correct the data to ensure the reliability of the data.
[0052] Multi-source data fusion and calibration steps: For data from different sources with differences, adopt the Bayesian estimation method for fusion and calibration. Let the data from different sources be D1, D2,..., D n , and their prior probabilities are P(D1), P(D2),..., P(D n ), and the likelihood functions are L(D1), L(D2),..., L(D n ). The fused data D f is calculated according to the formula. Improve the accuracy and consistency of the data.
[0053] During the fusion process, evaluate the credibility of the data, determine the credibility weight according to factors such as the historical accuracy of the data source and the data update frequency, mark the data with low credibility and further verify it to ensure the quality of the fused data.
[0054] Model self-learning and optimization steps: During the operation of the project cost prediction model, automatically collect new data samples, and adopt the incremental learning algorithm to update and optimize the model in real time. Let the new data sample set be S n , and the model adjusts the neuron connection weight W n according to the data in S ij, the adjustment formula is (η is the learning rate, δ i is the error term, x j is the input feature, α is the momentum factor, is the previous weight adjustment amount), enabling the model to adapt to market changes and new engineering types, and maintaining the accuracy and timeliness of predictions.
[0055] Regularly evaluate the performance of the model, using the confusion matrix and F1 value as evaluation metrics. When the F1 value is lower than the set threshold (such as 0.85), trigger the model reconstruction mechanism, reselect the algorithm or adjust the model structure to improve the generalization ability and robustness of the model.
[0056] In the data collection and integration step, a comprehensive and efficient multi-channel collection mechanism was constructed for the collection of building material price data. When connecting with major building material e-commerce platforms for data docking, a specially developed adaptable data interface program was used. This program can be customized and connected according to the API specifications of different e-commerce platforms. By sending accurate data request commands to the platform, multi-dimensional material price data information such as product name, specification model, price, inventory quantity, and shipping location can be obtained. At the same time, to ensure the real-time nature of the data, an intelligent data update trigger was set. Once there is a change in the price data on the platform, it can sense and capture the updated information within an extremely short time (accurate to the millisecond level).
[0057] When scraping data from the government price department website, advanced web crawler technology was adopted. First, in-depth analysis and modeling were carried out on the page structure, data storage method, and update rules of the price department website. The crawler program can accurately locate the web page area or database table where the material price data is located based on these models. During the scraping process, a multi-threaded concurrent scraping strategy was used to effectively improve the data scraping speed. For example, 10 - 20 threads can be simultaneously enabled for data acquisition operations. And, to cope with the possible anti-crawler mechanisms of the website, dynamic IP proxy technology and simulated browser behavior technology were adopted, enabling the crawler program to stably and continuously scrape data and ensuring the integrity of the data.
[0058] When collecting the official quotations of building material production enterprises, a data sharing and cooperation model was established with many building material production enterprises. Price information was obtained through the exclusive data interfaces provided by the enterprises or data files in formats such as Excel and XML regularly uploaded. For the parsing of data files, a general data parsing engine was developed, which can accurately identify key information such as material name, price, production batch, and validity period in different format files.
[0059] On this basis, a data traceability system is constructed using blockchain technology. A unique blockchain hash identifier is generated for each piece of collected building material price data, which contains key metadata such as data source information, collection time, collection location, etc. Through the distributed ledger feature of the blockchain, these data hash identifiers and their associated metadata are recorded on multiple nodes to form an immutable traceability chain. In the data fusion process, a multi-factor comprehensive evaluation method is adopted for the reliability assessment of data from different channels. In addition to considering the authority of the data source, factors such as historical accuracy, data update frequency, and data integrity of the data are also analyzed. For example, for the data of a building materials e-commerce platform with a data accuracy of over 98% and daily updates in the past year, its fusion weight may be determined to be 0.8. For some channels with occasional updates and relatively low data accuracy, their fusion weights will be correspondingly reduced, but the fusion weights of data from channels with high reliability are always not less than 0.6, so as to ensure the high quality and high credibility of the fused building material price data.
[0060] In the present invention, the engineering cost prediction model constructed in the data analysis and mining step adopts a combination of the long short-term memory network (LSTM) in the deep learning algorithm and the random forest algorithm. LSTM is used to learn the temporal characteristics of engineering cost data, and the random forest algorithm is used to classify and regression analyze the data, improving the model's processing ability and prediction accuracy for complex engineering cost data. The model training convergence speed is increased by more than 25% compared with a single algorithm. Through the feature importance evaluation algorithm, key features with greater influence on the engineering cost are screened out, reducing the data dimension and improving the model training efficiency.
[0061] In the present invention, in the information optimization and screening step, when the screening conditions set by the user are fuzzy or incomplete, the system adopts an intelligent recommendation algorithm to automatically supplement and improve the screening conditions according to the user's historical usage records and partial feature information of the current project. Using the fuzzy matching algorithm and the case-based reasoning method, the accuracy of the intelligent recommendation is not less than 85%, providing more accurate information screening results for the user.
[0062] In the integration and visualization presentation step, for the project cost information report, a multi-level annotation system is adopted in terms of data source annotation. Not only will it clearly record the specific website name, database name, or file source path from which the data is collected, but it will also detail the relevant information of the data provider, such as the enterprise name, institutional code (if applicable), and the specific department or person responsible for data collection, etc. For the data update timestamp, it is recorded accurately to the second level and in the international standard time format to facilitate unified time comparison and analysis globally. When users view the report, they can clearly see the detailed source description and accurate update time display next to each data item. The detailed degree of the annotation information is ensured to be no less than 95%. This high-detail annotation method provides a solid basis for users to trace the information source, enabling them to quickly locate the data source and also making it extremely convenient for users to intuitively judge the timeliness of the information, so that they can fully consider the reliability and effectiveness of the data when using the data for cost assessment and decision-making.
[0063] In the visualization chart part, an advanced front-end interaction technology framework is used to build the interaction function. When the user moves the mouse pointer over a data point in the chart, an immediate event listener is triggered, which quickly sends a request containing the data point coordinate information to the background server. The background server, within 0.2 seconds after receiving the request (the interaction response time does not exceed 0.2 seconds), based on the pre-built data index and query optimization algorithm, quickly locates the detailed data information associated with the data point, including the original data value, the engineering sub-project category to which the data belongs, the time period corresponding to the data, and the relevant calculation formula (if any), etc. Then these detailed information are presented in a concise and well-formatted pop-up window on the user interface. The style design of the pop-up window follows the best principles of user experience, with a clear title bar indicating the category of the data point detailed information, the main content presented in columns for each data item, and a quick button provided for users to copy the information conveniently.
[0064] In addition, a dedicated chart export engine has been developed for the export function of visual charts. When the user selects to export a visual chart, the engine first determines whether to export in high-definition image format or vector graphics format according to the user's needs. If the high-definition image format is selected, such as the common PNG format, the engine will use image rendering technology to draw the chart onto a virtual image canvas based on the size and resolution settings of the current chart. During the drawing process, it precisely controls the thickness of the lines, the filling accuracy of the colors, and the clarity of the text, etc., to ensure that the exported image has high fidelity and clarity, meeting the requirements for direct insertion and use in other documents without distortion. If the user selects to export in vector graphics format, such as SVG format, the engine will convert each graphic element in the chart, such as lines, rectangles, circles, text boxes, etc., into the corresponding vector graphic description language based on the drawing specifications of vector graphics. During the conversion process, it precisely encodes the geometric attributes, style attributes, and text content of the graphic elements to ensure that the exported vector graphic can be losslessly edited and scaled in software that supports the SVG format, facilitating the flexible use of visual chart resources by users in different document editing scenarios.
[0065] In the data security and privacy protection steps of the present invention, the industry-standard AES encryption algorithm is used to encrypt and store and transmit the collected project cost data. The AES encryption algorithm is renowned for its excellent security and high efficiency, and its key length is set to 256 bits. A key of this length can provide extremely strong encryption strength to resist the vast majority of brute-force cracking attacks. During the data encryption storage process, the project cost data is first divided into specific data block sizes, for example, each 128 bits is a data block, and then each data block is encrypted in turn. The encryption process is based on complex mathematical transformations, including operations such as byte substitution, row shift, column mixing, and round key addition. Through multiple rounds of iterative calculations, the original data is highly scrambled and diffused, thus ensuring that even if the data is illegally obtained, it is difficult for attackers to restore the original project cost data.
[0066] For the user's personal information and operation records, advanced anonymization techniques are adopted. By desensitizing the user identity identification information, such as performing irreversible hashing transformation or replacement operations on sensitive information like user names, ID numbers, contact information, etc., these information lose their direct identifiability after anonymization, and their identifiability is lower than 0.5%. At the same time, a complete data access permission management system is established. This system is built based on the role-based access control (RBAC) model, and multiple user roles are predefined, such as administrators, cost engineers, ordinary users, etc. For the administrator role, full read and write permissions to all data are granted so that they can perform operations such as system maintenance, data update and management. While ordinary users are only granted data query permissions, they can only browse and retrieve project cost information in specific interfaces of the system and cannot perform operations such as data modification, deletion or addition. When the user performs a login operation, the system will strictly verify the account password entered by the user. After verification, the corresponding permission menu and operation interface are loaded according to the role to which the user belongs. And when the user performs each data access operation, the system will check the user's permissions in real time to prevent the user from accessing data beyond their authority, thus effectively preventing the occurrence of data leakage and illegal access events.
[0067] In the present invention, a data quality monitoring and evaluation mechanism is established to regularly evaluate the data quality in the project cost information database. The evaluation indicators include data accuracy, integrity, consistency, etc. The principal component analysis (PCA) algorithm is used to perform dimensionality reduction on the data to extract the main features, and then the information entropy of the data is calculated p i is the probability distribution of the data features to evaluate the uncertainty of the data. When the data quality is lower than the set threshold, the data update and correction program is automatically started. The data quality monitoring period can be set between 1 day and 7 days, and the efficiency of data update and correction is not less than 92%, ensuring the reliability of project cost information.
[0068] In the present invention, in the data analysis and mining step, an outlier detection model for project cost data is established, which combines a statistics-based method and a clustering analysis method. Let the outlier judgment threshold be O t , and when the data deviates from the normal range by more than O t , it is determined as an outlier. The isolated forest algorithm is used to further analyze the outlier to determine the isolation degree of the outlier E(h(x)) is the average path length of the data point x in the isolated forest, and C(h(x)) is the average path length of the data point x in the completely random tree). The outlier is marked and analyzed to analyze the reasons for the outlier, such as data entry errors, abnormal market fluctuations, etc. The accuracy rate of outlier detection is not less than 96%, providing a basis for data cleaning and correction.
[0069] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A method for optimizing and integrating engineering cost information based on Internet big data services, characterized in that: The following steps are involved: S1. Data collection and integration steps: Collect project cost related data from multiple Internet data sources; Using intelligent crawler technology, accurately locate and collect data based on preset keywords and data features; Clean and integrate the collected data, remove duplicate data, erroneous data and abnormal data, and integrate data from different sources through a data matching algorithm based on semantic analysis. Set the data semantic similarity threshold as M t , when the data semantic similarity is greater than M t , M t =8, data fusion is performed to form a standardized basic information database of engineering cost; S2. Data analysis and mining steps: Use data mining algorithms to analyze the integrated project cost data, and use the Apriori association rule algorithm to mine the multi-dimensional correlation between building material prices and project geographical locations, seasonal factors, and project types. Set the minimum support as S min , S min =0.2, the minimum confidence level is C min , C min =0.6, mining out data item sets with strong association; The construction cost prediction model is constructed based on big data analysis. The neural network integrated with particle swarm optimization algorithm is used to construct the model. The input layer neurons of the neural network correspond to the engineering characteristic parameters, and the output layer is the engineering cost prediction value. The particle swarm optimization algorithm is used to optimize the weights and thresholds of the neural network. The particle swarm scale is N, N = 50, the inertia weight w decreases linearly from 0.9 to 0.4, and the learning factor c1 = c2 = 2; S3, information optimization and screening step: optimize and screen the project cost information according to user needs and project specific requirements. Users can set screening conditions, including project location, project budget range, project construction period, and project environmental protection requirements. The system screens out information that meets the requirements from the basic information database according to the set conditions; Prioritize the selected information, comprehensively consider the timeliness, reliability, authority of data sources, and the matching degree between data and user needs, and use a dynamic weighted sorting algorithm, setting the timeliness weight as W. t , the reliability weight is W r , the authority weight is W a , the matching weight is W m , where W t =0.25,W r =0.3,W a =0.2,W m =0.25; S4, integration and visualization step: integrating the optimized and screened engineering cost information to form a complete engineering cost information report; The project cost information is presented in the form of intuitive charts using visualization technology, including bar charts showing the proportion of different parts of the cost, line charts reflecting material price trends, and heat maps showing the cost differences in different regions. When the user hovers over a data point, the detailed information of the data point and related data information can be automatically displayed; Data update and verification steps: Establish a real-time data update mechanism and establish a long connection with the data source. When the data source has data updates, immediately obtain and process them. Use data fingerprint technology to verify the integrity of the updated data. Suppose the data fingerprint function is F(d), d is the data block, and compare the fingerprints of the data before and after the update. If the fingerprints match, perform the data update operation. Regularly verify the data in the information database, use the cross-validation method, divide the data into multiple groups, each group of data is used as a validation set and a training set, and calculate the data deviation rate by comparing the model prediction results with the actual data. P is the predicted value, A is the actual value, and when the deviation rate exceeds the set threshold of 5%, the data is re-collected and corrected; S5. Multi-source data fusion and calibration step: For data from different sources with differences, the Bayesian estimation method is used for fusion and calibration. Suppose the data from different sources are D1, D2, ..., D n , whose prior probabilities are P(D1), P(D2), …, P(D n ), the likelihood functions are L(D1), L(D2), …, L(D n ), the fused data D f Calculated according to the formula During the integration process, the credibility of the data is evaluated, and the credibility weight is determined based on the historical accuracy of the data source and the frequency of data update. The data with low credibility is marked and further verified; S6. Model self-learning and optimization steps: During the operation of the project cost prediction model, new data samples are automatically collected, and the incremental learning algorithm is used to update and optimize the model in real time. Suppose the new data sample set is S n , the model is based on S n The data in adjusts the neuron connection weights W ij , the adjustment formula is Where η is the learning rate, δ i is the error term, x j is the input feature, α is the momentum factor, is the amount of the last weight adjustment; The performance of the model is evaluated, and the confusion matrix and F1 value are used as evaluation indicators. When the F1 value is lower than the set threshold of 0.85, the model reconstruction mechanism is triggered to reselect the algorithm or adjust the model structure.
2. The method for optimizing and integrating engineering cost information based on Internet big data services according to claim 1 is characterized in that: In the data collection and integration steps, a multi-channel collection method is adopted for the collection of building materials price data, including data connection with major building materials e-commerce platforms, data capture from the government price department website, and collection of official quotations from building materials manufacturers. Blockchain technology is used to build a data traceability system, and the fusion weight of data from different channels is determined according to the reliability of the data source.
3. The method for optimizing and integrating construction cost information based on Internet big data services according to claim 1 is characterized in that: The engineering cost prediction model constructed in the data analysis and mining steps adopts a combination of the long short-term memory network LSTM in the deep learning algorithm and the random forest algorithm. LSTM is used to learn the time series characteristics of engineering cost data, and the random forest algorithm is used to classify and regress the data. The key features that have a greater impact on the engineering cost are screened out through the feature importance evaluation algorithm.
4. The method for optimizing and integrating construction cost information based on Internet big data services according to claim 1 is characterized in that: In the information optimization and screening steps, when the screening conditions set by the user are vague or incomplete, the system uses an intelligent recommendation algorithm to automatically supplement and improve the screening conditions based on the user's historical usage records and some characteristic information of the current project, using fuzzy matching algorithms and case-based reasoning methods.
5. The method for optimizing and integrating construction cost information based on Internet big data services according to claim 1 is characterized in that: In the integration and visualization steps, add data source annotations and data update timestamps in the engineering cost information report. The level of detail of the annotation information should be no less than 95%. Set interactive functions in the visualization charts. Users can obtain detailed information on data points in the charts by clicking the mouse, and export the visualization charts into high-definition images or vector graphics formats according to user needs.
6. The method for optimizing and integrating construction cost information based on Internet big data services according to claim 1 is characterized in that: It also includes data security and privacy protection steps. An encryption algorithm is used to encrypt, store and transmit the collected engineering cost data. The encryption algorithm uses the AES encryption algorithm with a key length of 256 bits. The user's personal information and operation records are anonymized. At the same time, a data access permission management system is established to assign different access permissions based on user roles and operation requirements.
7. The method for optimizing and integrating construction cost information based on Internet big data services according to claim 1 is characterized in that: Establish a data quality monitoring and evaluation mechanism to evaluate the data quality in the engineering cost information database, use the principal component analysis (PCA) algorithm to reduce the dimension of the data, extract features, and then calculate the information entropy of the data. p i The probability distribution of data features is used to evaluate the uncertainty of the data. When the data quality is lower than the set threshold, the data update and correction program is automatically started.
8. The method for optimizing and integrating construction cost information based on Internet big data services according to claim 1 is characterized in that: In the data analysis and mining step, an outlier detection model for engineering cost data is established, and a statistical method is combined with a cluster analysis method. The outlier judgment threshold is set as 0. t , when the data deviates from the normal range by more than O t When it is determined as an outlier, the isolation forest algorithm is used to analyze the outlier to determine the degree of isolation of the outlier E(h(x)) is the average path length of data point x in the isolation forest, and C(h(x)) is the average path length of data point x in the completely random tree. The outliers are marked and analyzed, and the causes of the outliers are analyzed.
Citation Information
Patent Citations
Cost allocation method based on power grid project spare parts, medium and system
CN117934042A
Engineering cost data calculation management system and method
CN119624559A
Engineering cost management method based on artificial intelligence
CN119648143A
Cited By
Elevator map generation method, system and equipment based on semantic recognition and image segmentation large model and medium
CN120894636A
Vehicle control method and system based on data acquisition of automobile data recorder
CN121062682A