A cost consulting data sorting and analysis system based on a cloud platform
By designing a cost consulting data sorting and analysis system based on cloud platform, the problems of data source heterogeneity, data volume explosion and data analysis strategies in the traditional cost consulting industry are solved, efficient and accurate data collection and analysis are achieved, and the reliability of analysis results is ensured.
Patent Information
- Application Number
- CN202510351098.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The traditional cost consulting industry faces the problems of data source heterogeneity, explosion of data volume, singularity of data analysis strategies and difficult to guarantee data accuracy.
Design a cost consulting data sorting and analysis system based on cloud platform, including data acquisition module, data clustering module, policy generation module, policy optimization module, execution engine module and data verification module. The system can collect data from multiple heterogeneous data sources in real time, perform format standardization processing, dynamic clustering, build multiple candidate analysis strategies, optimize strategies, perform data cleaning and association modeling, and perform multi-level verification of the results.
It effectively solves the problems of data source heterogeneity, data volume explosion and data analysis strategy singularity, improves the efficiency and accuracy of data acquisition and analysis, and ensures the accuracy and reliability of analysis results.
Smart Images

Figure CN119863286B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and analysis, and particularly to a cost consulting data sorting and analysis system based on a cloud platform. Background Art
[0002] In today's cost consulting industry, data processing and analysis face many challenges, and traditional cost data processing methods are difficult to meet the needs of industry development.
[0003] First of all, the heterogeneity of data sources has become a major problem. With the development of information technology, cost data sources are extensive, covering different formats of files (such as Excel, PDF, etc.), various databases (relational databases, non-relational databases), and numerous business system interfaces. These data sources vary greatly in data format, storage structure, and data standards, making it difficult to directly integrate and analyze the original cost data obtained from multiple sources. For example, cost data exported from architectural design software may be stored in a specific binary format, while cost data in the financial system is stored according to the table structure of a relational database, which results in complex and cumbersome data collection and preprocessing work, consuming a large amount of human and time costs.
[0004] Secondly, the explosive growth of data volume has brought huge pressure to data processing. With the continuous expansion of the scale and increasing number of construction projects, cost data has grown exponentially. Traditional single-machine processing methods are seriously insufficient in computing power and storage capacity when facing massive data. For example, for large-scale urban infrastructure construction projects, the cost data involved may contain thousands of detailed records. Single-machine processing is not only extremely slow but may even be unable to complete the data processing task due to insufficient memory, failing to meet the requirements of real-time and accuracy.
[0005] Furthermore, the singularity and staticity of data analysis strategies are difficult to adapt to complex and changeable business scenarios. In actual cost consulting work, different types of projects (such as residential, commercial, industrial, etc.) have different characteristics, and the distribution and laws of data are also different. Existing analysis strategies are often based on fixed rules and models, lacking adaptability to dynamic changes in data. For example, when analyzing the cost of a residential project, if the same analysis model as that of a commercial project is used, it may ignore the unique cost components in the residential project (such as the impact of housing type structure on cost), resulting in inaccurate analysis results and unable to provide effective support for decision-making.
[0006] In addition, it is difficult to guarantee the accuracy and reliability of data. During the processing of cost data, due to the complexity of data sources and the possibility of manual intervention, data quality problems occur frequently. Issues such as data entry errors, data missing, and data inconsistency will seriously affect the credibility of the analysis results. Moreover, the existing data verification mechanism is not perfect and cannot comprehensively and deeply detect logical errors and abnormal situations in the data. For example, when calculating the cost of building materials, if the unit price of the materials entered manually is incorrect and the system fails to detect and correct it in time, it will lead to deviations in the entire cost analysis results, bringing risks to project decisions. Summary of the Invention
[0007] The purpose of the present invention is to provide a cost consulting data sorting and analysis system based on a cloud platform to solve the problems raised in the above-mentioned background technology.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A cost consulting data sorting and analysis system based on a cloud platform, the system includes:
[0009] A data acquisition module, used to obtain original cost data from multiple heterogeneous data sources in real time and perform format standardization processing on the original cost data to generate a structured data set;
[0010] A data clustering module, used to dynamically cluster the structured data set according to preset classification dimensions to generate multiple data subsets and extract the key feature vectors of each data subset;
[0011] A strategy generation module, used to construct multiple candidate analysis strategies based on the key feature vectors of each data subset, and each candidate analysis strategy includes data cleaning rules, association models, and optimization parameter combinations;
[0012] A strategy optimization module, used to calculate the estimated computational complexity, estimated storage occupancy, and estimated response time of each candidate analysis strategy according to preset optimization goals, and screen out the preferred analysis strategies that meet the constraint conditions;
[0013] An execution engine module, used to perform data cleaning, association modeling, and parameter optimization operations on the corresponding data subsets according to the preferred analysis strategies to generate a standardized cost analysis result;
[0014] Among them, a mapping relationship between multiple types of standard data scenarios and candidate analysis strategies is pre-stored in the strategy generation module, and the optimization goals include the lowest computational complexity, the smallest storage occupancy, and the shortest response time.
[0015] Preferably, the data clustering module includes:
[0016] A dimension definition unit for defining classification dimensions according to cost item attributes, time series characteristics, and cost composition factors;
[0017] A clustering algorithm unit for using a hybrid clustering algorithm with adaptive weight allocation to calculate the clustering weight of each data point according to the following formula:
[0018] ;
[0019] Where is the clustering weight of the jth data point, represents the similarity between the data point and the project attribute, represents the continuity score of the data point in the time series, and are dynamic adjustment coefficients; n is the total number of valid data points in the current clustering cluster, represents the sum of the weighted comprehensive scores of all data points;
[0020] A subset generation unit for iteratively partitioning data points according to the clustering weight to generate the data subset.
[0021] Preferably, the strategy generation module includes:
[0022] A rule base unit storing an abnormal data detection rule set trained based on historical cost data;
[0023] A model base unit storing a linear regression model, a Bayesian network model, and a neural network model;
[0024] A strategy construction unit for matching candidate rules and models from the rule base unit and the model base unit according to the key feature vectors of the data subset and generating a parameter optimization space , and constructing the candidate analysis strategy.
[0025] Preferably, the formula for calculating the estimated computational complexity in the strategy optimization module is:
[0026] ;
[0027] Where represents the estimated computational complexity, represents the execution time of the th cleaning rule, represents the memory occupancy of the th association model, and are weight coefficients.
[0028] Preferably, the system further includes:
[0029] A data verification module for performing multi-level verification on the standardized cost analysis results, specifically including:
[0030] A consistency verification unit for eliminating result data beyond the threshold based on a preset cost index threshold range;
[0031] A logical verification unit for modeling the dependency relationships between cost parameters using a directed acyclic graph (DAG) and detecting logical conflicts in the result data;
[0032] A traceability unit for recording the generation path and data lineage of each analysis result.
[0033] Preferably, the modeling method of the dependency relationships in the logical verification unit includes:
[0034] Step A1: Extract the set of cost parameters ;
[0035] Step A2: Construct an edge set based on the conditional probability distribution between parameters analyzed from historical data , where is the probability threshold;
[0036] Step A3: Generate the directed acyclic graph according to and .
[0037] Preferably, the system further includes:
[0038] A dynamic adjustment module for monitoring changes in data distribution in real time during the data cleaning process and calculating the data distribution offset according to the following formula :
[0039] ;
[0040] where represents the occurrence frequency of the i-th data feature in the current data stream, represents the baseline frequency of the i-th data feature in the historical dataset, m is the total number of preset key features, is the preset threshold, and if then trigger the policy generation module to reconstruct the candidate analysis policy.
[0041] Preferably, the execution engine module includes:
[0042] A distributed computing unit for splitting the data subset into multiple data blocks and allocating them to the computing nodes of the cloud platform for parallel processing;
[0043] The fault-tolerant unit adopts a checkpoint mechanism based on redundant calculation, saves the intermediate state at preset time intervals, and resumes the calculation from the nearest checkpoint in case of node failure.
[0044] Preferably, the system further includes:
[0045] A knowledge graph module for constructing an entity relationship graph in the field of construction cost, specifically including:
[0046] An entity extraction unit for extracting construction cost-related entities and attributes from unstructured text;
[0047] A relationship reasoning unit for learning implicit association rules between entities using a graph neural network model;
[0048] A graph update unit for dynamically expanding entity nodes and relationship edges according to new data.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] The data acquisition module can obtain raw construction cost data from multiple heterogeneous data sources in real time, perform format standardization processing, and generate a structured data set. This function effectively solves the data integration problem caused by heterogeneous data sources in the traditional method, greatly improves the efficiency and accuracy of data acquisition. For example, whether it is data from design software, financial systems or other business platforms, it can be quickly and accurately collected and converted into a unified format, laying a solid foundation for subsequent analysis work, reducing the cumbersome work of manually processing data format differences, and reducing the error probability.
[0051] The data clustering module defines classification dimensions according to the attributes of construction cost projects, time series characteristics and cost composition factors, and uses a hybrid clustering algorithm with adaptive weight assignment for dynamic clustering to generate multiple data subsets and extract key feature vectors. This clustering method can more accurately reflect the characteristics of different types of construction cost data, making the subsequent analysis more targeted. Taking different types of construction projects as an example, through clustering, the data of residential, commercial and industrial projects can be classified separately, and the unique characteristics of each type of project can be analyzed, improving the accuracy and practicality of the analysis results.
[0052] The strategy generation module constructs multiple candidate analysis strategies based on the key feature vectors of the data subset, and pre-stores the mapping relationship between the standard data scenarios and the candidate analysis strategies, providing a rich variety of analysis solutions. The strategy optimization module filters out the optimal analysis strategies that meet the constraint conditions according to the preset optimization goals (the lowest computational complexity, the smallest storage occupancy rate, and the shortest response time). This not only improves the scientificity and rationality of the analysis strategies, but also optimizes the system performance under the resource constraints of the cloud platform. For example, when processing large-scale cost data, it can quickly select the optimal analysis strategy, reduce the waste of computing resources, and improve the operating efficiency of the system.
[0053] The data verification module conducts multi-level verification on the standardized cost analysis results, including consistency verification, logical verification, and traceability. The consistency verification eliminates abnormal data based on the preset threshold range of cost indicators. The logical verification uses a directed acyclic graph to model the dependency relationships between various cost parameters to detect logical conflicts. The traceability unit records the generation path and lineage of the data. These measures comprehensively ensure the accuracy and reliability of the analysis results. For example, when abnormal analysis results are found, the problem can be quickly located through traceability, errors can be corrected in a timely manner, providing reliable data support for cost consulting and reducing decision-making risks.
[0054] The dynamic adjustment module monitors the changes in data distribution in real time during the data cleaning process. When the data distribution offset exceeds the preset threshold, it triggers the strategy generation module to reconstruct the candidate analysis strategies. This enables the system to adapt to the dynamic changes of the data and always maintain the accuracy of the analysis results. For example, when there are large fluctuations in the prices of building materials in the market, resulting in changes in the cost data distribution, the system can adjust the analysis strategy in a timely manner to reflect the latest market conditions and provide timely and effective suggestions for project cost control.
[0055] The distributed computing unit of the execution engine module divides the data subset into multiple data blocks and distributes them to the computing nodes of the cloud platform for parallel processing, improving the data processing speed. The fault tolerance unit adopts a checkpoint mechanism based on redundant computing to save the intermediate state at preset time intervals and resume the calculation from the nearest checkpoint when a node fails, ensuring the continuity of data processing. This is crucial for processing large-scale and long-term cost data calculation tasks, ensuring the stable operation of the system in a complex computing environment.
[0056] The knowledge graph module constructs an entity relationship graph in the cost field. Through entity extraction, relationship reasoning, and graph update, it presents various cost-related information in the form of a graph. This helps users more intuitively and deeply understand the internal relationships between cost data, discover potential knowledge and rules, provides a more comprehensive and in-depth analysis perspective for cost consulting, and improves the intelligent level of the industry. Description of the Drawings
[0057] Figure 1 This is the working principle diagram of the cost consulting data sorting and analysis system based on the cloud platform according to the present invention;
[0058] Figure 2 This is the working flow chart of the data clustering module;
[0059] Figure 3 This is the working flow chart of the data verification module;
[0060] Figure 4 This is the working principle diagram of the execution engine module. Specific implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0062] Please refer to Figures 1-4 , the present invention provides a technical solution: a cost consulting data sorting and analysis system based on the cloud platform, and the system includes:
[0063] Data acquisition module: Real-time obtain the original cost data from multiple heterogeneous data sources, and these data sources may include files in different formats, databases, and various business system interfaces, etc. After obtaining the data, perform format standardization processing on the original cost data, convert it into a unified format, and finally generate a structured data set for subsequent analysis and processing.
[0064] Data clustering module: Dynamically cluster the structured data set according to the preset classification dimensions. Through the clustering operation, generate multiple data subsets, and then extract the key feature vectors of each data subset for subsequent construction of targeted analysis strategies.
[0065] Strategy generation module: Based on the key feature vectors of each data subset, construct multiple candidate analysis strategies. Each candidate analysis strategy includes data cleaning rules, association models, and optimization parameter combinations, providing multiple optional solutions for subsequent data processing and analysis.
[0066] Strategy optimization module: According to the preset optimization objectives, calculate the estimated computational complexity, estimated storage occupancy, and estimated response time of each candidate analysis strategy. Through these calculations, screen out the preferred analysis strategies that meet the constraint conditions to improve the operation efficiency and performance of the system.
[0067] Execution Engine Module: Perform data cleaning, association modeling, and parameter optimization operations on the corresponding data subset according to the preferred analysis strategy, and finally generate a standardized cost analysis result to provide strong data support for cost consulting.
[0068] Among them, the strategy generation module pre-stores the mapping relationships between multiple types of standard data scenarios and candidate analysis strategies, and the optimization objectives include the lowest computational complexity, the smallest storage occupancy, and the shortest response time.
[0069] The present invention will be further described below in conjunction with Embodiments 1 to 5:
[0070] Embodiment 1:
[0071] The data clustering module includes a dimension definition unit, a clustering algorithm unit, and a subset generation unit. The dimension definition unit defines classification dimensions according to the cost project attributes, time series characteristics, and cost composition factors. In practical applications, the cost project attributes cover project types (such as residential, commercial, industrial, etc.), project scales (building area, investment amount, etc.); the time series characteristics involve the project approval time, construction period, completion time, etc.; the cost composition factors include labor costs, material costs, equipment costs, etc. By comprehensively considering these factors to determine the classification dimensions, the clustering results can better meet the actual business needs.
[0072] The clustering algorithm unit adopts a hybrid clustering algorithm with adaptive weight assignment. When calculating the clustering weight of each data point, the formula
[0073] ;
[0074] Among them, is the clustering weight of the j-th data point, represents the similarity between the data point and the project attributes. For example, for a data point of a residential project, if its project type, scale, and other attributes are highly similar to those of other data points in a certain clustering cluster, then has a larger value; represents the continuity score of the data point in the time series. If the data point is closely related to the previous and subsequent data points in time and has good continuity, then has a higher value. and are dynamically adjusted coefficients. According to different business scenarios and data characteristics, the values of these two coefficients can be dynamically adjusted to balance the influence of project attribute similarity and time series continuity in the calculation of clustering weights. n is the total number of valid data points in the current clustering cluster, represents the sum of the weighted comprehensive scores of all data points.
[0075] The subset generation unit iteratively partitions the data points according to the clustering weights. First, the clustering clusters are initialized, and the data points are assigned to each clustering cluster according to the initial clustering weights. Then, the clustering weights of each data point are continuously calculated, and the belonging of the data points is readjusted according to the weights until the clustering result is stable and no obvious changes occur, and finally a data subset is generated.
[0076] The key feature vector is extracted from the data subset generated by clustering, and is a multi-dimensional vector that can characterize the core business attributes of the subset. Its dimension is jointly determined by the preset classification dimensions (such as cost item attributes, time series features, cost composition factors) and the calculation results of the dynamic weights in the clustering algorithm. It specifically includes the following elements:
[0077] Project attribute features: such as project type (residential, commercial, industrial, etc.), project scale (building area, investment amount range), etc.
[0078] Time series features: such as time continuity score (based on the tightness of the distribution of data points on the time axis), construction period length, completion time deviation, etc.
[0079] Cost composition features: such as the proportion of labor cost, the fluctuation coefficient of material cost, the ratio of equipment cost to total cost, etc.
[0080] Clustering weight features: based on the similarity calculated in the clustering algorithm and the time continuity score of the weighted result , reflecting the comprehensive importance of data points in the clustering process.
[0081] The steps for obtaining the key feature vector include:
[0082] Step 1: Dynamically cluster the structured data set through the data clustering module to generate multiple data subsets.
[0083] Step 2: Conduct statistical analysis on the data points in each data subset to extract specific indicators of the preset dimensions:
[0084] Project attribute features: Statistically analyze the distribution ratio of project types in the subset (such as the proportion of residential category is 100%), the mean or range of project scale (such as the mean building area is 128,000 square meters).
[0085] Time series features: Calculate the time continuity score , for example, evaluate the continuity by the standard deviation of the construction time interval between adjacent data points (the smaller the standard deviation, the higher the score).
[0086] Cost composition characteristics: Calculate the average value and proportion of each cost factor. For example, the average labor cost proportion is 35%, and the material cost fluctuation coefficient (standard deviation / mean) is 0.15.
[0087] Clustering weight feature: extract the weight value of each data point in the clustering algorithm , and calculate the mean of the weights in the subset (such as 0.85).
[0088] Step 3: Combine the above statistical results into a multi-dimensional vector as the key feature vector of the data subset.
[0089] Embodiment 2:
[0090] This embodiment describes in detail the workflow of the strategy generation module, which generates diversified candidate analysis strategies for different data subsets based on historical data and multiple models, thereby improving the accuracy and adaptability of the analysis strategies.
[0091] The strategy generation module includes a rule base unit, a model base unit, and a strategy construction unit. The rule base unit stores a set of abnormal data detection rules trained based on historical cost data. These rules are obtained through analysis and learning of a large amount of historical cost data. For example, through statistical analysis of various cost data such as labor costs and material costs in historical data, the reasonable fluctuation range of various costs under a specific project type and scale is determined. When new data exceeds this range, it can be judged as abnormal data. These rules can effectively identify errors, duplications, or unreasonable data in the data, providing a basis for subsequent data cleaning.
[0092] The model library unit stores linear regression models, Bayesian network models and neural network models. The linear regression model is suitable for analyzing the linear relationship between variables in the cost data, such as predicting the construction cost based on factors such as building area and number of floors; the Bayesian network model can handle uncertainty and complex dependencies between variables, and has advantages in analyzing cost risks; the neural network model is good at processing nonlinear and complex data patterns, and can be used to mine deep-level cost data features.
[0093] The strategy building unit matches candidate rules and models from the rule base unit and the model base unit according to the key feature vector of the data subset. For example, for a data subset containing a large amount of time series data and data features showing a certain linear trend, the strategy building unit may select a linear regression model as the association model and select rules related to time series data anomaly detection from the rule base unit. At the same time, the parameter optimization space is generated. . Taking the linear regression model as an example, the parameter optimization space may include the value range of regression coefficients, regularization parameters, etc. By constructing the parameter optimization space, the parameters of the model can be adjusted and optimized to improve the performance of the model and the accuracy of the analysis results, and finally a candidate analysis strategy can be constructed.
[0094] Example 3:
[0095] This example mainly illustrates the calculation method of the strategy optimization module and the data verification function of the system, which is used to ensure that the selected analysis strategy is better in performance and at the same time ensure the accuracy and reliability of the generated standardized cost analysis results.
[0096] In the strategy optimization module, the formula for calculating the estimated computational complexity is
[0097] ;
[0098] Among them, represents the estimated computational complexity, represents the execution time of the th cleaning rule. For example, when cleaning a certain data subset, the time required for the rule to remove duplicate data to be executed once; represents the memory occupancy of the th association model, such as the memory size occupied by a linear regression model during operation. and are weight coefficients. According to the emphasis of the system on computational time and memory occupancy, the values of these two weight coefficients can be adjusted. Through this formula, the computational complexity of each candidate analysis strategy can be quantitatively evaluated, providing a basis for screening and optimizing analysis strategies.
[0099] The system also includes a data verification module for multi-level verification of the standardized cost analysis results. The consistency verification unit eliminates the result data that exceeds the threshold based on the preset cost index threshold range. For example, in construction project cost, there is usually a reasonable range for the cost per square meter, which is determined according to factors such as different regions and project types. If the cost per square meter of a certain project in the analysis results exceeds the preset threshold range, the consistency verification unit will eliminate this part of the data to ensure the numerical rationality of the analysis results.
[0100] The logical verification unit uses a directed acyclic graph (DAG) to model the dependency relationships between each cost parameter and detects the logical conflicts in the result data. During the modeling process, first extract the set of cost parameters , and these parameters may include building area, number of floors, labor cost, material cost, etc. Then, based on historical data analysis of the conditional probability distribution between the parameters, construct the edge set , where is the probability threshold. For example, if historical data analysis reveals that when the building area increases, the probability of the material cost increasing exceeds a certain threshold (such as 0.8), then an edge is added to the edge set from the building area parameter node to the material cost parameter node. Finally, based on and a directed acyclic graph is generated. Through this directed acyclic graph, the dependency relationships between cost parameters can be clearly shown. When data that does not conform to this dependency relationship appears in the analysis results, a logical conflict can be judged, and thus the result data can be corrected or eliminated.
[0101] The traceability unit records the generation path of each analysis result and the data lineage. This means that every step in the entire process from the collection of raw data to data processing, analysis, and finally to the generation of analysis results, as well as the data sources used, are detailedly recorded. For example, when a problem is found in a certain analysis result, the traceability unit can be used to trace back to which step of data processing had a deviation, or whether there was an error in the raw data itself, facilitating problem troubleshooting and result correction.
[0102] Example 4:
[0103] This example focuses on introducing the dynamic adjustment mechanism of the system and the fault tolerance function of the execution engine module. Its unique role is to enable the system to adapt to the dynamic changes of data, while ensuring the continuity and accuracy of data processing when node failures occur during the calculation process.
[0104] The system includes a dynamic adjustment module that monitors the change in data distribution in real time during the data cleaning process. Calculate the data distribution offset according to the following formula :
[0105] ;
[0106] where represents the occurrence frequency of the i-th data feature in the current data stream, represents the reference frequency of the i-th data feature in the historical dataset, is the total number of preset key features. For example, in a cost dataset, the key features include the prices of different materials, man-hours, etc. By counting the occurrence frequencies of these key features in the current data stream and comparing them with the reference frequencies in the historical dataset, the data distribution offset is calculated. If ( is the preset threshold), then the strategy generation module is triggered to reconstruct the candidate analysis strategy. This enables the system to timely detect changes in data distribution and adjust the analysis strategy according to the changes, ensuring the accuracy of the analysis results.
[0107] The execution engine module includes a fault-tolerant unit that adopts a checkpoint mechanism based on redundant computing. During the operation of the system, the fault-tolerant unit saves the intermediate state at preset time intervals. For example, every certain period (such as 10 minutes), it saves the intermediate state information such as the processing progress of the currently processed data subset and the completed calculation results. When a node failure occurs, the system can resume the calculation from the nearest checkpoint instead of starting from scratch. This can greatly reduce the waste of calculation time caused by node failures, ensure the continuity and accuracy of data processing, and improve the reliability and stability of the system.
[0108] Embodiment 5:
[0109] The knowledge graph module includes an entity extraction unit, a relationship reasoning unit, and a graph update unit. The entity extraction unit extracts cost-related entities and attributes from unstructured text. The unstructured text may come from project documents, contract documents, industry reports, etc. For example, it extracts entities such as "project name", "construction unit", "construction company", etc., and attributes such as "project scale", "cost amount" from project documents. Through natural language processing techniques, such as named entity recognition algorithms, the text is analyzed and processed to accurately identify these entities and attributes.
[0110] The relationship reasoning unit uses a graph neural network model to learn the implicit association rules between entities. Based on the constructed entity set, the relationship between entities is learned and inferred through the graph neural network model. For example, by analyzing a large amount of project data and related texts, the graph neural network model can find that there may be a cooperation relationship between the "construction unit" and the "construction company", and there is an association relationship between the "project name" and the "cost amount", etc. These implicit association rules are obtained through model learning and are used to construct the entity relationship graph.
[0111] The graph update unit dynamically expands entity nodes and relationship edges according to new data. When new cost data is added to the system, the graph update unit first determines whether the new data contains new entities or relationships. If there are new entities, such as a new project participant, a new entity node will be added to the graph; if there is a new relationship, such as a business transaction relationship emerging between two previously unassociated entities, the corresponding relationship edge will be added to the graph. In this way, the knowledge graph can be continuously updated and improved as new data is added, always maintaining the latest reflection of knowledge in the cost field and providing users with more accurate and comprehensive knowledge support.
[0112] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus.
[0113] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A cost consulting data collation and analysis system based on a cloud platform, characterized in that: The system comprises: A data acquisition module is used to obtain original cost data from multiple heterogeneous data sources in real time, and perform format standardization processing on the original cost data to generate a structured data set; A data clustering module, used to dynamically cluster the structured data set according to a preset classification dimension, generate multiple data subsets, and extract key feature vectors of each data subset; A strategy generation module is used to construct multiple candidate analysis strategies based on the key feature vectors of each data subset. Each candidate analysis strategy includes data cleaning rules, association models, and optimization parameter combinations. The strategy optimization module is used to calculate the estimated computational complexity, estimated storage occupancy rate and estimated response time of each candidate analysis strategy according to the preset optimization goal, and select the preferred analysis strategy that meets the constraint conditions; An execution engine module, used to perform data cleaning, association modeling and parameter optimization operations on the corresponding data subset according to the preferred analysis strategy to generate a standardized cost analysis result; The strategy generation module pre-stores mapping relationships between multiple types of standard data scenarios and candidate analysis strategies, and the optimization objectives include minimum computational complexity, minimum storage occupancy, and shortest response time; The system further comprises: The data verification module is used to perform multi-level verification on the standardized cost analysis results, specifically including: The consistency check unit removes the result data exceeding the threshold value based on the preset cost index threshold value range; The logic verification unit uses a directed acyclic graph (DAG) to model the dependency relationship between various cost parameters and detect logical conflicts in the result data; The traceability unit records the generation path and data lineage of each analysis result; The method for modeling the dependency relationship in the logic verification unit includes: Step A1: Extract cost parameter set ; Step A2: Analyze the conditional probability distribution between parameters based on historical data , construct edge sets ,in is the probability threshold; Step A3, according to and Generating the directed acyclic graph; The system further comprises: The dynamic adjustment module is used to monitor data distribution changes in real time during data cleaning and calculate the data distribution offset according to the following formula: : ; in, Indicates the frequency of occurrence of the i-th data feature in the current data stream, represents the benchmark frequency of the ith data feature in the historical data set, m is the total number of preset key features, is the preset threshold, if The strategy generation module is then triggered to reconstruct the candidate analysis strategy; The execution engine module includes: Distributed computing unit, used to split the data subset into multiple data blocks and distribute them to computing nodes of the cloud platform for parallel processing; The fault-tolerant unit adopts a checkpoint mechanism based on redundant calculations to save intermediate states at preset time intervals and resume calculations from the most recent checkpoint when a node fails.
2. According to the cloud platform-based cost consulting data collation and analysis system of claim 1, it is characterized in that: The data clustering module comprises: Dimension definition unit, used to define classification dimensions according to cost item attributes, time series characteristics and cost component factors; The clustering algorithm unit is used to use a hybrid clustering algorithm with adaptive weight allocation to calculate the clustering weight of each data point according to the following formula: ; in, is the clustering weight of the jth data point, Indicates the similarity between the data point and the item attribute, Represents the continuity score of the data point in the time series, and is the dynamic adjustment coefficient; n is the total number of valid data points in the current cluster, represents the sum of weighted comprehensive scores of all data points; The subset generation unit is used to iteratively divide the data points according to the clustering weights to generate the data subsets.
3. According to the cloud platform-based cost consulting data collation and analysis system of claim 2, it is characterized in that: The strategy generation module includes: A rule base unit stores a set of abnormal data detection rules obtained through training based on historical cost data; A model library unit stores linear regression models, Bayesian network models and neural network models; A strategy building unit, used to match candidate rules and models from the rule base unit and the model base unit according to the key feature vector of the data subset, and generate a parameter optimization space , construct the candidate analysis strategy.
4. According to the cloud platform-based cost consulting data collation and analysis system of claim 3, it is characterized in that: The formula for calculating the estimated computational complexity in the strategy optimization module is: ; in, Indicates the estimated computational complexity, Indicates The execution time of each cleaning rule is Indicates The memory usage of the associated model, and is the weight coefficient.
5. According to the cloud platform-based cost consulting data collation and analysis system of claim 1, it is characterized in that: The system further comprises: The knowledge graph module is used to construct an entity relationship graph in the cost field, including: Entity extraction unit, extracting cost-related entities and attributes from unstructured text; The relational reasoning unit uses a graph neural network model to learn implicit association rules between entities; The graph update unit dynamically expands entity nodes and relationship edges based on new data.
Citation Information
Patent Citations
Cost consultation data arrangement and analysis method and system based on cloud platform
CN118798939A
Abnormal data identification and cleaning method and system
CN119272016A