Engineering cost big data management and analysis system

By employing multi-dimensional feature extraction and a hybrid storage architecture, combined with a spatiotemporal anchor diffusion filling strategy, the problems of low retrieval efficiency and insufficient semantic understanding in engineering cost data management are solved. This enables the identification and intelligent analysis of deep-level relationships, thereby improving the accuracy and efficiency of data management.

CN120873058AActive Publication Date: 2025-10-31GUANGZHOU ZHUJIAN ENG COST CONSULTING CO LTD

Patent Information

Application Number
CN202510941244.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-31
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing technologies in engineering cost data management suffer from low retrieval efficiency, insufficient semantic understanding capabilities, difficulty in identifying deep-seated relationships between different cost projects, and inadequate intelligent data analysis and similarity recognition capabilities.

Method used

By extracting and vectorizing multidimensional features, adopting a hybrid storage architecture and multidimensional index structure, and combining a spatiotemporal anchor diffusion filling strategy, semantic recognition and similarity calculation of cost data are achieved, thereby improving the accuracy and intelligence of data management.

Benefits of technology

It enables semantic association recognition and large-scale data retrieval of cost data, improves the integrity and usability of the dataset, identifies deep-seated relationships between projects, and enhances data retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873058A_ABST
    Figure CN120873058A_ABST
Patent Text Reader

Abstract

The invention provides a project cost big data management and analysis system, and relates to the technical field of data management, and the system comprises a data collection and preprocessing module which is used for collecting original cost data from a heterogeneous data source, and carrying out the preprocessing of the original cost data, and obtaining the preprocessed cost data; the semantic feature extraction module is used for converting the preprocessed cost data into a multi-dimensional feature vector based on a multi-level feature extraction system; the similarity calculation module is used for calculating similarities among different cost data based on the multi-dimensional feature vectors to obtain a similarity matrix; and the data storage and management module is used for storing the cost data, the multi-dimensional feature vector and the similarity matrix by adopting a mixed storage architecture, and providing retrieval and recommendation functions of cost projects based on a multi-level feature space index structure. According to the method, the limitation that a traditional method only depends on keyword matching is solved, and the system can recognize the deep incidence relation between the items.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and in particular to a big data management and analysis system for engineering cost. Background Technology

[0002] Construction cost data, as a crucial resource in the construction industry, is characterized by its large volume, diverse sources, and complex structure. With the rapid development and increasing informatization of the construction industry, construction cost data exhibits multi-dimensional and multi-layered characteristics, encompassing various attributes such as basic project information, cost composition, technical parameters, geographical location, and time stamps. This data not only differs in numerical characteristics but also demonstrates high heterogeneity in text descriptions, classification coding, and units of measurement. Furthermore, the data exhibits significant timeliness and regional characteristics, posing technical challenges to unified data management and intelligent analysis.

[0003] Currently, engineering cost data management technology mainly adopts the traditional storage method based on relational databases, classifying and retrieving data through keyword matching and simple numerical comparisons. In terms of data quality control, existing technologies primarily rely on pre-set rule templates for data verification, ensuring data quality through methods such as field integrity checks, numerical range verification, and logical consistency verification. Regarding data storage and retrieval, traditional methods are based on relational database indexing technology, using traditional index structures such as B+ trees to support data queries. However, these traditional methods suffer from low retrieval efficiency and insufficient semantic understanding capabilities when handling large-scale, multi-dimensional feature data, making it difficult to accurately identify deep-seated relationships between different cost projects.

[0004] Chinese invention patent CN114490604B discloses a method for managing engineering cost data, which includes four core steps: data collection, data verification and storage, data classification and processing, and visualization. Its technical feature lies in establishing a five-rule verification system, including template fusion rules, integrity rules, cross-reference rules, interval verification rules, and trend verification rules. It determines the data processing method through a comparison mechanism between the "difference of the first type of data" and the "probability of the occurrence of the nth type of data," and uses a node partitioning method to store cost data in the corresponding database. However, this method has certain limitations in its technical implementation: its data classification is mainly based on preset data type identifiers for mechanical classification, with relatively simple classification criteria and a lack of in-depth mining of the inherent semantic relationships of the data; its data retrieval mechanism is not perfect, making it difficult for users to quickly locate relevant historical data according to complex business needs; although this method solves the data quality control problem, it is insufficient in terms of intelligent data analysis and similarity recognition capabilities, and its main functions are limited to standardized data storage and visualization, failing to meet the needs of modern cost management for in-depth data analysis and intelligent recommendation. Summary of the Invention

[0005] In view of this, the present invention provides an engineering cost big data management and analysis system. This system realizes semantic recognition and similarity calculation of cost data through multi-dimensional feature extraction and vectorization processing; improves the storage and retrieval efficiency of large-scale cost data through a hybrid storage architecture and multi-dimensional index structure; and enhances the accuracy and intelligence level of cost data management through data quality control and adaptive weight adjustment.

[0006] The technical solution of this invention is implemented as follows: This invention provides a big data management and analysis system for engineering cost, comprising: The data acquisition and preprocessing module is used to acquire raw cost data from heterogeneous data sources and preprocess the raw cost data to obtain preprocessed cost data. The semantic feature extraction module is used to convert the preprocessed cost data into multi-dimensional feature vectors based on a multi-level feature extraction system. The similarity calculation module is used to calculate the similarity between different cost data based on multi-dimensional feature vectors, and obtain a similarity matrix; The data storage and management module is used to store cost data, multi-dimensional feature vectors, and similarity matrices using a hybrid storage architecture, and provides retrieval and recommendation functions for cost projects based on a multi-level feature space index structure.

[0007] Preferably, the data acquisition and preprocessing module includes: The data acquisition unit is used to collect raw cost data from the project management system, cost software database, and bidding platform via API interface, database connection, and file import, and to add a timestamp to each cost record. The data preprocessing unit is used to perform format conversion, data cleaning, and missing value imputation on the original cost data to obtain preprocessed cost data.

[0008] Preferably, the data preprocessing unit uses a spatiotemporal anchor diffusion imputation strategy to impute missing values, and the process is as follows: For the cost data to be filled Anchor point sets are identified from historical cost data based on anchor point selection criteria. ; Each anchor point Cost data to be filled Intensity of influence between for:

[0009] in, For standardization With anchor point Time difference, For standardization With anchor point Spatial distance, anchor point Data credibility weight, The spatiotemporal coupling coefficient is determined through cross-validation of historical cost data; against Missing fields This will affect the intensity. As anchor point The fill weights are calculated using a weighted average. fill value :

[0010] in anchor point Corresponding missing fields The known value, For the set of anchor points The number of anchor points; Based on the calculated fill-in values ​​for each missing field, the filled-in cost data is generated. .

[0011] Preferably, the anchor point selection criteria include: time window conditions. Spatial range conditions, Integrity condition, anchor point exist The missing field has a complete value; where for timestamp, anchor point timestamp, As a time threshold, for With anchor point Geographical distance between them This is the spatial threshold.

[0012] Preferably, the semantic feature extraction module includes: The feature extraction unit is used to extract the original feature vector from the preprocessed cost data based on a four-level feature extraction system, including the project attribute feature layer, cost structure feature layer, technical process feature layer, and text semantic feature layer. The standardization unit is used to standardize the original feature vectors using an adaptive standardization strategy. Z-score standardization is used for normally distributed features, and quantile standardization is used for skewed features. This combines the feature vectors from each layer into a unified multidimensional feature vector representation. .

[0013] Preferably, the four-level feature extraction system includes: The project attribute feature layer extracts project type, building area, number of floors, and structural form from basic project information. The project attribute feature vector is represented as follows: ,in This represents the total number of dimensions for the project attributes. The cost structure feature layer calculates the proportion of each cost item to the total cost. The cost structure feature vector is represented as follows: ,in For the first Item cost amount, The total cost of the project; The technical process feature layer identifies technical keywords through engineering technology dictionary matching and calculates weight values. The technical feature vector is represented as follows: ,in For the first The weight value of each technical keyword; The text semantic feature layer extracts semantic features from the project description text based on a domain dictionary and context analysis. The text semantic feature vector is represented as follows: ,in denoted as the number of dimensions of the semantic features.

[0014] Preferably, the execution process of the standardized unit includes: The distribution characteristics of each feature layer are determined by the Shapiro-Wilk normality test, and each feature layer is identified as having normal or skewed distribution characteristics. The Z-score is used to standardize the characteristics of the normal distribution. The calculation formula is as follows:

[0015] in These are the original eigenvalues. The characteristic mean, Standard deviation; The quantile standardization is applied to the skewed distribution characteristics, and the calculation formula is as follows:

[0016] in and These are the minimum and maximum values ​​of the eigenvalues, respectively. Combine the feature vectors from each dimension into a complete multidimensional feature vector. All standardized feature values ​​are mapped to the interval [0,1].

[0017] Preferably, the similarity calculation module constructs a similarity matrix using a pairwise comparison strategy, and the process is as follows: For a dataset containing N cost data points Each cost data Corresponding to a multidimensional feature vector ; Iterate through all cost data using a double loop. ,in, and : Cost data is calculated based on similarity algorithms. Similarity values ​​between:

[0018] in, For cost data Similarity value between them and These represent the cost data respectively. and In the Normalized feature values ​​on each feature dimension For the first Weight coefficients for each feature dimension The total number of feature dimensions. and These represent the cost data respectively. and timestamp, For feature distance sensitivity parameters, Control parameters affected by time; After calculation, all similarity values ​​are indexed by row and column. Stored in In the similarity matrix.

[0019] Preferably, a multi-level feature space index structure is used to provide retrieval and recommendation functions for cost projects, including: Spatial index nodes are constructed by KD-tree segmentation of multidimensional feature vectors. Each index node records the feature space range it covers and the corresponding list of cost project identifiers. At the same time, a mapping relationship table between index nodes and row and column coordinates of similarity matrix is ​​established. When searching for cost projects, the target index node is located in the KD-tree according to the query feature vector, and the candidate project identifiers covered by the node and its neighboring nodes are obtained. The candidate project identifiers are converted into coordinate positions in the similarity matrix through the mapping relationship table. The corresponding similarity values ​​are read in batches from the similarity matrix and sorted in descending order of similarity value to generate a recommended project list.

[0020] Preferably, the hybrid storage architecture of the data storage and management module includes: a relational database for storing preprocessed cost data and a mapping table between index nodes and row and column coordinates of the similarity matrix; and a vector database for storing multidimensional feature vectors, similarity matrices, and multi-level feature space index structures.

[0021] The present invention has the following advantages over the prior art: (1) This invention achieves semantic association recognition and large-scale data retrieval of cost data through the collaborative work of multi-dimensional feature extraction, similarity calculation and hybrid storage architecture. The system converts cost data into multi-dimensional feature vectors for similarity calculation, which solves the limitation of traditional methods that rely solely on keyword matching, and enables the system to identify deep-level relationships between projects; (2) The data preprocessing module uses a spatiotemporal anchor diffusion filling strategy, combining time difference, spatial distance, and data credibility weights to calculate the influence intensity, thereby achieving intelligent filling of missing data. This method utilizes the spatiotemporal distribution patterns of engineering cost data, avoiding data bias that may occur with simple statistical filling methods, and improving the completeness and usability of the dataset; (3) The semantic feature extraction module establishes a four-level feature extraction system, extracting features from four dimensions: project attributes, cost structure, technology and text semantics, and processing feature data of different types and dimensions through an adaptive standardization strategy. This method enables the system to comprehensively capture the multi-dimensional feature information of cost data, providing a unified data foundation for subsequent similarity calculation; (4) The data storage module adopts a hybrid storage architecture, storing cost data in a relational database and feature vectors in a vector database, and constructing a multi-level feature space index through a KD-tree. This architecture fully utilizes the performance advantages of different database types, reduces the amount of data for similarity calculation through spatial partitioning and index mapping, and improves the retrieval performance of large-scale cost data. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the system framework of the present invention; Figure 2 This is a schematic diagram illustrating the technical implementation of the present invention; Figure 3 This is a schematic diagram of the filling strategy of the present invention. Detailed Implementation

[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0025] like Figure 1 and Figure 2 As shown, the present invention provides a big data management and analysis system for engineering cost, comprising: The data acquisition and preprocessing module is used to acquire raw cost data from heterogeneous data sources and preprocess the raw cost data to obtain preprocessed cost data. The semantic feature extraction module is used to convert the preprocessed cost data into multi-dimensional feature vectors based on a multi-level feature extraction system. The similarity calculation module is used to calculate the similarity between different cost data based on multi-dimensional feature vectors, and obtain a similarity matrix; The data storage and management module is used to store cost data, multi-dimensional feature vectors, and similarity matrices using a hybrid storage architecture, and provides retrieval and recommendation functions for cost projects based on a multi-level feature space index structure.

[0026] Specifically, in one embodiment of the present invention, the data acquisition and preprocessing module includes: The data acquisition unit is used to collect raw cost data from the project management system, cost software database, and bidding platform via API interface, database connection, and file import, and adds a timestamp to each cost record.

[0027] In this embodiment, corresponding access methods are adopted based on the characteristics of different data sources: the API interface method is suitable for modern systems that provide standard APIs, and data is pulled periodically through REST or SOAP interfaces; the database connection method is suitable for traditional database systems, and data is synchronized directly through JDBC or ODBC; the file import method is suitable for systems that provide data in the form of Excel, CSV, etc., and imports data through file parsing. During the data acquisition process, the system automatically adds a timestamp to each cost record, recording the data generation time, acquisition time, and processing time. The timestamp uses the UTC standard format and is accurate to the millisecond level.

[0028] The data preprocessing unit is used to perform format conversion, data cleaning, and missing value imputation on the original cost data to obtain preprocessed cost data.

[0029] In this embodiment, the preprocessing mainly includes: Format conversion: The system establishes a standard format definition for engineering cost data and processes heterogeneous data from different data sources through a format conversion engine, including field name mapping, data type conversion, and character encoding unification, converting text data with different encoding formats into UTF-8 encoding.

[0030] Data cleaning mainly involves identifying and processing outliers, duplicates, and formatting errors in the data; removing obvious entry errors such as negative quantities or unit prices exceeding reasonable ranges; identifying and merging duplicate records; and correcting formatting errors such as inconsistent date formats or non-numeric characters in numerical values.

[0031] Standardization processing primarily targets and normalizes numerical and categorical data. A unit conversion matrix is ​​established to convert various length, area, weight, and currency units into unified standard units. Unit identification is performed through regular expression matching and contextual analysis. For numerical values ​​lacking unit information, unit inference is made based on business semantics and numerical range. Simultaneously, a standard system and mapping table for project classification coding, material coding, and process coding are established, and code conversion is processed using precise matching and string similarity algorithms.

[0032] like Figure 3 As shown, the data preprocessing unit employs a spatiotemporal anchor point diffusion imputation strategy to fill in missing values. Traditional missing value imputation methods mainly rely on simple statistical methods, such as mean imputation, median imputation, or correlation-based linear interpolation. These methods only consider the statistical characteristics of the data, ignoring the spatial and geographical characteristics and temporal evolution patterns of engineering cost data. The spatiotemporal anchor point diffusion imputation strategy proposed in this invention fully considers the spatiotemporal distribution characteristics of engineering cost data. By establishing a spatiotemporal coupling influence intensity model, it achieves accurate imputation of missing data, exhibiting higher accuracy and rationality. The core idea of ​​this strategy is to find historical data similar to the data to be imputed in both time and space dimensions as anchor points. By calculating the spatiotemporal distance and influence intensity between the anchor points and the data to be imputed, a weighted average is used to calculate the imputation value, ensuring both the rationality of the imputation result and reflecting the spatiotemporal correlation characteristics of the cost data. The process is as follows: For the cost data to be filled Anchor point sets are identified from historical cost data based on anchor point selection criteria. ; Each anchor point Cost data to be filled Intensity of influence between for:

[0033] in, For standardization With anchor point Time difference, For standardization With anchor point Spatial distance, anchor point Data credibility weight, The spatiotemporal coupling coefficient is determined through cross-validation of historical cost data; against Missing fields This will affect the intensity. As anchor point The fill weights are calculated using a weighted average. fill value :

[0034] in anchor point Corresponding missing fields The known value, For the set of anchor points The number of anchor points; Based on the calculated fill-in values ​​for each missing field, the filled-in cost data is generated. .

[0035] In this embodiment, the anchor point selection criteria include: time window conditions. Spatial range conditions, Integrity condition, anchor point exist The missing field has a complete value; where for timestamp, anchor point timestamp, As a time threshold, for With anchor point Geographical distance between them This is the spatial threshold.

[0036] The anchor point selection criteria are designed to reflect the spatiotemporal characteristics of engineering cost data and the needs of business logic. The time window condition ensures that the anchor points are relevant to the data to be filled in the time dimension; data with excessively large time intervals has limited reference value, therefore a time threshold is set. Historical data with excessively long time intervals is filtered out. Spatial range conditions ensure that anchor points are geographically close to the data to be filled. Considering that project costs are affected by geographical factors such as regional economic level, material transportation distance, and labor costs, projects that are too far apart may have significantly different cost levels; therefore, a spatial threshold is set. Filter out historical data that is geographically too far away. Integrity conditions ensure that the anchor point has a complete and valid value in the field to be filled. This is a basic requirement for missing value imputation and avoids using missing or incomplete data in the imputation process.

[0037] The specific implementation process of anchor point filtering is as follows: The system first retrieves all project records from the historical cost database to construct a candidate anchor point set; then, it applies three filtering conditions in sequence. For the time window condition, the system calculates the absolute value of the time difference between each candidate anchor point and the data to be filled, retaining those with a time difference less than or equal to... For spatial range conditions, the system calculates the spatial distance between candidate anchor points and the data to be filled based on the project's geographical coordinates, and uses the Haversian formula to calculate the spherical distance, retaining distances less than or equal to... The system identifies anchor points; for integrity conditions, it checks the data integrity of each candidate anchor point in the field to be filled, excluding anchor points where the corresponding field contains null or outlier values. After three layers of filtering, the system obtains the final set of anchor points. If the number of filtered anchor points is less than the preset threshold, the system will appropriately relax the time window or spatial range conditions to ensure that enough anchor points participate in the filling process.

[0038] Regarding the spatiotemporal coupling coefficient The system employs the following method to determine the test set: A certain proportion of complete data is randomly selected from historical cost data; some fields of this data are manually set to be missing; and then different... A filling experiment was conducted. By comparing the error between the filled result and the true value, a grid search method was used to find the optimal value in the interval [0.1, 2.0] with a step size of 0.1. Value. Specifically, for each candidate Calculate the mean imputation error of all test samples and select the one with the smallest error. The value is used as the final parameter. The cross-validation process employs 5-fold cross-validation to ensure the robustness and generalization ability of the parameter selection. The physical meaning of the spatiotemporal coupling coefficient lies in describing the moderating effect of the interaction between the time and space dimensions on the intensity of the influence. When the value is large, the product of spatiotemporal distance has a more significant weakening effect on the intensity of the influence, and vice versa.

[0039] The preprocessed data is output in a unified standard format, containing complete fields such as basic project information, cost structure, technical parameters, geographical location, and timestamps, ensuring data quality and consistency.

[0040] Specifically, in one embodiment of the present invention, the semantic feature extraction module includes: The feature extraction unit is used to extract the original feature vector from the preprocessed cost data based on a four-level feature extraction system, including the project attribute feature layer, cost structure feature layer, technical process feature layer, and text semantic feature layer. The standardization unit is used to standardize the original feature vectors using an adaptive standardization strategy. Z-score standardization is used for normally distributed features, and quantile standardization is used for skewed features. This combines the feature vectors from each layer into a unified multidimensional feature vector representation. .

[0041] The four-level feature extraction system includes: The project attribute feature layer extracts project type, building area, number of floors, and structural form from basic project information. The project attribute feature vector is represented as follows: ,in This represents the total number of dimensions for the project attributes.

[0042] In this embodiment, classification attributes such as project type and structural form are converted into numerical vectors using one-hot encoding. The system pre-defines common building engineering types, including major categories such as residential, commercial, industrial, and public facilities, with each category corresponding to a coding bit. For continuous numerical attributes such as building area and number of floors, they are directly used as feature values. The system performs a reasonableness check on these values ​​and removes obviously abnormal values.

[0043] The cost structure feature layer calculates the proportion of each cost item to the total cost. The cost structure feature vector is represented as follows: ,in For the first Item cost amount, This represents the total cost of the project.

[0044] In this embodiment, during the cost structure feature extraction process, the system decomposes costs into major categories such as material costs, labor costs, machinery costs, management fees, and profits according to industry standards for cost estimation. The system then calculates the proportion of each cost category to the total cost to reflect the cost distribution characteristics of the project.

[0045] The technical process feature layer identifies technical keywords through engineering technology dictionary matching and calculates weight values. The technical feature vector is represented as follows: ,in For the first The weight value of each technical keyword.

[0046] In this embodiment, the system establishes an engineering technology dictionary, containing professional terms such as construction techniques, material specifications, and equipment models. The dictionary is organized in a hierarchical structure, including primary categories such as materials, processes, and equipment, and secondary categories such as steel, concrete, masonry, and decoration. Technical keywords are identified through dictionary matching, and their weight values ​​are calculated based on their frequency of occurrence in the project's technical description, their positional importance, and their terminology authority. For each identified technical keyword, the system uses the TF-IDF algorithm to calculate its importance weight in the current project, while also considering the keyword's rarity across the entire dataset.

[0047] The text semantic feature layer extracts semantic features from the project description text based on a domain dictionary and context analysis. The text semantic feature vector is represented as follows: ,in denoted as the number of dimensions of the semantic features.

[0048] In this embodiment, the system establishes a professional dictionary for engineering cost estimation, covering professional terms such as material names, construction techniques, equipment types, and engineering locations. Each entry includes standard terms, synonyms, and related words. Text feature extraction is achieved through a combination of dictionary matching and semantic analysis. First, the project description text is segmented and tagged with parts of speech. Then, it is matched with the professional dictionary to identify professional terms. For the identified terms, a weight value is calculated based on their position, frequency, and importance in the text. The system uses the term frequency-inverse document frequency method, combined with the contextual semantic relationship of the words and the importance score of the professional field.

[0049] The execution process of the standardized unit includes: The distribution characteristics of each feature layer are determined by the Shapiro-Wilk normality test, and each feature layer is identified as having normal or skewed distribution characteristics. The Z-score is used to standardize the characteristics of the normal distribution. The calculation formula is as follows:

[0050] in These are the original eigenvalues. The characteristic mean, Standard deviation; The quantile standardization is applied to the skewed distribution characteristics, and the calculation formula is as follows:

[0051] in and These are the minimum and maximum values ​​of the eigenvalues, respectively. Combine the feature vectors from each dimension into a complete multidimensional feature vector. All standardized feature values ​​are mapped to the interval [0,1].

[0052] The specific implementation process of the Shapiro-Wilk normality test is as follows: For each feature layer's feature value sequence, the system calculates the Shapiro-Wilk statistic. The calculation formula is:

[0053] in For the first A series of ordinal statistics The Shapiro-Wilk coefficient, This is the sample mean.

[0054] By calculating The value is compared with the critical value to determine whether the characteristic sequence follows a normal distribution. When the value is greater than the critical value, the characteristic sequence is considered to follow a normal distribution; when... When the value is less than the critical value, the characteristic sequence is considered to follow a skewed distribution.

[0055] Based on the characteristics of engineering cost data, continuous numerical features such as building area and number of floors in the project attribute feature layer typically follow a normal distribution because these values ​​have relatively stable distribution patterns in engineering practice. The cost structure feature layer typically exhibits a skewed distribution of cost percentages, as some cost items may dominate in specific types of projects, leading to uneven distribution. Keyword weights in the technology and process feature layer typically exhibit a skewed distribution because a few core technical keywords have higher weights, while most technical keywords have lower weights. The semantic feature weights in the text semantic feature layer also exhibit a skewed distribution, reflecting the uneven distribution of technical terms in project descriptions.

[0056] For sparsely distributed features, i.e., feature dimensions containing a large number of zero values, the system employs a non-zero value normalization strategy, standardizing only non-zero feature values ​​while leaving zero values ​​unchanged. This approach avoids the influence of a large number of zero values ​​on the standardization results, ensuring the effectiveness of sparse features. After standardization, all feature values ​​are mapped to the [0,1] interval, ensuring the comparability of features of different types and dimensions. For anomalous feature values ​​exceeding the normal range, the system uses a truncation method, restricting values ​​outside the [0,1] range to boundary values.

[0057] After feature extraction and standardization, the data is output as a standardized multi-dimensional feature vector. Each cost project corresponds to a feature vector of a unified dimension, where the sub-vectors are concatenated in a predefined order to form a complete feature vector. The output vector contains both feature dimension identifiers and weight information, providing a standardized data foundation for subsequent similarity calculations.

[0058] Specifically, in one embodiment of the present invention, the similarity calculation module constructs a similarity matrix using a pairwise comparison strategy, the process of which is as follows: For a dataset containing N cost data points Each cost data Corresponding to a multidimensional feature vector ; Iterate through all cost data using a double loop. ,in, and : Cost data is calculated based on similarity algorithms. Similarity values ​​between:

[0059] in, For cost data Similarity value between them and These represent the cost data respectively. and In the Normalized feature values ​​on each feature dimension For the first Weight coefficients for each feature dimension The total number of feature dimensions. and These represent the cost data respectively. and timestamp, For feature distance sensitivity parameters, Control parameters affected by time; After calculation, all similarity values ​​are indexed by row and column. Stored in In the similarity matrix.

[0060] In this embodiment, the similarity calculation formula is designed based on the feature space analysis and time decay theory of engineering cost data. Traditional similarity calculation methods, such as Euclidean distance and cosine similarity, suffer from the curse of dimensionality when processing high-dimensional sparse data and cannot effectively incorporate the influence of time factors. The similarity calculation method proposed in this invention converts feature similarity measurement into angular distance, mapping the angle between feature vectors to an angular distance using an inverse cosine function. The interval makes similarity calculation more stable and accurate. At the same time, the introduction of the exponential function ensures that the similarity value remains within a certain range. Within the interval, and exhibiting nonlinear response characteristics to feature differences and time differences, it can better distinguish between different degrees of similarity.

[0061] The concept of this formula is reflected in three aspects: First, it uses weighted angular distance instead of traditional Euclidean distance, through... Calculating the weighted inner product of eigenvectors and then converting it to angular distance using the inverse cosine function is a better way to handle high-dimensional sparse vectors and avoid the curse of dimensionality. Secondly, a time decay factor is introduced. The method comprehensively considers the cumulative effect of feature differences and the influence of time decay, enabling similarity calculation to reflect the timeliness characteristics of engineering cost data. Finally, through the nonlinear transformation of the exponential function, the similarity value has better discriminative power and can effectively identify different degrees of similarity.

[0062] The similarity calculation module uses a fixed weight allocation strategy, determining the specific values ​​based on the information content and importance of the four types of features. The weight of the cost structure feature is set as follows: Because cost composition directly reflects the economic characteristics and resource allocation pattern of a project, the proportion of each cost item to the total cost is the most important indicator for judging project similarity, directly reflecting the project's cost level and cost structure characteristics. The weights for project attribute characteristics are set as follows: Because basic attributes such as project type, building area, number of floors, and structural form determine the scale, complexity, and fundamental characteristics of a project, they are important criteria for judging project similarity. The weight of technical and process characteristics is set as follows: Because technical characteristics such as construction techniques, material specifications, and equipment models reflect the project's technical level and construction methods, they have a significant impact on cost, but their importance is slightly lower than that of cost structure and project attributes. The text semantic feature weights are set to... Although textual semantic features contain rich information such as material names, construction techniques, equipment types, and engineering locations, their weight is relatively low due to the strong subjectivity of textual descriptions and the certain overlap with technical and process features. They are mainly used to assist in judgment and supplement technical feature information.

[0063] Feature distance sensitivity parameter The value range is 1.5-2.5, and the recommended value is [value missing]. The parameter was determined based on statistical analysis and cross-validation experiments of historical cost data. The system analyzes different... To determine the optimal parameter values, the discriminative power and accuracy of the similarity calculation results are assessed. The specific determination process is as follows: Project pairs with clearly defined similarity relationships are selected from historical cost data as the validation set, and... A grid search is performed with a step size of 0.1 for values ​​in the range of 1.0-3.0, and each... Calculate the accuracy and F1 score of similarity at each value, and select the value that optimizes the evaluation metric. value.

[0064] Time affects control parameters The value range is 0.1-0.3, and the recommended value is [value missing]. The determination of this parameter is based on the analysis of the timeliness characteristics of engineering cost data and practical application needs. The system determines the appropriate time decay rate by analyzing the changing patterns of project similarity over different time intervals. The specific determination process is as follows: statistically analyze the time distribution characteristics of similar projects in historical cost data, analyze the impact of time intervals on the accuracy of similarity judgment, and determine the reasonable range of time decay through regression analysis. The value of this parameter implies that the similarity decreases by approximately 18% for every additional year in the time interval. This degree of decay aligns with the time-sensitive nature of engineering cost data, ensuring the priority of recent data while not completely excluding the reference value of historical data. (Time difference) To eliminate the influence of time units on the calculation results, the system converts the time difference into a value in days and sets an effective range for the time difference according to application requirements. Items that exceed the range will not participate in the similarity calculation.

[0065] The module calculates the similarity scores for all item pairs in the dataset and generates a symmetric similarity matrix. , of which elements Indicates project and projects The similarity between them. The similarity value ranges from 1 to 2. A higher similarity score indicates greater similarity between items. For any query item, the system sorts all other items based on their similarity scores, generating a list of similar items. Users can set a similarity threshold to filter out items with similarity scores exceeding the threshold as recommended results.

[0066] Specifically, in one embodiment of the present invention, the data storage and management module is used to store cost data, multi-dimensional feature vectors, and similarity matrices using a hybrid storage architecture, and provides retrieval and recommendation functions for cost projects based on a multi-level feature space index structure. The hybrid storage architecture of this module includes two core storage units: a relational database and a vector database.

[0067] The relational database storage unit stores preprocessed cost data and a mapping table between index nodes and the row and column coordinates of the similarity matrix. The system stores structured data such as basic project information, cost details, and timestamps in the relational database. The relational database ensures transactional consistency and integrity constraints, and supports standard SQL queries and data management operations. Basic project information includes key fields such as project number, project name, construction unit, project type, building area, number of floors, structural form, construction location, and construction time. Cost details include various cost categories and their specific amounts, such as material costs, labor costs, machinery costs, management fees, and profit. Timestamps include time-related information such as project start time, completion time, data collection time, and data update time.

[0068] At the same time, relational databases maintain a mapping table between index nodes and the row and column coordinates of the similarity matrix. The mapping table is implemented using a hash table data structure: for each item identifier The system in the similarity matrix Find its corresponding row and column coordinates in the middle. Establish mapping relationship The mapping table is implemented using a hash table data structure to ensure... The search time complexity is [not specified]. Simultaneously, the system maintains the reverse mapping relationship. It supports fast conversion from similarity matrix coordinates to item identifiers.

[0069] The vector database storage unit stores multidimensional feature vectors, similarity matrices, and multi-level feature space index structures. Multidimensional feature vectors output from the semantic feature extraction module are stored in a dedicated vector database. This database is optimized for storing high-dimensional feature vectors and for similarity retrieval, supporting fast vector similarity calculation and nearest neighbor queries. Each cost project corresponds to a unique vector record, with the vector dimension defined by the feature extraction module. ,in For project attribute feature dimensions, As a dimension of cost structure characteristics, As a dimension of technical process characteristics, This refers to the semantic features dimension of the text. The vector database adopts a distributed storage architecture, ensuring high availability and access performance through data sharding and replication mechanisms.

[0070] Similarity matrix As The symmetric matrix is ​​stored in the vector database, where For the total number of projects, This represents the similarity value between item i and item j. Considering the symmetry and sparsity of the similarity matrix, the system adopts a compressed storage format, storing only the upper triangular part of the matrix, and uses sparse matrix storage technology to save only non-zero similarity values, significantly reducing storage space requirements.

[0071] The multi-level feature space index structure is also stored in the vector database. Spatial index nodes are constructed by splitting multi-dimensional feature vectors into KD trees. Each index node records the feature space range it covers and the corresponding list of cost project identifiers. At the same time, a mapping relationship table between index nodes and row and column coordinates of the similarity matrix is ​​established. When searching for cost projects, the target index node is located in the KD tree according to the query feature vector. The candidate project identifiers covered by the node and its neighboring nodes are obtained. The candidate project identifiers are converted into coordinate positions in the similarity matrix through the mapping relationship table. The corresponding similarity values ​​are read in batches from the similarity matrix and sorted in descending order of similarity values ​​to generate a list of recommended projects.

[0072] The construction process of the multi-level feature space index structure is as follows: Based on the distribution pattern of historical feature vectors, the system uses the KD-tree algorithm to recursively divide the high-dimensional feature space into multiple subspace regions. The KD-tree construction starts from the root node, and each time the feature dimension with the largest variance is selected as the split dimension, and the median of that dimension is selected as the split point, dividing the dataset into two subsets. The splitting process is recursively executed until the number of items contained in each leaf node is less than a preset threshold. Each index node Record the feature space range it covers and the corresponding list of cost item identifiers ,in This represents the number of items contained in this node.

[0073] The data association mechanism uses the project's unique identifier ID. project Establish a connection between the relational database and the vector database. The system assigns a unique project identifier to each cost project, which remains consistent across both databases. When performing a similarity query, the system first finds similar project identifiers in the vector database based on the index structure, and then retrieves complete project details from the relational database based on the identifiers, thus achieving cross-database data association and querying.

[0074] The specific implementation process of cost project retrieval is as follows: After the user inputs the basic information of the project to be queried, the system calls the semantic feature extraction module to convert it into a standardized feature vector. Based on the query feature vector A recursive search is performed in the KD-tree, starting from the root node. The search compares the feature values ​​of the query vector in the current split dimension with the split point, selecting the appropriate subtree to continue the search until a leaf node is reached, at which point the target index node is determined. Obtain the candidate item identifiers covered by the target index node and its neighboring nodes. Neighboring nodes are determined by calculating the spatial distance between nodes, and the one closest to the target node is selected. Each node is considered an adjacent node. These are preset parameters. Utilizing a mapping table... The process of converting candidate item identifiers into coordinate positions in a similarity matrix is ​​as follows: For each candidate item identifier... Search This allows us to obtain its position in the similarity matrix.

[0075] The system reads the corresponding similarity values ​​in batches from the similarity matrix, and then determines the row coordinates of the query items in the similarity matrix. Read the similarity values ​​corresponding to the column coordinates of all candidate items in this row. ,in The column coordinates represent the candidate items. A list of recommended items is generated by sorting the candidates in descending order of similarity value. The system quickly sorts the candidate items by similarity value and selects the top ones with the highest similarity. One item was selected as the recommendation result, among which The maximum number of recommendations that can be set for a user.

[0076] The system also supports setting similarity thresholds. and time range Search parameters are specified. When the user sets a similarity threshold of 0.7, the system only returns items with a similarity greater than 0.7. When the user specifies a time range of 2020-2022, the system prioritizes returning similar items within that time period and adjusts the similarity value according to a time decay factor. Search results are sorted from highest to lowest similarity, and each result item includes a similarity value, basic item information, and key differences to help users quickly understand the basis for recommendations.

[0077] To improve retrieval efficiency, the system employs a query result caching mechanism, caching user query results, including query vectors, lists of similar items, and similarity scores. When a user initiates the same or similar query, the results are directly returned from the cache, avoiding duplicate calculations. The cache uses an LRU eviction policy, prioritizing the retention of frequently accessed query results.

[0078] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A big data management and analysis system for engineering cost, characterized in that, include: The data acquisition and preprocessing module is used to acquire raw cost data from heterogeneous data sources and preprocess the raw cost data to obtain preprocessed cost data. The semantic feature extraction module is used to convert the preprocessed cost data into multi-dimensional feature vectors based on a multi-level feature extraction system. The similarity calculation module is used to calculate the similarity between different cost data based on multi-dimensional feature vectors, and obtain a similarity matrix; The data storage and management module is used to store cost data, multi-dimensional feature vectors, and similarity matrices using a hybrid storage architecture, and provides retrieval and recommendation functions for cost projects based on a multi-level feature space index structure.

2. The engineering cost big data management and analysis system according to claim 1, characterized in that, The data acquisition and preprocessing module includes: The data acquisition unit is used to collect raw cost data from the project management system, cost software database, and bidding platform via API interface, database connection, and file import, and to add a timestamp to each cost record. The data preprocessing unit is used to perform format conversion, data cleaning, and missing value imputation on the original cost data to obtain preprocessed cost data.

3. The engineering cost big data management and analysis system according to claim 2, characterized in that, The data preprocessing unit employs a spatiotemporal anchor diffusion imputation strategy to fill in missing values, and the process is as follows: For the cost data to be filled Anchor point sets are identified from historical cost data based on anchor point selection criteria. ; Each anchor point Cost data to be filled Intensity of influence between for: , in, For standardization With anchor point Time difference, For standardization With anchor point Spatial distance, anchor point Data credibility weight, The spatiotemporal coupling coefficient is determined through cross-validation of historical cost data; against Missing fields This will affect the intensity. As anchor point The fill weights are calculated using a weighted average. fill value : , in anchor point Corresponding missing fields The known value, For the set of anchor points The number of anchor points; Based on the calculated fill-in values ​​for each missing field, the filled-in cost data is generated. .

4. The engineering cost big data management and analysis system according to claim 3, characterized in that, The anchor point selection criteria include: time window conditions. Spatial range conditions, Integrity condition, anchor point exist The missing field has a complete value; where for timestamp, anchor point timestamp, As a time threshold, for With anchor point Geographical distance between them This is the spatial threshold.

5. The engineering cost big data management and analysis system according to claim 1, characterized in that, The semantic feature extraction module includes: The feature extraction unit is used to extract the original feature vector from the preprocessed cost data based on a four-level feature extraction system, including the project attribute feature layer, cost structure feature layer, technical process feature layer, and text semantic feature layer. The standardization unit is used to standardize the original feature vectors using an adaptive standardization strategy. Z-score standardization is used for normally distributed features, and quantile standardization is used for skewed features. This combines the feature vectors from each layer into a unified multidimensional feature vector representation. .

6. The engineering cost big data management and analysis system according to claim 5, characterized in that, The four-level feature extraction system includes: The project attribute feature layer extracts project type, building area, number of floors, and structural form from basic project information. The project attribute feature vector is represented as follows: ,in This represents the total number of dimensions for the project attributes. The cost structure feature layer calculates the proportion of each cost item to the total cost. The cost structure feature vector is represented as follows: ,in For the first Item cost amount, The total cost of the project; The technical process feature layer identifies technical keywords through engineering technology dictionary matching and calculates weight values. The technical feature vector is represented as follows: ,in For the first The weight value of each technical keyword; The text semantic feature layer extracts semantic features from the project description text based on a domain dictionary and context analysis. The text semantic feature vector is represented as follows: ,in denoted as the number of dimensions of the semantic features.

7. The engineering cost big data management and analysis system according to claim 6, characterized in that, The execution process of the standardized unit includes: The distribution characteristics of each feature layer are determined by the Shapiro-Wilk normality test, and each feature layer is identified as having normal or skewed distribution characteristics. The Z-score is used to standardize the characteristics of the normal distribution. The calculation formula is as follows: , in These are the original eigenvalues. The characteristic mean, Standard deviation; The quantile standardization is applied to the skewed distribution characteristics, and the calculation formula is as follows: , in and These are the minimum and maximum values ​​of the eigenvalues, respectively. Combine the feature vectors from each dimension into a complete multidimensional feature vector. All standardized feature values ​​are mapped to the interval [0,1].

8. The engineering cost big data management and analysis system according to claim 1, characterized in that, The similarity calculation module constructs a similarity matrix using a pairwise comparison strategy, as follows: For a dataset containing N cost data points Each cost data Corresponding to a multidimensional feature vector ; Iterate through all cost data using a double loop. ,in, and : Cost data is calculated based on similarity algorithms. Similarity values ​​between: , in, For cost data Similarity value between them and These represent the cost data respectively. and In the Normalized feature values ​​on each feature dimension For the first Weight coefficients for each feature dimension The total number of feature dimensions. and These represent the cost data respectively. and timestamp, For feature distance sensitivity parameters, Control parameters affected by time; After calculation, all similarity values ​​are indexed by row and column. Stored in In the similarity matrix.

9. The engineering cost big data management and analysis system according to claim 1, characterized in that, Based on a multi-level feature space index structure, this system provides retrieval and recommendation functions for cost estimation projects, including: Spatial index nodes are constructed by KD-tree segmentation of multidimensional feature vectors. Each index node records the feature space range it covers and the corresponding list of cost project identifiers. At the same time, a mapping relationship table between index nodes and row and column coordinates of similarity matrix is ​​established. When searching for cost projects, the target index node is located in the KD-tree according to the query feature vector, and the candidate project identifiers covered by the node and its neighboring nodes are obtained. The candidate project identifiers are converted into coordinate positions in the similarity matrix through the mapping relationship table. The corresponding similarity values ​​are read in batches from the similarity matrix and sorted in descending order of similarity value to generate a recommended project list.

10. The engineering cost big data management and analysis system according to claim 9, characterized in that, The hybrid storage architecture of the data storage and management module includes: a relational database for storing preprocessed cost data and a mapping table between index nodes and row and column coordinates of the similarity matrix; and a vector database for storing multidimensional feature vectors, similarity matrices, and multi-level feature space index structures.

Citation Information

Patent Citations

  • A method for managing engineering cost data

    CN114490604B

  • Identifying critical features in ordered scale space

    CA2509580A1

  • Digital banking platform and architecture

    CA3045736A1

  • Transparent characterization method for element universe anchor point network

    CN117040850A

  • Engineering cost data management system

    CN117454225A

Cited By

  • Intelligent construction method and system of cost database

    CN121051223A

  • Intelligent construction method and system of cost database

    CN121051223B

  • Multi-level data collection method, system and equipment and storage medium

    CN121210521A

  • Construction engineering cost data identification and standardization method and system

    CN121579583A