An engineering cost auditing method and system based on multi-source data fusion
By employing a multi-source data fusion-based engineering cost auditing method, and utilizing a multimodal feature coding model and Mahalanobis distance calculation, the problem of insufficient multi-source data fusion is solved, and efficient, accurate, and traceable results for engineering cost auditing are generated.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OCEAN UNIVERSITY
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-01
AI Technical Summary
Existing engineering cost auditing methods suffer from insufficient integration of multi-source data, limited accuracy in feature extraction and matching, weak anomaly detection and tracing capabilities, and low efficiency in utilizing historical compliance data, resulting in low accuracy and efficiency of audit results.
By acquiring bill of quantities, spatiotemporal images, and construction log text data, a multimodal feature coding model is used to map them into semantically consistent feature vectors. The correlation score is calculated and a semantic fusion feature matrix is generated. Anomaly detection is performed using Mahalanobis distance, and detailed audit results are generated.
It achieves efficient fusion and accurate matching of multi-source data, improves the accuracy and efficiency of engineering cost audit, can accurately locate anomalies and trace their causes, and generate clear audit results.
Smart Images

Figure CN121279966B_ABST
Abstract
Description
A method and system for engineering cost auditing based on multi-source data fusion Technical Field
[0001] This invention relates to the field of engineering cost auditing, and in particular to an engineering cost auditing method and system based on multi-source data fusion. Background Technology
[0002] In engineering project management, cost auditing is a crucial step in controlling investment costs and ensuring project quality and compliance. Its core objective is to verify whether the cost items declared in the bill of quantities are consistent with the actual construction conditions on site, avoiding waste of funds or compliance risks caused by false or incorrect reporting. As engineering projects expand in scale and increase in process complexity, the data sources for cost auditing are becoming increasingly diversified. In addition to traditional bill of quantities data, spatiotemporal image data of the construction site (including time and space correlation information of the construction process) and construction log text data (recording the actual construction content, personnel, material usage, etc.) have become important empirical evidence. How to efficiently integrate and utilize these multi-source heterogeneous data has become a core requirement for improving the accuracy and efficiency of auditing.
[0003] However, existing methods for reviewing construction costs still face several technical bottlenecks: First, the integration of multi-source data is insufficient. The cost declaration data in the bill of quantities and the physical reality data from spatiotemporal images and construction logs belong to different modalities (text and visual). Existing technologies mostly rely on manual comparison or simple data splicing, making it difficult to uncover the deep semantic relationships between the data, resulting in a "disconnect between the cost declaration and the actual construction situation." The problems are prominent; secondly, the accuracy of feature extraction and matching is limited, and there is a lack of a dedicated feature encoding mechanism for engineering cost scenarios, which makes it impossible to effectively map multimodal data into feature vectors with semantic consistency, resulting in deviations in the determination of the correlation between cost declarations and empirical data; thirdly, the ability to identify and trace anomalies is weak. Existing methods mostly rely on a single threshold or human experience to judge anomalies, which not only makes it difficult to accurately locate specific cost items and corresponding construction periods that lack empirical support, but also makes it impossible to trace the cause of anomalies after they are identified, resulting in insufficient practicality and guidance of the audit results; fourthly, the utilization efficiency of historical compliance data is low, and standardized process feature benchmarks have not been formed. The audit process relies too much on the professional experience of the auditors, resulting in low audit efficiency and strong subjectivity of the results. These shortcomings make it impossible for existing technologies to accurately and efficiently audit engineering costs. Summary of the Invention
[0004] This invention provides a method and system for auditing engineering costs based on multi-source data fusion, in order to solve the problem that existing technologies cannot accurately and efficiently audit engineering costs.
[0005] Firstly, this application provides a method for engineering cost auditing based on multi-source data fusion, including:
[0006] Obtain the bill of quantities data, spatiotemporal image data of the construction site, and text data of the construction log for the project to be reviewed;
[0007] The bill of quantities data, spatiotemporal image data, and construction log text data are input into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector, and maps the spatiotemporal image data and the construction log text data into a physical reality feature vector.
[0008] The multimodal feature encoding model is based on multimodal data pairs of similar sub-projects of historical compliant projects, and is trained by minimizing the semantic distance between the historical cost declaration vector and the historical physical reality vector using a contrastive loss function.
[0009] Calculate the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time segment feature; extract the time segment feature with the highest score from the correlation score sequence as the empirical segment feature;
[0010] The cost declaration feature vector and the empirical fragment feature are fused to generate a semantic fusion feature matrix;
[0011] The sub-item engineering category to which the bill of quantities data belongs is identified, and the process feature prototype vector generated by clustering based on historical compliance data under the sub-item engineering category to which the bill of quantities data belongs is obtained. The Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vector is calculated.
[0012] If the Mahalanobis distance is greater than the preset anomaly tolerance threshold, it is determined that the current sub-item of the project to be reviewed has an anomaly; based on the location information of the empirical fragment features, an audit result is generated that includes the cost item corresponding to the sub-item with the anomaly and the corresponding construction period lacking empirical information.
[0013] This application overcomes the limitations of traditional auditing methods that rely on single data sources by acquiring three types of multi-source heterogeneous data: bill of quantities data, spatiotemporal imagery data of construction sites, and textual construction log data. This provides a comprehensive and complementary empirical foundation for comparing cost declarations with actual construction conditions. By inputting multi-source data into a multi-modal feature encoding model trained on multi-modal data of similar sub-items from historical compliant projects, and employing a contrastive loss function to minimize the semantic distance between historical cost declaration vectors and historical physical reality vectors, different modalities of data can be mapped to semantically consistent cost declaration feature vectors and physical reality feature vectors. This effectively solves the problems of insufficient multi-source data fusion and inadequate semantic association mining, significantly improving the accuracy of feature matching. Finally, by calculating the correlation scores of cost declaration feature vectors and physical reality feature vectors for each time-series segment and extracting the highest-scoring features, the application provides a more comprehensive and complementary empirical foundation for comparing cost declaration feature vectors and actual construction conditions. The empirical fragment features enable precise location of the construction site fragments most relevant to the cost declaration, avoiding audit bias caused by the mismatch between empirical data and the time of cost items. By fusing the cost declaration feature vector with the empirical fragment features to generate a semantic fusion feature matrix, the semantic association between the two is further strengthened, providing more comprehensive feature support for anomaly detection. By identifying the categories of sub-items and obtaining the process feature prototype vectors generated by clustering historical compliance data, and combining the sensitivity of Mahalanobis distance to feature distribution differences for anomaly detection, the reliance on manual experience is eliminated, achieving objective quantitative identification of anomalies. Finally, based on the location information of the empirical fragment features, audit results containing abnormal cost items and corresponding construction periods are generated, accurately tracing specific links lacking empirical support, effectively solving the problem that existing technologies cannot accurately and efficiently audit engineering costs.
[0014] Furthermore, the acquisition of the bill of quantities data for the project to be reviewed, as well as the spatiotemporal image data of the construction site and the text data of the construction log as empirical evidence, specifically includes:
[0015] Extract bill of quantities data from the project management file of the project to be reviewed. The bill of quantities data includes the specific name, declared quantity, declared unit price and declared total price of each sub-item.
[0016] Based on the pre-set construction area of the construction site, the camera equipment with positioning function is deployed to collect image data of the construction site at fixed time intervals, and simultaneously record the specific time and the location of the equipment at the site for each collection operation, and integrate them to obtain spatiotemporal image data of the construction site.
[0017] Obtain the daily construction log text data submitted by the construction site. The construction log text data includes the date of construction, the actual construction content, the number of personnel involved in the construction, and the actual amount of construction materials used.
[0018] This application ensures the compliance of the cost declaration data source and the completeness of core information by explicitly extracting the bill of quantities data from the project management file of the project to be reviewed, including the specific name, declared quantity, declared unit price, and declared total price of each sub-item. This avoids subsequent feature mapping deviations caused by unclear data extraction channels or missing key fields. By deploying positioning-enabled cameras in the pre-defined construction area at the construction site, and collecting image data at fixed time intervals while simultaneously recording the specific time and location, this application integrates spatiotemporal image data. This ensures both the standardization and continuity of empirical data collection and endows the image data with precise spatiotemporal attributes, solving the problem of traditional image data lacking spatiotemporal correlation and unable to accurately correspond to construction progress and cost items. Through obtaining... The daily construction log text data submitted from the construction site is used, clearly specifying the date of construction, actual construction content, number of personnel involved, and actual amount of construction materials used. This ensures the timeliness and completeness of core elements of the construction data, enabling a quantitative record of the construction process. The acquisition methods for all three types of data clearly define the source, collection specifications, and core content, allowing the subsequent multimodal feature coding model to obtain high-quality, highly relevant input data. This provides a solid foundation for the accurate mapping of cost declaration feature vectors and physical reality feature vectors, thereby improving the accuracy of subsequent correlation score calculations, feature fusion, and anomaly detection. From the data source, this ensures the reliability and objectivity of the engineering cost audit process, effectively reducing audit errors caused by data quality issues.
[0019] Furthermore, the step of inputting the bill of quantities data, spatiotemporal image data, and construction log text data into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector, and maps the spatiotemporal image data and the construction log text data into a physical reality feature vector, specifically:
[0020] The bill of quantities data is input into the text encoding module of the multimodal feature encoding model, so that the text encoding module can extract the bill of quantities data and convert it into vector form, and output the cost declaration feature vector.
[0021] Spatiotemporal image data is input into the visual encoding module of the multimodal feature encoding model so that the visual encoding module can extract construction scene features, material features and equipment features from the spatiotemporal image data, while incorporating the spatiotemporal information corresponding to the image. The extracted features are then vectorized and fused to obtain the feature vector corresponding to the image.
[0022] The construction log text data is segmented into words; the processed construction log text data is then input into the text encoding module of the multimodal feature encoding model, so that the text encoding module can extract features and transform vectors from the processed log text data to obtain the feature vectors corresponding to the logs.
[0023] The feature vectors corresponding to the images and the feature vectors corresponding to the logs are concatenated and fused to obtain the physical reality feature vector.
[0024] This application inputs bill of quantities data into the text encoding module of a multimodal feature encoding model. The text encoding module then extracts the core cost information from the data and converts it into vector form. This ensures that the cost declaration feature vector accurately carries the key information for each sub-item of the project, establishing a clear semantic benchmark for subsequent comparison with actual construction conditions. Spatiotemporal image data is input into the visual encoding module, which not only extracts visual features such as construction scenes, materials, and equipment, but also integrates the corresponding spatiotemporal information and performs vector transformation and fusion. This fully leverages the real-world representation value of visual data and solves the problem of pure visual features lacking temporal and locational references through spatiotemporal information association, enabling the feature vector corresponding to the image to accurately map the construction status at a specific time and location. Construction log text data is first segmented before being input into the text encoding module. This module effectively filters out key construction information from logs, avoiding interference from redundant text in feature extraction and ensuring that the feature vectors corresponding to the logs accurately reflect the core real-time elements of the construction process. Finally, by splicing and fusing the feature vectors corresponding to the images and the logs, complementary fusion of visual and textual real-time data is achieved. This allows the generated physical real-time feature vectors to simultaneously cover multi-dimensional information such as intuitive construction scenes, specific construction elements, and quantitative construction records, forming a semantically comparable feature correspondence with the cost declaration feature vectors. This completely solves the problems of incomparable features and insufficient fusion caused by modal differences in traditional multi-source data, providing a high-quality and highly consistent feature foundation for subsequent correlation score calculation, feature fusion, and anomaly detection, thereby significantly improving the accuracy and reliability of engineering cost auditing.
[0025] Furthermore, the multimodal feature encoding model is based on multimodal data pairs of similar sub-projects from historical compliant projects, and is trained by minimizing the semantic distance between the historical cost declaration vector and the historical physical reality vector using a contrastive loss function. Specifically:
[0026] Obtain engineering files for multiple historical compliant projects, classify the engineering files according to the categories of sub-items, extract the corresponding historical bill of quantities data, historical spatiotemporal image data and historical construction log text data from each category, and combine them to form multimodal data pairs of the same type of sub-items.
[0027] Based on the preset text encoding module and the preset visual encoding module, combined with the preset initialization parameters, an initial multimodal feature encoding model is constructed.
[0028] The multimodal data pairs are input into the initial multimodal feature encoding model to obtain the historical cost declaration feature vector corresponding to the historical bill of quantities data in each multimodal data pair, as well as the historical physical reality feature vector after fusing historical construction site spatiotemporal image data and historical construction log text data.
[0029] Based on the contrastive loss function, calculate each first semantic distance between each group of historical cost declaration feature vectors and the corresponding historical physical condition feature vectors of the same group, and calculate each second semantic distance between each group of historical cost declaration feature vectors and the historical physical condition feature vectors of different groups.
[0030] Calculate the contrastive loss value based on each first semantic distance and each second semantic distance;
[0031] Based on the contrast loss value, the parameters of the text encoding module and the visual encoding module in the initial multimodal feature encoding model are adjusted until the contrast loss value drops to a preset convergence threshold, thus obtaining the multimodal feature encoding model.
[0032] This application selects multimodal data pairs of similar sub-items from historical compliant projects as model training data. It first categorizes project files according to sub-item categories before extracting corresponding historical data to form data pairs. This ensures the matching of training data with the project under review category and data compliance, avoiding the problem of insufficient model generalization ability caused by cross-category data interference. This lays a precise data foundation for the model to learn the inherent semantic relationship between cost declarations and construction realities of similar projects. By constructing an initial multimodal feature encoding model based on text encoding modules, visual encoding modules, and preset initialization parameters, the core architecture and initial configuration of the model are clarified. This ensures that the model can specifically handle multi-source heterogeneous data of text types (bills of quantities, construction logs) and visual types (spatiotemporal images), avoiding the limitations of single-modal encoding models that cannot adapt to multi-source data. By inputting multimodal data pairs into the initial model to generate historical cost declaration feature vectors and historical physical reality feature vectors, it provides a basis for subsequent semantic distance calculation and model optimization. The model directly manipulates the target object; by comparing the loss function to calculate the first semantic distance of the same group of feature vectors and the second semantic distance of different groups of feature vectors, it can effectively strengthen the semantic correlation between cost declarations and corresponding construction realities in similar projects, while weakening semantic interference with non-corresponding construction realities, enabling the model to learn more discriminative feature mapping rules; by calculating and comparing the loss value based on the two types of semantic distances and adjusting the model parameters until the loss value converges to the preset threshold, it ensures that the model continuously optimizes its encoding ability during training. The resulting multimodal feature encoding model has stable and accurate feature mapping performance, which can efficiently transform multi-source data of projects to be reviewed into semantically consistent and closely related feature vectors. This completely solves the problem of large feature mapping deviations and incomparable semantics of multimodal data caused by the lack of targeted training in traditional encoding models, providing highly reliable technical support for subsequent correlation score calculation, feature fusion, and anomaly detection, thereby significantly improving the accuracy, stability, and intelligence level of project cost review.
[0033] Furthermore, the step of calculating the contrastive loss value based on each first semantic distance and each second semantic distance is specifically as follows:
[0034] The second semantic distance with the smallest value is selected from multiple second semantic distances as the comparison distance; the difference between the comparison distance and the first semantic distance is calculated, and the difference is compared with a preset marginal value;
[0035] If the difference is less than the marginal value, the marginal value minus the difference is taken as a single loss component; if the difference is greater than or equal to the marginal value, the single loss component is recorded as zero.
[0036] The individual loss components corresponding to all historical cost declaration vectors are summed to obtain the comparative loss value.
[0037] This application, by selecting the smallest value from multiple second semantic distances as the comparison distance, can accurately identify the non-corresponding historical physical reality vectors most easily confused with the current historical cost declaration vector. This allows the loss calculation to focus on the "same-type interference items" that the model finds most difficult to distinguish, avoiding the problem of insufficient targeting of the loss function caused by randomly selecting comparison vectors. This provides a precise target for the model to learn feature mapping rules with strong discriminative power. By calculating the difference between the comparison distance and the first semantic distance and comparing it with a preset marginal value, the difference in semantic correlation between the same group of vectors and the semantic correlation between the most similar dissimilar vectors can be quantified, clarifying the current boundary of the model's discriminative ability. When the difference is less than the marginal value, it indicates that the model has not yet fully distinguished between corresponding and non-corresponding vectors, and parameter adjustments are required through loss components. When the difference reaches the target, no additional optimization is needed, ensuring the effectiveness of training while avoiding overfitting caused by overtraining. By summing the individual loss components corresponding to all historical cost declaration vectors to obtain the total comparison loss value, the overall discriminative performance of the model on all training samples can be comprehensively reflected, ensuring that parameter adjustments are based on the global training effect rather than the bias of a single sample, making model optimization more holistic and stable. The entire comparison loss value calculation process is achieved by "accurately identifying interference items—" The logic of "quantifying differences and globally integrating loss" allows the loss function to efficiently guide the model to strengthen the semantic aggregation of corresponding vectors and the semantic separation of non-corresponding vectors. This solves the problems of unclear target differentiation and low training efficiency in traditional loss calculation. It enables the trained multimodal feature encoding model to have more accurate semantic mapping capabilities, providing a highly reliable feature foundation for the subsequent determination of the correlation between cost declaration and construction reality, thereby improving the accuracy and anti-interference ability of engineering cost review.
[0038] Furthermore, the step of calculating the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time-series segment feature; and extracting the time-series segment feature with the highest score from the correlation score sequence as the empirical segment feature, specifically involves:
[0039] The physical reality feature vector is divided into multiple time-series segment features of equal length according to the time sequence, and the time interval corresponding to each time-series segment feature of equal length is recorded;
[0040] The correlation score between each time segment feature and the cost declaration feature vector is obtained by summing the element-wise multiplications of the cost declaration feature vector and each time segment feature.
[0041] Arrange all correlation scores sequentially according to the time sequence of the time segments to form a correlation score sequence;
[0042] Obtain the score with the largest value in the correlation score sequence, and determine the time segment feature corresponding to the largest score as the empirical segment feature.
[0043] This application solves the problem of inaccurate matching between cost declarations and construction periods caused by the lack of temporal segmentation in traditional empirical data by dividing the physical reality feature vector into multiple equal-length time-series feature segments and recording the corresponding time intervals. This ensures the temporal integrity of the construction reality data and assigns a clear time anchor point to each segment, providing a precise temporal dimension reference for subsequent correlation determination. By using an element-wise multiplication and summation operation to calculate the correlation score between the cost declaration feature vector and each time-series feature segment, the application can directly quantify the semantic fit between the two types of feature vectors. This operation logic is simple and highly targeted, avoiding the complications of complex calculations. This process reduces efficiency losses while ensuring the intuitiveness and accuracy of correlation judgments, eliminating the possibility of incorrect selection of empirical segments due to fuzzy judgments. By arranging all correlation scores into a sequence according to the time order of time segments, a one-to-one correspondence between scores and construction time is established, facilitating the rapid location of the construction period most closely related to the cost statement and improving the efficiency of subsequent empirical segment screening. By extracting the features of the time segment with the highest score as empirical segment features, the core construction data with the highest semantic fit with the cost statement can be accurately identified, effectively eliminating interference from irrelevant or low-correlation time segments and ensuring the relevance and effectiveness of the empirical data. The entire process, through a logical closed loop of "time segmentation to define intervals - precise calculation to calculate correlations - orderly arrangement to establish correspondences - high-score screening to lock in the core," solves the problems of fuzzy empirical segment positioning, inaccurate correlation judgments, and low efficiency in extracting effective information in traditional methods. This allows subsequent feature fusion and anomaly judgment to be based on the most representative empirical data, significantly improving the accuracy, relevance, and efficiency of engineering cost auditing.
[0044] Furthermore, the step of identifying the sub-item engineering category to which the bill of quantities data belongs, and obtaining the process feature prototype vector generated based on historical compliance data clustering under the sub-item engineering category to which the bill of quantities data belongs, specifically involves:
[0045] Obtain the category identification information of each sub-item of the bill of quantities data, and determine the category of the sub-item of the data to be reviewed by combining it with the preset engineering category classification standards.
[0046] Obtain multimodal data pairs of similar historical compliant sub-items under the sub-item engineering category to which the data to be audited belongs. Input each multimodal data pair into the multimodal feature encoding model to obtain the corresponding historical cost declaration feature vector and historical physical condition feature vector.
[0047] Element-wise weighted summation is performed on each set of historical cost declaration feature vectors and historical physical condition feature vectors to obtain the corresponding historical fusion feature vector for each set of multimodal data pairs;
[0048] Based on the clustering algorithm and the preset number of clusters, and combining all historical fusion feature vectors, the historical fusion feature vectors with the highest similarity are grouped into one class to obtain multiple process clusters.
[0049] Wherein, the number of clusters is the number of process types;
[0050] Calculate the average value of each element corresponding to all historical fusion feature vectors in each cluster, and use the vector composed of the average values as the process feature prototype vector corresponding to each cluster.
[0051] This application determines the category of each sub-item of the project to be reviewed by obtaining category identification information from the bill of quantities data and combining it with preset project category classification standards. This ensures the accuracy and standardization of category identification and avoids the problem of subsequent historical data matching deviation caused by ambiguous category determination, laying the foundation for the targeted acquisition of subsequent process feature prototype vectors. By obtaining historical compliant multimodal data of similar sub-items and inputting it into a trained multimodal feature encoding model, historical cost declaration feature vectors and historical physical condition feature vectors are generated. This ensures the semantic consistency and comparability of historical feature vectors and feature vectors of the project to be reviewed, avoiding feature deviations caused by inconsistent encoding logic. At the same time, the compliance benchmark of the prototype vector is ensured by relying on historical compliant data. By performing element-level weighted summation on each group of historical two-class feature vectors, the following is obtained: By integrating historical feature vectors, a deep correlation between historical cost declarations and actual construction features is achieved. This provides a more comprehensive reflection of the technological essence of similar projects than a single feature, offering a more representative feature foundation for subsequent clustering. By pre-setting the number of clusters based on the clustering algorithm and the number of technological types, the historical fusion feature vectors with the highest similarity are grouped into multiple technological clusters. This accurately distinguishes the feature differences between different technological types in similar projects, avoiding the problem of insufficient generalization ability of prototype vectors caused by confusion of different technological features. By calculating the average value of the corresponding elements of all feature vectors in each cluster to generate technological feature prototype vectors, common feature benchmarks for each technological type are extracted, giving the prototype vectors the standard representation ability of similar technologies. The entire process involves "precise classification - compliant data matching - feature fusion - technological clustering - commonality extraction". The logical closed loop solves the problems of traditional process benchmarks lacking specificity, insufficient feature representativeness, and being disconnected from the semantics of the project to be reviewed. The generated process feature prototype vector can serve as a standard feature reference for similar compliant projects, providing a highly reliable comparison benchmark for the subsequent semantic fusion feature matrix and prototype vector Mahalanobis distance calculation. This significantly improves the accuracy and objectivity of anomaly judgment, ensuring that the review results are more in line with the actual process situation of the project.
[0052] Furthermore, if the Mahalanobis distance is greater than a preset anomaly tolerance threshold, it is determined that the current sub-item of the project to be reviewed has an anomaly; based on the location information of the empirical fragment features, an audit result is generated that includes the cost item corresponding to the sub-item with the anomaly and the corresponding construction period lacking empirical information, specifically:
[0053] The calculated Mahalanobis distance is compared with the preset anomaly tolerance threshold. If the Mahalanobis distance is less than or equal to the anomaly tolerance threshold, the current sub-item project is determined to be without anomalies, and an anomaly-free audit result is generated.
[0054] If the Mahalanobis distance is greater than the anomaly tolerance threshold, it is determined that there is an anomaly in the current sub-item project. The time interval and on-site location information corresponding to the empirical fragment features are extracted, and the first spatiotemporal image data and the first construction log text data within the time interval are obtained.
[0055] The cost information of the current sub-item of the project in the bill of quantities data is broken down into verifiable elements including materials, quantity and process.
[0056] Using the first spatiotemporal image data and the first construction log text data as an empirical dataset, we analyze whether there is supporting data in the empirical dataset corresponding to each verifiable element.
[0057] If a verifiable element lacks corresponding supporting data, the verifiable element is determined to be a cost item lacking empirical support, and the corresponding construction period is recorded.
[0058] Based on the verifiable elements lacking empirical support and the corresponding construction periods, an audit result is generated for the cost items and corresponding construction periods of the abnormal sub-projects that lack empirical information.
[0059] This application determines anomalies by comparing Mahalanobis distance with a preset anomaly tolerance threshold. Mahalanobis distance fully considers the feature distribution differences between the semantic fusion feature matrix and the process feature prototype vector, and compared with traditional single numerical comparison, it more accurately reflects the essential fit between the two types of features, ensuring the scientific nature and accuracy of anomaly determination. This avoids normal projects being misjudged as anomalies due to simple thresholds, and also prevents anomaly projects from being missed due to differences in feature distribution. When anomalies are determined to exist, the application extracts the time interval and on-site location information corresponding to the empirical segment features and obtains the first spatiotemporal image data and the first construction log text data within that interval, achieving precise locking of the empirical verification scope and avoiding... The inefficiency and excessive interference caused by comprehensive data retrieval provide a focused empirical foundation for subsequent precise verification. By breaking down the cost information of sub-items of the project into verifiable elements such as materials, quantities, and processes, the abstract cost declaration content is transformed into specific, verifiable core dimensions. This solves the problem of general cost item verification and lack of clear focus in traditional audits, making empirical matching more targeted. By using a defined empirical dataset as a basis and analyzing the supporting data for each verifiable element, specific cost items lacking empirical support can be accurately located, rather than simply determining that the project is abnormal. This achieves "abnormality identification - precise problem location." The process deepens the understanding of project cost audits. Finally, based on verifiable elements lacking empirical support and corresponding construction periods, audit results are generated. This not only clarifies abnormal projects and cost items but also clearly identifies the specific stages and periods lacking empirical evidence. This solves the problems of vagueness and lack of traceability in traditional audit results, providing a clear basis for subsequent rectification and review. The entire process, through a logical closed loop of "accurately identifying abnormalities - locking in the scope of empirical evidence - refining verification dimensions - accurately locating problems - generating clear results," significantly improves the accuracy, objectivity, and practicality of project cost audits, effectively avoiding the waste of funds caused by false or incorrect reporting, and strengthening the guidance and operability of audit results.
[0060] Furthermore, after determining that there is an anomaly in the current sub-item of the project, the process also includes:
[0061] The feature deviation vector is obtained by calculating the difference between the empirical fragment features and the process feature prototype vector in each dimension.
[0062] Identify the top N dimensions with the largest absolute values in the feature deviation vector, and retrieve the historical process features with the largest differences from the empirical fragment features in the top N dimensions from the historical compliance data used to generate the process feature prototype vector; where N is a positive integer;
[0063] Based on the retrieved historical construction details corresponding to the historical process characteristics, a cause analysis report for the anomaly is generated.
[0064] After identifying anomalies in sub-projects, this application calculates feature deviation vectors by subdividing empirical fragment features and process feature prototype vectors dimension by dimension. This allows for precise location of the specific differences between the process features of the project under review and the benchmark features of similar compliant projects in each dimension. This overcomes the limitation of traditional anomaly analysis, which can only determine the existence of anomalies but cannot identify the core dimensions of the differences, providing precise feature-level evidence for subsequent cause tracing. By identifying the top N dimensions with the largest absolute values in the feature deviation vectors, the application focuses on the most critical deviation directions, avoiding the obscuring of core issues caused by treating all deviation dimensions equally. This significantly improves the targeting and efficiency of cause analysis, eliminating the need to waste resources on irrelevant or secondary difference dimensions. Furthermore, by retrieving the top N dimensions from historical compliance data that generated the process feature prototype vectors... Historical process features that differ most significantly from empirical fragments in key dimensions are identified. By accurately comparing historical compliant process features with the features of the project under review, abstract feature deviations are transformed into specific differences in construction processes. Since each historical process feature corresponds to specific historical construction content, it can be directly linked to compliance standards in the actual construction process, solving the problem of feature deviations being disconnected from actual construction scenarios. Finally, an anomaly cause analysis report is generated based on the retrieved historical construction content. This ensures that the audit results not only include anomaly identification and problem location but also specific and traceable explanations of the causes of anomalies. This completely changes the traditional status quo of engineering cost audits, which "only judges anomalies without analyzing causes." It provides the auditee with a clear direction for rectification and a clear basis for audit review, further enhancing the depth, guidance, and practicality of engineering cost audits and effectively promoting the extension of audit work from "identifying problems" to "solving problems."
[0065] Secondly, this application provides an engineering cost auditing system based on multi-source data fusion. The engineering cost auditing system based on multi-source data fusion includes:
[0066] The acquisition module is used to acquire the bill of quantities data, spatiotemporal image data of the construction site, and construction log text data of the project to be reviewed;
[0067] The mapping module is used to input the bill of quantities data, spatiotemporal image data and construction log text data into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector and the spatiotemporal image data and the construction log text data into a physical reality feature vector.
[0068] The multimodal feature encoding model is based on multimodal data pairs of similar sub-projects of historical compliant projects, and is trained by minimizing the semantic distance between the historical cost declaration vector and the historical physical reality vector using a contrastive loss function.
[0069] The first calculation module is used to calculate the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time segment feature; and to extract the time segment feature with the highest score from the correlation score sequence as the empirical segment feature;
[0070] The fusion module is used to fuse the cost declaration feature vector with the empirical fragment features to generate a semantic fusion feature matrix;
[0071] The second calculation module is used to identify the sub-item engineering category to which the bill of quantities data belongs, obtain the process feature prototype vector generated by clustering based on historical compliance data under the sub-item engineering category to which the bill of quantities data belongs, and calculate the Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vector.
[0072] The review module is used to determine that there is an anomaly in the current sub-item of the project to be reviewed if the Mahalanobis distance is greater than a preset anomaly tolerance threshold; and to generate a review result that includes the cost item corresponding to the sub-item with the anomaly and the corresponding construction period lacking empirical information based on the location information of the empirical fragment features.
[0073] The acquisition module of this application specifically collects bill of quantities data, spatiotemporal image data of construction sites, and text data of construction logs, providing a comprehensive and complementary multi-source data foundation for review. This solves the problem of insufficient empirical support caused by the reliance on single data in traditional review, ensuring sufficient information sources for subsequent feature processing. The mapping module relies on a multimodal feature encoding model trained on multimodal data of similar sub-items of historical compliant projects. Through the training logic of minimizing the semantic distance of similar types of loss functions, it achieves accurate mapping of multimodal data to semantically consistent cost declaration feature vectors and physical reality feature vectors, breaking down the semantic barriers between different modal data (text, visual) and laying a comparable foundation for subsequent correlation analysis. The first calculation module calculates the temporal segment correlation scores of the two types of feature vectors and extracts the empirical segment features with the highest scores, accurately identifying the core construction reality data most relevant to the cost declaration. According to the data, the system effectively eliminates interference from irrelevant time-series segments, improving the relevance of empirical data and the efficiency of subsequent analysis. The fusion module merges the cost declaration feature vector with the empirical segment features to generate a semantic fusion feature matrix, further strengthening the semantic connection between the cost declaration and the core empirical data, making the feature representation more comprehensive and more identifiable. The second calculation module first identifies the categories of sub-items of the project and obtains the process feature prototype vector generated by clustering based on historical compliance data. Then, it calculates the distribution difference between the fusion feature matrix and the prototype vector through Mahalanobis distance. The sensitivity of Mahalanobis distance to feature distribution ensures the scientific nature of anomaly judgment and avoids the subjectivity and misjudgment risk of traditional single threshold comparison. The audit module judges anomalies based on the comparison results of Mahalanobis distance and anomaly tolerance threshold, and generates audit results containing abnormal cost items and corresponding construction periods by combining the location information of empirical segment features, realizing accurate anomaly judgment and clear source tracing of problems. Attached Figure Description
[0074] Figure 1: A schematic diagram of an embodiment of the engineering cost auditing method based on multi-source data fusion provided in this application;
[0075] Figure 2: A schematic diagram of an embodiment of the engineering cost auditing system based on multi-source data fusion provided in this application. Detailed Implementation
[0076] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0077] Example 1
[0078] Please refer to Figure 1. In order to solve the problem that the existing technology cannot accurately and efficiently audit the project cost, this embodiment of the invention provides a project cost auditing method based on multi-source data fusion, including steps S01-S06.
[0079] S01: Obtain the bill of quantities data, spatiotemporal image data of the construction site, and text data of the construction log for the project to be reviewed.
[0080] In a preferred embodiment of this invention, the acquisition of the bill of quantities data, spatiotemporal image data of the construction site, and construction log text data of the project to be reviewed specifically includes:
[0081] For bill of quantities data, it can be extracted through the project management file system of the project to be reviewed. This file system needs to store the official bill of quantities jointly confirmed by the construction unit, the construction unit and the supervision unit in advance. When extracting, it is necessary to filter out all cost-related data corresponding to the current review scope according to the current industry standard for classifying sub-items of the project. The extracted bill of quantities data should clearly include the specific name of each sub-item (such as "C30 concrete pouring" "reinforcement binding", etc.), the declared quantity (the unit must be consistent with the industry standard, such as cubic meters, tons, square meters, etc.), the declared unit price and the declared total price, to ensure that the core information of each cost item is not omitted. For spatiotemporal image data of the construction site, high-definition camera equipment with GPS positioning function needs to be deployed in the pre-set construction area of the construction site (including key areas such as the main structure construction area, decoration and finishing construction area, and material storage area). The equipment resolution should be no less than 1080P, and it should support automatic image acquisition at fixed time intervals. The acquisition interval can be flexibly set according to the construction progress, usually 30 minutes / time. During the acquisition process, the equipment needs to synchronously record the specific time (accurate to the second) and the on-site location (latitude and longitude coordinates) of each acquisition operation. After the acquisition is completed, the image data is transmitted to the back-end server through a wireless transmission module. The server will then associate and integrate the image data with spatiotemporal information in chronological order to form a structured spatiotemporal image dataset, ensuring that each frame of the image can accurately correspond to the construction period and site location. For construction log text data, it is necessary to collect the original construction logs submitted daily by the construction unit. The original logs can be electronic ledgers or scanned digital text. After collection, they need to be preprocessed according to unified standardized rules. First, the text is segmented to remove redundant information such as meaningless auxiliary words and modal particles. Then, the core elements are extracted, including the date of construction, the actual construction content (the specific procedures of the corresponding sub-items of the project must be clearly defined), the number of personnel involved in the construction (statistics are categorized by trade, such as carpenters, bricklayers, etc.), and the actual amount of construction materials used (the material model, specifications, and specific usage values must be noted). Finally, the construction log text data with a unified format and complete information is organized to provide high-quality input data for subsequent multimodal feature coding.
[0082] S02: Input the bill of quantities data, spatiotemporal image data and construction log text data into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector, and maps the spatiotemporal image data and the construction log text data into a physical reality feature vector;
[0083] The multimodal feature encoding model is based on multimodal data pairs of similar sub-projects from historical compliant projects, and is trained by minimizing the semantic distance between the historical cost declaration vector and the historical physical reality vector using a contrastive loss function.
[0084] In a preferred embodiment of this example, the step of inputting the bill of quantities data, spatiotemporal image data, and construction log text data into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector, and maps the spatiotemporal image data and construction log text data into a physical reality feature vector, specifically:
[0085] In the training phase of the multimodal feature coding model, 150 historical engineering project archives that have been completed and accepted with compliant audit conclusions are first collected. Based on the classification standards for sub-items in the "Construction Engineering Quantity List Pricing Specification," the historical project archives are divided into multiple categories such as building construction, decoration and renovation, and installation engineering. From each category, corresponding historical quantity list data, historical spatiotemporal image data, and historical construction log text data are extracted. Each set of data is combined one-to-one to form multimodal data pairs for the same type of sub-item, ensuring the category relevance and compliance benchmark of the training data. Subsequently, an initial multimodal feature coding model is constructed. This model includes independent text coding and visual coding modules. The text coding module can adopt a BERT-based pre-trained model architecture; the visual coding module can adopt a ResNet-50 model architecture. The output dimension of the last fully connected layer is adjusted to 2048. The model initialization parameters are set as follows: learning rate of 0.0001 for the text coding module, learning rate of 0.0002 for the visual coding module, dropout probability of 0.15, and weight decay coefficient of 0.001. After inputting the aforementioned multimodal data into the initial model, the text encoding module performs word segmentation on the historical bill of quantities data (using an industry-specific dictionary containing engineering terms, material names, and process names). The segmentation results are converted into word vectors via an embedding layer, and then the Transformer encoder extracts contextual semantic features, ultimately outputting a historical cost declaration feature vector with a dimension of 768. The visual encoding module performs preprocessing on the historical spatiotemporal image data, including normalization (pixel values normalized to the [0,1] interval) and random cropping. Multi-scale visual features are extracted through convolutional and pooling layers. Simultaneously, the timestamps corresponding to the images are converted into time-series features (normalized to the [0,1] interval based on Unix timestamps), and latitude and longitude coordinates are converted into spatial vectors (normalized after Gaussian projection). The visual and spatiotemporal features are fused and input into a fully connected layer, outputting a historical image feature vector with a dimension of 2048. The text encoding module simultaneously performs feature extraction and vector transformation on the preprocessed historical construction log text data (word segmentation and stop word removal), outputting a feature vector with a dimension of 768. The historical log feature vector is obtained by horizontally concatenating the historical image feature vector with the historical log feature vector to obtain a historical physical reality feature vector with a dimension of 2816.
[0086] When calculating semantic distance using the contrastive loss function, for each set of historical cost statement feature vectors, cosine distance is used to calculate the first semantic distance (positive sample pair distance, representing the semantic fit of true association) between it and the historical physical condition feature vectors in the same set. Simultaneously, the second semantic distance (negative sample pair distance, representing the semantic difference of no association) is calculated between this historical cost statement feature vector and all other sets of historical physical condition feature vectors. The smallest second semantic distance is selected as the contrastive distance (i.e., the negative sample distance most easily confused with positive samples, ensuring that contrastive learning focuses on the most difficult-to-distinguish sample pairs and improves the model's discriminative ability). To achieve the contrastive learning objective of "positive sample aggregation and negative sample separation," a preset marginal value of 0.35 is used (this value is optimized by 10...). The pre-training validation group is used to determine the optimal semantic discrimination threshold for multimodal data in the field of engineering cost (which can balance the generalization and accuracy of the model). The difference between the contrast distance and the first semantic distance is calculated: if the difference is less than 0.35, it means that the semantic discrimination between the positive sample pair and the closest negative sample pair has not met expectations. In this case, the difference should be subtracted from 0.35 as a single loss component to force the model to reduce the distance between positive samples and increase the distance between negative samples. If the difference is greater than or equal to 0.35, it means that the semantic discrimination between positive and negative samples has met the requirements. The single loss component is recorded as zero to avoid overfitting caused by overtraining. The single loss components corresponding to all historical cost declaration vectors are summed to obtain the global contrast loss value. The stochastic gradient descent (SGD) algorithm is used to adjust the parameters of the text encoding module and the visual encoding module based on the backpropagation of the global contrast loss value. Each iteration processes 64 multimodal data pairs in batches. The training is stopped and the model parameters are saved when the global contrast loss value drops below 0.008 and the fluctuation of the loss value does not exceed 0.0005 in 8 consecutive iterations. The innovative value of contrastive learning: Compared with traditional single-modal coding or unsupervised coding methods, the contrastive learning in this application maps cost declaration data and construction reality data to the same semantic space through the training logic of "positive sample alignment and negative sample differentiation". This solves the semantic gap between text-based cost data and visual / text-based reality data, so that the output feature vector retains the core semantics of cost declaration and has accurate comparability with construction reality, laying a highly reliable foundation for subsequent correlation calculation and anomaly detection.
[0087] In the data mapping stage, the preprocessed bill of quantities data to be reviewed (with redundant remarks removed and core cost items retained) is input into the text encoding module of the trained multimodal feature encoding model. After word segmentation, word embedding, and semantic feature extraction, it is transformed into a cost declaration feature vector with a dimension of 768 and output. The spatiotemporal image data to be reviewed is normalized and features are extracted according to the preprocessing standards of the training stage. After fusing the corresponding spatiotemporal information, an image feature vector with a dimension of 2048 is output. The construction log text data is segmented and stop words (such as words without substantial meaning, such as "today" and "in progress") are removed. The text is then input into the text encoding module to extract semantic features and transformed into a log feature vector with a dimension of 768. The image feature vector and the log feature vector are horizontally concatenated to obtain a physical reality feature vector with a dimension of 2816. This achieves accurate mapping of multi-source heterogeneous data to a unified semantic space feature vector, providing a reliable feature foundation for subsequent correlation calculations and anomaly detection.
[0088] S03: Calculate the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time segment feature; extract the time segment feature with the highest score from the correlation score sequence as the empirical segment feature.
[0089] In a preferred embodiment of this invention, the step of calculating the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time-series segment feature, and extracting the time-series segment feature with the highest score from the correlation score sequence as the empirical segment feature, specifically involves:
[0090] First, based on the time span of the spatiotemporal image data and construction log text data corresponding to the physical reality feature vector, the physical reality feature vector is divided into multiple time-series segment features of equal length in chronological order. The segmentation duration needs to be adapted to the data collection interval (e.g., 30 minutes / time) and the construction process cycle to ensure that each time-series segment can fully cover the key construction stages of a single process. For example, if each time-series segment feature is set to correspond to a 1-hour construction period, and the total construction time of the sub-item project to be reviewed is 24 hours, then the physical reality feature vector will be divided into 24 time-series segment features. Each segment corresponds to a unique time interval (accurate to the minute, such as "2024-05-20 08:00-09:00"), and the correlation between each time-series segment feature and the construction site location information is recorded simultaneously to ensure that the segment can be traced back to the specific construction area.
[0091] Subsequently, the correlation score is calculated. Since the dimension of the cost declaration feature vector is 768, and the dimensions of the physical reality feature vector and each time segment feature are both 2816, the cost declaration feature vector needs to be expanded to 2816 dimensions using zero-padding to ensure that the dimensions of the two are consistent. For each time segment feature, the correlation score is calculated by summing the element-wise multiplications: that is, iterating through each corresponding dimension of the cost declaration feature vector and the time segment feature, multiplying the two vector values of dimension i to obtain the product, and then summing the product results of all dimensions to obtain the correlation score between the time segment feature and the cost declaration feature vector. The higher the score, the higher the semantic fit between the construction reality and the cost declaration content during that period.
[0092] After calculating the correlation scores for all time-series segment features, the correlation scores are arranged sequentially according to the construction time order corresponding to the time-series segments, forming a complete correlation score sequence (e.g., [0.62, 0.78, 0.85, ..., 0.71]). Finally, the score with the largest value in this score sequence is retrieved, and its corresponding time-series segment feature is determined as the empirical segment feature. If multiple time-series segment features have the same correlation score and are all the maximum value, the time-series segment feature with the earliest time interval is selected as the empirical segment feature. If it is necessary to prioritize matching the time periods corresponding to key processes, process priority weights can be pre-configured, and filtering can be performed according to the weights to ensure the timeliness and relevance of the empirical data, providing core construction data support most relevant to the cost declaration for subsequent feature fusion and anomaly detection.
[0093] S04: The cost declaration feature vector and the empirical fragment feature are fused to generate a semantic fusion feature matrix.
[0094] In a preferred embodiment of this invention, the step of fusing the cost declaration feature vector with the empirical fragment features to generate a semantic fusion feature matrix specifically involves:
[0095] Feature fusion needs to be based on the semantic correlation between the cost declaration feature vector and the empirical fragment features. A semantic fusion feature matrix containing the core information of both is generated through standardized calculations, ensuring that the fused data retains both the key semantics of the cost declaration and the construction reality features of the empirical fragments. First, the dimensional parameters of the two types of features are defined: the cost declaration feature vector output by the multimodal feature coding model has a dimension of 768, and the empirical fragment features (i.e., the core temporal fragments of the physical reality feature vector) have a dimension of 2816. Before fusion, the two types of feature vectors need to be normalized, mapping all element values to the [0,1] interval to avoid imbalance in fusion weights due to differences in dimensional value ranges. Then, fusion weights are set based on feature importance analysis: considering that the cost declaration feature vector is the core benchmark for cost review, it is assigned a weight of 0.4; the empirical fragment features are a direct representation of the construction reality, and are assigned a weight of 0.6. The weight values can be flexibly adjusted according to the review focus of different sub-items of the project, but the sum of the weights of the two types of features must be 1.
[0096] The fusion operation employs a combination of weighted element-wise summation and dimensional expansion: first, the normalized cost declaration feature vector is scaled with a weight of 0.4, and the empirical fragment feature vector is scaled with a weight of 0.6. Then, the two scaled vectors are summed element-wise to obtain a 2816-dimensional fusion feature vector (where the cost declaration feature vector is expanded to 2816 dimensions through zero-padding before participating in the operation; the zero-padding values do not affect the effective information of the empirical fragment features). To form a semantic fusion feature matrix, this 2816-dimensional fusion feature vector needs to be further divided according to semantic categories: based on the feature output rules of the multimodal feature coding model, the first 768 dimensions of the fusion feature vector are divided into "cost declaration semantic sub-vectors," corresponding to the core declaration information of the cost item; the middle 1024 dimensions are divided into "construction scene visual sub-vectors," corresponding to the visual features of construction scenes, materials, and equipment in the empirical fragments; and the last 1024 dimensions are divided into "construction scene text sub-vectors," corresponding to the core information of the construction logs in the empirical fragments. Finally, the three sub-vectors are stacked row by row to form a 3×2816-dimensional semantic fusion feature matrix. Each row of the matrix corresponds to a type of core semantic information, and each column corresponds to a specific value of the feature dimension. This matrix not only fully preserves the semantic relationship between the cost statement and the actual construction situation, but also provides structured input data for the subsequent Mahalanobis distance calculation with the prototype vector of the process feature through the matrix structure, ensuring the accuracy of subsequent anomaly detection.
[0097] S05: Identify the sub-item engineering category to which the bill of quantities data belongs, obtain the process feature prototype vector generated by clustering based on historical compliance data under the sub-item engineering category to which the bill of quantities data belongs, and calculate the Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vector.
[0098] In a preferred embodiment of this invention, the step of identifying the sub-item engineering category to which the bill of quantities data belongs, obtaining the process feature prototype vector generated based on historical compliance data clustering under the sub-item engineering category to which the bill of quantities data belongs, and calculating the Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vector, specifically involves:
[0099] First, the category identification of each sub-item of the project is carried out: the category identification information of each sub-item of the project is filtered out from the extracted bill of quantities data. This identification information can be determined by the first four digits of the bill of quantities code (according to the "Construction Engineering Bill of Quantities Pricing Specification"). For example, the category identification corresponding to the bill of quantities code "010401" is "0104". Combined with the preset engineering category classification standard, the category of the sub-item of the data to be reviewed is determined to be "concrete and reinforced concrete engineering", ensuring that the category identification is consistent with the industry standard.
[0100] Subsequently, the prototype vector of the process features was obtained: At least 50 sets of multimodal data pairs (no fewer than 50 sets) of similar sub-items from all historical compliant projects under the category of "Concrete and Reinforced Concrete Engineering" were collected. Each data pair was input into a pre-trained multimodal feature encoding model, which outputs the corresponding historical cost declaration feature vector (768 dimensions) and historical physical condition feature vector (2816 dimensions). Element-wise weighted summation was performed on each set of historical cost declaration and historical physical condition feature vectors, with the weight of the historical cost declaration feature vector set to 0.3 and the weight of the historical physical condition feature vector set to 0.7 (the weights can be adjusted according to the project type, the core being to highlight the importance of construction condition features), resulting in a 2816-dimensional historical fusion feature vector for each data pair. The K-Means clustering algorithm was used, and the number of clusters was preset based on the actual number of process types under the category (e.g., "Concrete and Reinforced Concrete Engineering" includes C30 concrete pouring, C40 concrete pouring, and rebar tying). For each core process (with a cluster size of 3), all historical fusion feature vectors are clustered according to similarity to form 3 process clusters, each corresponding to a specific process type. The average value of the corresponding elements of all historical fusion feature vectors in each cluster is calculated. For example, the average value of the i-th dimension element of all vectors in cluster 1 is x_i. The average values of all dimensions are combined in order to obtain the 2816-dimensional process feature prototype vector corresponding to the cluster. Finally, 3 process feature prototype vectors corresponding to different process types are generated.
[0101] Finally, the Mahalanobis distance is calculated: First, the semantic fusion feature matrix (3×2816 dimensions) is preprocessed by averaging the columns to transform it into a 2816-dimensional feature vector (ensuring consistency with the dimension of the process feature prototype vector); all historical fusion feature vectors under this category are collected, and their covariance matrix and inverse matrix are calculated (the covariance matrix dimension is 2816×2816, implemented using the cov and inv functions in Python's numpy library); let the preprocessed semantic fusion feature vector be X, a certain process feature prototype vector be μ, and the inverse matrix of the covariance matrix be Σ. -1 The Mahalanobis distance between the semantic fusion feature vector and the prototype feature vector of the process is calculated using the Mahalanobis distance formula. If the project to be audited has multiple prototype vectors corresponding to the process type (e.g., 3), then calculate the Mahalanobis distance with each prototype vector separately, and take the minimum value as the final Mahalanobis distance result to ensure that the distance calculation can reflect the characteristic differences between the project to be audited and the closest compliant process.
[0102] S06: If the Mahalanobis distance is greater than the preset anomaly tolerance threshold, it is determined that the current sub-item of the project to be reviewed has an anomaly; based on the location information of the empirical fragment features, an audit result is generated that includes the cost item corresponding to the sub-item with the anomaly and the corresponding construction period lacking empirical information.
[0103] In a preferred embodiment of this example, if the Mahalanobis distance is greater than a preset anomaly tolerance threshold, it is determined that the current sub-item of the project to be audited has an anomaly; based on the location information of the empirical fragment features, an audit result is generated that includes the cost item corresponding to the sub-item with the anomaly and the corresponding construction period lacking empirical information, specifically:
[0104] First, the determination of the preset anomaly tolerance threshold needs to be statistically calibrated based on historical compliant project data: collect semantic fusion feature vectors and process feature prototype vectors of more than 100 historical compliant projects under the corresponding sub-item engineering category, calculate the Mahalanobis distance of each group and statistically analyze the distribution range, and take the 95th percentile of the distribution as the anomaly tolerance threshold (for example, the anomaly tolerance threshold for the "concrete and reinforced concrete engineering" category is set to 2.8), to ensure that the threshold can both cover the feature differences of normal projects and effectively identify abnormal situations that exceed the compliance range.
[0105] The calculated Mahalanobis distance is compared with the preset anomaly tolerance threshold: if the Mahalanobis distance is less than or equal to 2.8, it indicates that the semantic fusion features of the project under review are within a reasonable range from the process feature benchmark of similar compliant projects, and the current sub-item project is determined to be without anomalies, generating an audit result without anomalies that includes the project name, sub-item category, audit conclusion (no anomalies), and audit time; if the Mahalanobis distance is greater than 2.8, it is determined that the current sub-item project has anomalies, at which point the location information corresponding to the empirical fragment features needs to be extracted, including the precise time interval (e.g., "2024-05-20 10:00-11:00") and the on-site location coordinates (e.g., "116.35°E, 39.92°N"), and based on this location information, the first spatiotemporal image data (high-definition continuous image sequence) and the first construction log text data (the original record and preprocessed structured data corresponding to this time period) of the corresponding time interval and construction area are retrieved from the background database.
[0106] Subsequently, the cost information of the current sub-item of the bill of quantities is broken down into three verifiable elements according to a unified standard: materials, quantity, and process. For example, the cost item of "C30 concrete pouring" is broken down into: materials (C30 commercial concrete), quantity (120 cubic meters), and process (formwork → pouring → vibration → curing). The retrieved first spatiotemporal image data and first construction log text data were used as empirical datasets. For each verifiable element, the existence of corresponding supporting data was analyzed one by one: For material elements, it was checked whether there were continuous images of C30 concrete transport vehicles entering and unloading in the spatiotemporal images, and whether there were material arrival acceptance records and usage statistics in the construction log; For quantity elements, the actual usage was estimated by combining the construction area area and component size in the images with industry loss standards, and compared with the declared quantity. At the same time, quantitative data such as material requisition forms and pouring records in the construction log were checked; For process elements, it was checked whether there were construction images of each stage of formwork, pouring, vibration, and curing in the spatiotemporal images, and whether there were start and end times of each process and records of signature confirmation by construction personnel in the construction log.
[0107] If the analysis reveals that a verifiable element lacks supporting data (e.g., the declared 120 cubic meters of C30 concrete lacks corresponding arrival images and usage statistics), then this verifiable element is identified as a cost item lacking empirical support, and its corresponding construction period is recorded simultaneously (e.g., "2024-05-20 10:00-11:00"). If multiple verifiable elements lack supporting data, each element and its corresponding time period are recorded separately. Finally, a standardized audit result is generated. This result must clearly include the name of the project to be audited, the category and name of the abnormal sub-items, the specific content of each abnormal cost item (e.g., "C30 concrete pouring"), the corresponding verifiable element (e.g., "quantity: 120 cubic meters"), the type of lacking empirical evidence (e.g., "no material arrival images, no usage statistics"), and the specific construction period, ensuring that the audit result is clear and traceable, providing a clear basis for subsequent rectification, review, and responsibility determination.
[0108] Furthermore, after determining that there are anomalies in the current sub-item of the project, this application needs to further analyze the causes of the anomalies. Specifically, the following steps are taken: First, the empirical fragment features and the corresponding process feature prototype vectors are compared dimension by dimension to obtain the feature deviation vector. For example, if the value of the i-th dimension of the empirical fragment feature is a_i and the value of the i-th dimension of the process feature prototype vector is b_i, then the deviation value of this dimension is |a_i-b_i|. The deviation values of all dimensions are combined in order to form the feature deviation vector. Then, N=5 (N can be adjusted to a positive integer between 3 and 10 according to the complexity of the project) to identify the top 5 dimensions with the largest absolute values in the feature deviation vector and lock them as the core deviation dimensions. Then, these 5 dimensions are retrieved from the historical compliance data that generated the process feature prototype vectors. The historical process features that differ most significantly from the empirical segment features under the core deviation dimension are identified. For example, if the core deviation dimension is "concrete pouring thickness" or "rebar spacing," then the historical process features with the most significant differences between the numerical range of that dimension and the empirical segment features are retrieved from the historical compliance data (e.g., the historical compliance process requires a rebar spacing of 15cm, while the corresponding feature value of the empirical segment is 25cm). Finally, the historical construction content corresponding to the retrieved historical process features (e.g., "C30 concrete pouring thickness must be ≥10cm, rebar spacing must be controlled between 12-15cm") is compared with the empirical segment features of the project to be audited to clarify the cause of the anomaly (e.g., "concrete pouring thickness is only 8cm, rebar spacing is 25cm, which does not meet the process standards of similar compliant projects"). Based on this, a structured anomaly cause analysis report is generated. The report must include the core deviation dimension, historical compliance process standards, actual characteristics of the project to be audited, conclusions on the cause of the anomaly, and rectification suggestions.
[0109] In summary, this application overcomes the limitations of traditional auditing methods that rely on single data sources by acquiring three types of multi-source heterogeneous data: bill of quantities data, spatiotemporal imagery data of construction sites, and textual construction log data. This provides a comprehensive and complementary empirical foundation for comparing cost declarations with actual construction conditions. By inputting multi-source data into a multi-modal feature encoding model trained on multi-modal data of similar sub-items from historical compliant projects, and by employing a contrastive loss function to minimize the semantic distance between historical cost declaration vectors and historical physical reality vectors, different modalities of data can be mapped to semantically consistent cost declaration feature vectors and physical reality feature vectors. This effectively solves the problems of insufficient multi-source data fusion and inadequate semantic association mining, significantly improving the accuracy of feature matching. Finally, by calculating and extracting the correlation scores of cost declaration feature vectors and physical reality feature vectors for each time-series segment, the application demonstrates its effectiveness in addressing these issues. The highest-quality empirical fragment features enable precise location of the construction data fragments most relevant to the cost declaration, avoiding audit bias caused by the mismatch between empirical data and the time of cost items. By fusing the cost declaration feature vector with the empirical fragment features to generate a semantic fusion feature matrix, the semantic association between the two is further strengthened, providing more comprehensive feature support for anomaly detection. By identifying the categories of sub-items and obtaining the process feature prototype vectors generated by clustering historical compliance data, and combining the sensitivity of Mahalanobis distance to feature distribution differences for anomaly detection, the reliance on manual experience is eliminated, achieving objective quantitative identification of anomalies. Finally, based on the location information of empirical fragment features, audit results containing abnormal cost items and corresponding construction periods are generated, accurately tracing specific links lacking empirical support, effectively solving the problem that existing technologies cannot accurately and efficiently audit engineering costs.
[0110] Example 2
[0111] Please refer to Figure 2, which shows an engineering cost audit system based on multi-source data fusion provided in an embodiment of this application.
[0112] In this embodiment, the engineering cost auditing system based on multi-source data fusion includes an acquisition module 10, a mapping module 20, a first calculation module 30, a fusion module 40, a second calculation module 50, and an auditing module 60.
[0113] Module 10 is used to acquire the bill of quantities data, spatiotemporal image data of the construction site, and text data of the construction log of the project to be reviewed.
[0114] The mapping module 20 is used to input the bill of quantities data, spatiotemporal image data and construction log text data into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector and maps the spatiotemporal image data and the construction log text data into a physical reality feature vector.
[0115] The multimodal feature encoding model is based on multimodal data pairs of similar sub-projects of historical compliant projects, and is trained by minimizing the semantic distance between the historical cost declaration vector and the historical physical reality vector using a contrastive loss function.
[0116] The first calculation module 30 is used to calculate the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time segment feature; and to extract the time segment feature with the highest score from the correlation score sequence as the empirical segment feature;
[0117] The fusion module 40 is used to fuse the cost declaration feature vector with the empirical fragment features to generate a semantic fusion feature matrix;
[0118] The second calculation module 50 is used to identify the sub-item engineering category to which the bill of quantities data belongs, obtain the process feature prototype vector generated by clustering based on historical compliance data under the sub-item engineering category to which the bill of quantities data belongs, and calculate the Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vector.
[0119] The audit module 60 is used to determine that there is an anomaly in the current sub-item of the project to be audited if the Mahalanobis distance is greater than a preset anomaly tolerance threshold; and to generate an audit result that includes the cost item corresponding to the sub-item with the anomaly and the corresponding construction period lacking empirical information based on the location information of the empirical fragment features.
[0120] For ease of description and brevity, the system embodiments of the present invention include all the implementation methods described in the above embodiments of the engineering cost audit method based on multi-source data fusion, and will not be repeated here.
[0121] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for engineering cost auditing based on multi-source data fusion, characterized in that, include: The process involves acquiring the bill of quantities data, spatiotemporal image data of the construction site, and construction log text data of the project to be reviewed; inputting these data into a pre-defined multimodal feature encoding model, which maps the bill of quantities data into a cost declaration feature vector and the spatiotemporal image data and construction log text data into a physical reality feature vector; wherein the multimodal feature encoding model is trained based on multimodal data pairs of similar sub-items of historical compliant projects, and by minimizing the semantic distance between historical cost declaration feature vectors and historical physical reality feature vectors using a contrastive loss function; calculating the correlation score sequence between each temporal segment feature of the cost declaration feature vector and the physical reality feature vector; extracting the temporal segment feature with the highest score from the correlation score sequence as the empirical segment feature; and fusing the cost declaration feature vector and the empirical segment feature to generate a semantic fusion feature matrix. The engineering quantity list data is identified to its respective category, and process feature prototype vectors generated based on historical compliance data clustering are obtained for each category. The Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vectors is calculated. If the Mahalanobis distance is greater than a preset anomaly tolerance threshold, it is determined that the current engineering quantity list of the project to be reviewed has an anomaly. Based on the location information of the empirical fragment features, an audit result is generated that includes the cost item corresponding to the abnormal engineering quantity list and the corresponding construction period lacking empirical information.
2. The engineering cost auditing method based on multi-source data fusion according to claim 1, characterized in that, The acquisition of the bill of quantities data for the project to be reviewed, as well as the spatiotemporal image data and construction log text data of the construction site as empirical evidence, specifically involves: extracting the bill of quantities data from the project management file of the project to be reviewed, wherein the bill of quantities data includes the specific name, declared quantity, declared unit price, and declared total price of each sub-item; using camera equipment with positioning function deployed in a preset construction area of the construction site, collecting image data of the construction site at fixed time intervals, and simultaneously recording the specific time and location of the equipment for each collection operation, and integrating them to obtain the spatiotemporal image data of the construction site; and acquiring the daily construction log text data submitted by the construction site, wherein the construction log text data includes the date of construction, the actual construction content, the number of personnel involved in the construction, and the actual amount of construction materials used.
3. The engineering cost auditing method based on multi-source data fusion according to claim 1, characterized in that, The process of inputting the bill of quantities data, spatiotemporal image data, and construction log text data into a preset multimodal feature coding model, so that the multimodal feature coding model maps the bill of quantities data into a cost declaration feature vector and the spatiotemporal image data and the construction log text data into a physical reality feature vector, specifically involves: inputting the bill of quantities data into the text coding module of the multimodal feature coding model, so that the text coding module extracts and converts the bill of quantities data into vector form, and outputs a cost declaration feature vector; inputting the spatiotemporal image data into the visual coding module of the multimodal feature coding model, so that the visual coding module extracts construction scene features, material features, and equipment features from the spatiotemporal image data, while incorporating the spatiotemporal information corresponding to the image, performing vector transformation and fusion on the extracted features to obtain the feature vector corresponding to the image; performing word segmentation on the construction log text data; inputting the processed construction log text data into the text coding module of the multimodal feature coding model, so that the text coding module extracts features and performs vector transformation on the processed log text data to obtain the feature vector corresponding to the log; and concatenating and fusing the feature vector corresponding to the image and the feature vector corresponding to the log to obtain the physical reality feature vector.
4. The engineering cost auditing method based on multi-source data fusion according to claim 1, characterized in that, The multimodal feature encoding model is based on multimodal data pairs of similar sub-items from historical compliant projects, and is trained by minimizing the semantic distance between historical cost declaration feature vectors and historical physical condition feature vectors using a contrastive loss function. Specifically, it involves: acquiring project files from multiple historical compliant projects; classifying these files according to sub-item categories; extracting corresponding historical bill of quantities data, historical spatiotemporal image data, and historical construction log text data from each category; combining these to form multimodal data pairs of similar sub-items; constructing an initial multimodal feature encoding model based on preset text encoding and visual encoding modules, combined with preset initialization parameters; and inputting the multimodal data pairs into the initial multimodal feature encoding model to obtain the multimodal data pairs... The system generates historical cost declaration feature vectors corresponding to historical bill of quantities data, and historical physical reality feature vectors fused from historical construction site spatiotemporal image data and historical construction log text data. Based on the contrastive loss function, it calculates each first semantic distance between each group of historical cost declaration feature vectors and the corresponding historical physical reality feature vectors of the same group, and calculates each second semantic distance between each group of historical cost declaration feature vectors and historical physical reality feature vectors of different groups. Based on each first semantic distance and each second semantic distance, it calculates the contrastive loss value. Based on the contrastive loss value, it adjusts the parameters of the text encoding module and the visual encoding module in the initial multimodal feature encoding model until the contrastive loss value drops to a preset convergence threshold, thus obtaining the multimodal feature encoding model.
5. The engineering cost auditing method based on multi-source data fusion according to claim 4, characterized in that, The step of calculating the contrast loss value based on each first semantic distance and each second semantic distance specifically involves: selecting the second semantic distance with the smallest value from multiple second semantic distances as the contrast distance; calculating the difference between the contrast distance and the first semantic distance; and comparing the difference with a preset marginal value. If the difference is less than the marginal value, the marginal value minus the difference is taken as a single loss component; if the difference is greater than or equal to the marginal value, the single loss component is recorded as zero; the single loss components corresponding to all historical cost declaration feature vectors are summed to obtain the comparative loss value.
6. The engineering cost auditing method based on multi-source data fusion according to claim 1, characterized in that, The calculation of the correlation score sequence between the cost declaration feature vector and the physical reality feature vector for each time segment feature; extracting the time segment feature with the highest score from the correlation score sequence as the empirical segment feature, specifically involves: dividing the physical reality feature vector into multiple time segment features of equal length according to time order, and recording the time interval corresponding to each time segment feature of equal length; performing element-wise multiplication and summation on the cost declaration feature vector and each time segment feature respectively to obtain the correlation score between each time segment feature and the cost declaration feature vector; arranging all correlation scores sequentially according to the time order of the time segments to form a correlation score sequence; obtaining the score with the largest value in the correlation score sequence, and determining the time segment feature corresponding to the largest score as the empirical segment feature.
7. The engineering cost auditing method based on multi-source data fusion according to claim 1, characterized in that, The step of identifying the sub-item engineering category to which the bill of quantities data belongs and obtaining the process feature prototype vector generated based on historical compliance data clustering under the sub-item engineering category to which the bill of quantities data belongs specifically involves: obtaining the category identification information of each sub-item engineering in the bill of quantities data, and determining the sub-item engineering category to which the data to be reviewed belongs by combining the preset engineering category classification standard. Obtain multimodal data pairs of similar historical compliant sub-items under the sub-item engineering category to which the data to be audited belongs. Input each multimodal data pair into a multimodal feature encoding model to obtain the corresponding historical cost declaration feature vector and historical physical condition feature vector. Perform element-wise weighted summation on each set of historical cost declaration feature vectors and historical physical condition feature vectors to obtain the historical fusion feature vector corresponding to each set of multimodal data pairs. Based on a clustering algorithm and a preset number of clusters, combine all historical fusion feature vectors and group the historical fusion feature vectors with the highest similarity into one class to obtain multiple process clusters. The number of clusters is the number of process types. Calculate the average value of each element corresponding to all historical fusion feature vectors in each cluster, and use the vector composed of the average values as the process feature prototype vector corresponding to each cluster.
8. The engineering cost auditing method based on multi-source data fusion according to claim 1, characterized in that, If the Mahalanobis distance is greater than a preset anomaly tolerance threshold, it is determined that the current sub-item of the project to be reviewed has an anomaly; based on the location information of the empirical fragment features, an audit result is generated that includes the cost item corresponding to the sub-item with anomaly and the corresponding construction period lacking empirical information. Specifically, the calculated Mahalanobis distance is compared with the preset anomaly tolerance threshold. If the Mahalanobis distance is less than or equal to the anomaly tolerance threshold, it is determined that the current sub-item has no anomaly, and an audit result without anomalies is generated. If the Mahalanobis distance is greater than the anomaly tolerance threshold, an anomaly is determined to exist in the current sub-item project. The time interval and on-site location information corresponding to the empirical segment features are extracted, and the first spatiotemporal image data and first construction log text data within the time interval are obtained. The cost item information of the current sub-item project in the bill of quantities data is decomposed into verifiable elements including materials, quantities, and processes. The first spatiotemporal image data and the first construction log text data are used as an empirical dataset to analyze whether there is supporting data corresponding to each verifiable element. If a verifiable element lacks corresponding supporting data, it is determined that the verifiable element is a cost item lacking empirical support, and the corresponding construction period is recorded. Based on the verifiable element lacking empirical support and the corresponding construction period, an audit result is generated for the cost item corresponding to the anomaly sub-item project and the corresponding construction period lacking empirical information.
9. The engineering cost auditing method based on multi-source data fusion according to claim 8, characterized in that, After determining that there is an anomaly in the current sub-item project, the method further includes: calculating the difference between the empirical fragment feature and the process feature prototype vector dimension by dimension to obtain a feature deviation vector; identifying the top N dimensions with the largest absolute value in the feature deviation vector; retrieving the historical process feature with the largest difference from the empirical fragment feature in the top N dimensions from the historical compliance data used to generate the process feature prototype vector; where N is a positive integer; and generating a cause analysis report of the anomaly based on the historical construction content corresponding to the retrieved historical process feature.
10. An engineering cost auditing system based on multi-source data fusion, characterized in that, include: The module includes an acquisition module for acquiring bill of quantities data, spatiotemporal image data of the construction site, and construction log text data of the project to be reviewed; a mapping module for inputting the bill of quantities data, spatiotemporal image data, and construction log text data into a preset multimodal feature encoding model, so that the multimodal feature encoding model maps the bill of quantities data into a cost declaration feature vector and the spatiotemporal image data and the construction log text data into a physical reality feature vector; wherein, the multimodal feature encoding model is trained based on multimodal data pairs of similar sub-items of historical compliant projects and by minimizing the semantic distance between historical cost declaration feature vectors and historical physical reality feature vectors using a contrastive loss function; and a first calculation module for calculating the correlation score sequence between each temporal segment feature of the cost declaration feature vector and the physical reality feature vector. The system extracts the time-series segment features corresponding to the highest scores from the correlation score sequence as empirical segment features; a fusion module is used to fuse the cost declaration feature vector with the empirical segment features to generate a semantic fusion feature matrix; a second calculation module is used to identify the sub-item engineering category to which the bill of quantities data belongs, obtain the process feature prototype vector generated by clustering based on historical compliance data under the sub-item engineering category to which the bill of quantities data belongs, and calculate the Mahalanobis distance between the semantic fusion feature matrix and the process feature prototype vector; an audit module is used to determine that the current sub-item engineering of the project to be audited has an anomaly if the Mahalanobis distance is greater than a preset anomaly tolerance threshold; based on the location information of the empirical segment features, an audit result is generated that includes the cost item corresponding to the sub-item engineering with anomalies and the corresponding construction period lacking empirical information.
Citation Information
Patent Citations
Engineering cost investment intelligent control system and method fusing multi-source data
CN120894059A
Engineering cost anomaly identification method and system based on multi-modal feature fusion
CN121033544A