A construction engineering cost data management method, system and medium

By identifying semantic associations and implementing hierarchical encryption of unstructured documents in the construction project cost data management system, the problem of unstructured documents being easily used to infer structured data in existing technologies has been solved. This achieves collaborative encryption protection of structured and unstructured data, improving overall data security and resistance to inference.

CN121706119BActive Publication Date: 2026-05-26HEBEI ZHICHENG ENGINEERING PROJECT MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEBEI ZHICHENG ENGINEERING PROJECT MANAGEMENT CO LTD
Filing Date
2025-12-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies encrypt only the core structured cost data while ignoring unstructured documents that are semantically strongly related to it. This allows attackers to infer core cost information from plaintext peripheral data, rendering overall data security ineffective.

Method used

By identifying unstructured related fragments in unstructured documents that have semantic connections with structured cost core data, a first set of related fragments with a correlation greater than or equal to a preset security threshold and a second set of related fragments with a correlation less than the preset security threshold are generated. The structured cost core data and the first set of related fragments are encrypted based on a key mapping database while keeping the format and non-sensitive content unchanged. The encryption result is then encrypted a second time based on the cost correlation relationship.

Benefits of technology

By blocking the semantic association between peripheral data and core data, the overall cost data security is improved, ensuring data format and context consistency, and enhancing resistance to inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706119B_ABST
    Figure CN121706119B_ABST
Patent Text Reader

Abstract

The application discloses a kind of construction engineering cost data management method, system and medium, related to data management technical field, the method includes: receiving structured cost core data and identifying semantic association fragment in unstructured document;Analysis the cost correlation between them, generate high correlation first fragment set and low correlation second fragment set;Based on second fragment set retrieval historical data, difference comparison is executed in combination with first fragment set and core data, generate key mapping database;Accordingly, core data and first fragment set are encrypted, and secondary encryption is carried out based on correlation.The present application solves the technical problem that only structured cost core data is encrypted in the prior art, and the unstructured document with strong semantic association is ignored, leading to the attacker can reverse infer core cost information through plaintext peripheral data, and the overall data protection fails, so as to improve the technical effect of overall cost data security protection ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and in particular to a method, system and medium for managing construction project cost data. Background Technology

[0002] With the continuous improvement of building information technology (BIT) and digitalization, a large amount of data related to construction costs is generated at each stage of construction projects, including project initiation, design, bidding, construction, and settlement. This data includes both structured core cost data in the form of bills of quantities, unit price lists, and cost summary tables, and a large number of unstructured documents in the form of design drawings, contract documents, meeting minutes, and change orders. Existing construction cost data management technologies mainly focus on access control or encryption of cost values ​​stored in structured databases, while unstructured documents, which are highly related to them, are often only subject to simple access control or plaintext storage. Since unstructured documents often contain descriptions of unit price sources, basis for quantity calculations, technical specifications, and reasons for changes, they have significant semantic and business relationships with the structured core cost data. When attackers or unauthorized personnel obtain these unstructured documents, even if they cannot directly access the encrypted core cost data, they may still infer core cost information through semantic analysis and cross-comparison, thus causing the overall data security protection to fail.

[0003] Furthermore, existing encryption schemes mostly use a unified or static key to process cost data, lacking differentiation on the correlation strength between different cost data sets, and failing to fully utilize historical cost data for dynamic obfuscation during the encryption process. This approach not only makes it difficult to resist inference attacks based on correlation relationships, but may also disrupt the original data format and contextual consistency, affecting the normal use of cost data in statistical analysis, audit verification, and system compatibility. Summary of the Invention

[0004] This invention provides a method, system, and medium for managing construction project cost data. It addresses the technical problem in existing technologies where only structured core cost data is encrypted while unstructured documents with strong semantic associations are ignored. This allows attackers to infer core cost information from plaintext peripheral data, resulting in the failure of overall data protection. The invention achieves the technical effect of improving the overall cost data security by using collaborative encryption protection of structured and unstructured data to block the semantic association between peripheral and core data.

[0005] In a first aspect, the present invention provides a method for managing construction project cost data, wherein the method for managing construction project cost data includes:

[0006] The system receives structured cost core data of the target building project and automatically identifies unstructured related segments in unstructured documents that have semantic connections with the structured cost core data. It analyzes the cost correlation between the structured cost core data and the unstructured related segments, generating a first set of related segments with a correlation greater than or equal to a preset security threshold and a second set of related segments with a correlation less than the preset security threshold. Based on the second set of related segments, it retrieves a historical cost data set. In the historical cost data set, it performs a difference comparison between the structured cost core data and the first set of related segments, generating a key mapping database for the structured cost core data and the first set of related segments. Based on the key mapping database, it encrypts the structured cost core data and the first set of related segments while maintaining the original format and non-sensitive content, and then performs a second encryption on the encryption result based on the cost correlation.

[0007] Secondly, the present invention also provides a construction project cost data management system, wherein the construction project cost data management system comprises:

[0008] Semantic association recognition module: Receives structured cost core data of the target building project and automatically identifies unstructured related segments in unstructured documents that have semantic associations with the structured cost core data; Cost association analysis module: Analyzes the cost association relationship between the structured cost core data and the unstructured related segments, generating a first set of related segments with an association greater than or equal to a preset security threshold and a second set of related segments with an association less than the preset security threshold; Difference comparison module: Retrieves a historical cost data set based on the second set of related segments, performs a difference comparison on the historical cost data set in combination with the structured cost core data and the first set of related segments, and generates a key mapping database for the structured cost core data and the first set of related segments; Data encryption module: Encrypts the structured cost core data and the first set of related segments based on the key mapping database while keeping the format and non-sensitive content unchanged, and then performs a second encryption on the encryption result based on the cost association relationship.

[0009] Thirdly, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a construction project cost data management method provided by the present invention.

[0010] This invention discloses a method, system, and medium for managing construction project cost data, comprising: receiving structured core cost data of a target construction project and automatically identifying unstructured related segments in unstructured documents that have semantic association with the structured core cost data; analyzing the cost association relationship between the structured core cost data and the unstructured related segments, generating a first set of related segments with an association greater than or equal to a preset safety threshold and a second set of related segments with an association less than the preset safety threshold; retrieving a historical cost data set based on the second set of related segments, and performing a difference comparison in the historical cost data set, combining the structured core cost data and the first set of related segments, to generate information about the structured core cost data and the first set of related segments. The invention discloses a key mapping database; based on the key mapping database, the structured cost core data and the first associated fragment set are encrypted without changing the format and non-sensitive content; and then the encryption result is encrypted again based on the cost association relationship. The invention solves the technical problem in the prior art that only the structured cost core data is encrypted while ignoring the unstructured documents that are strongly related to it semantically, which allows attackers to infer the core cost information through plaintext peripheral data and the overall data protection fails. It achieves the technical effect of blocking the semantic association inference between peripheral data and core data through the collaborative encryption protection of structured and unstructured data, thereby improving the overall cost data security protection capability. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating a construction project cost data management method according to the present invention.

[0012] Figure 2 This is a schematic diagram of the structure of a construction project cost data management system according to the present invention.

[0013] Figure labeling: Semantic association identification module 11, cost association analysis module 12, difference comparison module 13, data encryption module 14. Detailed Implementation

[0014] The above technical solutions will now be described in detail with reference to the accompanying drawings and specific embodiments to provide a better understanding of them. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments used only to explain the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Furthermore, it should be noted that, for ease of description, only the parts related to the present invention are shown in the drawings, not all of them.

[0015] Example 1, as Figure 1 This is a flowchart illustrating a method for managing construction project cost data according to the present invention, wherein the method for managing construction project cost data includes:

[0016] Receive the structured cost core data of the target building project and automatically identify unstructured related segments in unstructured documents that have semantic association with the structured cost core data.

[0017] Specifically, the process begins by receiving structured cost core data corresponding to the target construction project. This structured cost core data consists of pre-generated structured documents carrying definite price data, including at least the project number, sub-item names, quantities, unit prices, total prices, and cost composition ratios, stored in the form of database tables or standardized data files. Next, unstructured documents related to the target construction project are acquired as analysis objects. These unstructured documents include at least CAD drawing files with unit price annotations, scanned copies of meeting minutes containing cost analysis content, and OCR recognition results of relevant contract texts. Optical character recognition and paragraph segmentation are performed on these unstructured documents to form a unified initial set of text fragments. Then, core entity keywords related to prices are extracted from the structured cost core data, including but not limited to project names, sub-item names, cost names, and semantic tags corresponding to prices. Corresponding core entity keywords are constructed, and the semantic similarity between each text fragment in the initial set of text fragments and the core entity keywords is calculated. This semantic similarity calculation can be based on word vector matching, contextual semantic association, or a pre-trained semantic model. When the semantic similarity between a text fragment and at least one core entity keyword is greater than or equal to a preset first threshold, it is determined that the text fragment has a semantic relationship with the structured cost core data. At this time, the text fragment will be marked as an unstructured related fragment, thereby realizing the automatic identification of content in unstructured documents that has a semantic relationship with the cost core data, for use in subsequent cost relationship analysis and encryption protection processes.

[0018] In some embodiments, structured cost core data of a target building project is received, and unstructured related segments in unstructured documents that have semantic association with the structured cost core data are automatically identified, including:

[0019] The system acquires drawings, contracts, and change orders related to the target construction project as unstructured document input sources. It then performs optical character recognition (OCR) and paragraph segmentation on the input unstructured documents to obtain an initial set of text fragments. The system extracts core entity keywords from the structured cost data and calculates the semantic similarity between each initial text fragment and the core entity keywords. Initial text fragments with a semantic similarity greater than or equal to a first threshold are marked as unstructured related fragments.

[0020] Specifically, the process begins by acquiring multi-source unstructured documents corresponding to the target construction project as input sources. These unstructured documents include at least engineering design drawings, engineering contract texts, and engineering change notices. Drawings include CAD format or its exported format files; contract texts include electronic documents or scanned copies; and engineering change notices include scanned files or image files. Content parsing and text processing are uniformly performed on unstructured documents from different sources and in different formats. Specifically, for CAD drawings, the format is parsed, and the layer structure and object attribute information of the drawings are read. Then, according to preset layer rules, text layers, annotation layers, and comment layers in the drawings are filtered. Text layers include text objects describing component names, material specifications, and cost descriptions; annotation layers include dimension lines, leader lines, and corresponding annotation text; and comment layers include unit price descriptions, price remarks, and labor or material cost annotations. Subsequently, the selected layers are traversed one by one to extract the text content, while simultaneously recording the spatial location information of the text objects in the drawing, the layer identifier, and the information of adjacent graphic objects. Then, based on the numerical characteristics of the text content, the characteristics of the pricing unit, and the spatial relationship with the dimension annotations or component graphics, the text content related to unit price, quantity, or cost description is identified, and the identified text content is output as candidate text information, thereby realizing the structured extraction of cost-related text annotations, dimension annotations, and unit price annotations in CAD drawings.

[0021] For scanned contract texts and engineering change notices, the input image files undergo image preprocessing. This preprocessing includes at least image grayscale conversion, noise removal, tilt correction, and contrast enhancement to improve the accuracy of subsequent character recognition. Then, based on layout analysis algorithms, the preprocessed image is segmented to identify text, table, and non-text regions, dividing the text regions into line-level or paragraph-level text blocks. Next, optical character recognition (OCR) is performed on the segmented text blocks to convert the characters in the image into corresponding editable text information, while preserving the character order, paragraph structure, and line position information. The recognized text information is further processed through text normalization, including removing noise characters, standardizing units of measurement, and standardizing number and symbol formats, thereby generating editable text results that are semantically consistent with the original scanned document content. This provides a reliable text foundation for subsequent paragraph segmentation and semantic similarity analysis. Next, the text-processed content is divided into paragraphs and sentences. Specifically, the text content undergoes layout structure analysis, identifying line breaks, blank lines, indentation marks, and page separators, and text positions that meet preset layout separation conditions are used as candidate paragraph boundaries. For text content originating from contracts and engineering change orders, consecutive line breaks, blank lines, or numbered clauses are used as the primary basis for paragraph separation. For content originating from meeting minutes or explanatory texts, paragraph boundaries are confirmed by combining title keywords, colon endings, and list symbols. For continuous text that cannot be directly separated by layout structure, semantic integrity-based segmentation is performed. This semantic integrity uses punctuation marks such as periods, exclamation marks, and question marks as segmentation references, and combines syntactic structure to determine whether they constitute independent semantic units. When the length of a text paragraph exceeds a preset maximum character threshold, the text paragraph is further split using punctuation marks representing integrity, dividing it into multiple semantically relatively complete sub-paragraphs. When the length of a text paragraph is less than a preset minimum character threshold, it is merged with adjacent paragraphs to avoid generating fragments with insufficient semantic information. After completing the paragraph boundary identification, segmentation, and merging processes described above, each identified paragraph or sub-paragraph is treated as an initial text fragment, and a unique fragment identifier is assigned to each initial text fragment. At the same time, the location information, source document type, and page number information of the initial text fragment in the original document are recorded, ultimately forming an initial text fragment set containing multiple initial text fragments, which are used for subsequent semantic similarity calculation and related fragment identification processing.

[0022] After obtaining the initial set of text fragments, core entity keywords are extracted from the structured cost data according to preset key names. These core entity keywords include at least the names of sub-items, cost items, pricing basis, materials, and semantic tags corresponding to unit price and total price. Then, based on a preset cost domain dictionary, these core entity keywords are part-of-speech tagging is performed to identify nouns, quantifiers, and proper nouns with cost semantics, filtering out function words and stop words that lack semantic distinction, retaining only valid terms that indicate the meaning of project cost. Subsequently, vector embedding processing is performed on the valid terms based on the cost domain corpus, mapping each valid term to a corresponding low-dimensional continuous vector space. For core entity keywords containing multiple valid terms, the vectors of each valid term are combined according to a preset weighting rule. This weighting rule is based on the frequency of term occurrence or the field weights in the structured cost data, thereby generating the initial semantic vector representation of the core entity keyword. Next, vector normalization is performed on these initial semantic vectors to give the semantic vectors of different core entity keywords a uniform scale range, thereby reducing the impact of noisy features on subsequent similarity calculations. After the above processing, the obtained normalized vectors are used as the semantic feature vectors corresponding to the core entity keywords and stored for use in subsequent semantic similarity calculations.

[0023] For each initial text segment in the initial text fragment set, text preprocessing is performed, including word segmentation, part-of-speech tagging, and irrelevant stop word filtering, to obtain a set of candidate keywords for that text segment. Then, the contextual semantic information of each candidate keyword is extracted from the initial text fragments, including the word combinations before and after the keyword, modification relationships, and syntactic structure information, to reflect the actual semantic meaning of the keyword in a specific context. Subsequently, based on a corpus in the cost estimation field, the extracted candidate keywords and their contextual semantics are mapped into word vector representations. For text segments containing multiple keywords, the word vectors corresponding to each keyword are weighted and summed according to keyword importance and position weight to generate a text segment semantic vector reflecting the overall semantics of the text segment. Subsequently, the semantic similarity between the semantic feature vector of the text fragment and the semantic feature vector of the core entity keyword is calculated based on the distance or similarity between vectors. For example, cosine similarity can be used as a semantic similarity metric. By calculating the dot product of the semantic feature vector of the text fragment and the semantic feature vector of the core entity keyword, and dividing each by the product of the magnitudes of the two vectors, a cosine value ranging from 0 to 1 is obtained to quantify their semantic closeness. When the semantic similarity between an initial text fragment and at least one core entity keyword is greater than or equal to a preset first threshold, it is determined that the initial text fragment has a direct or indirect semantic association with the structured cost core data. In this case, the initial text fragment is marked as a non-structured associated fragment. When the semantic similarity is less than the first threshold, no marking is performed. Through the above process, the automatic screening and identification of content related to the core data of structured cost in unstructured documents is realized. This solves the problem that the semantic relationship between unstructured engineering documents and structured cost data is difficult to identify accurately, relies on manual screening, or is easily overlooked. It improves the accuracy of identification of related information, the degree of automation, and the overall data security in the process of building engineering cost data management.

[0024] Analyze the cost correlation between the structured core cost data and the unstructured related segments to generate a first set of related segments with a correlation greater than or equal to a preset safety threshold and a second set of related segments with a correlation less than the preset safety threshold.

[0025] Specifically, after obtaining unstructured related fragments, various predefined types of associations reflecting cost estimation risks are activated, including price formation basis relationships, engineering quantity source relationships, and technical specification limitation relationships. The corresponding association analysis models are then trained based on historical engineering cost data. Subsequently, for each unstructured related fragment, the semantic vector of the initial text fragment it represents is concatenated with the semantic feature vector of the structured cost core data to construct a multi-dimensional feature vector characterizing the relationship between the two. This multi-dimensional feature vector is then input into the activated association analysis model for calculation, outputting the association type between the unstructured related fragment and the structured cost core data, along with its decryption association score. This decryption association score quantifies the degree to which the unstructured related fragment poses an inference risk to the structured cost core data in plaintext. Next, the cracking correlation score is compared with a preset security threshold. When the cracking correlation score is greater than or equal to the preset security threshold, the corresponding unstructured correlation segment is assigned to the first correlation segment set, and is determined to be a cost correlation segment with high correlation and high inference risk. When the cracking correlation score is less than the preset security threshold, the corresponding unstructured correlation segment is assigned to the second correlation segment set, and is determined to be a cost correlation segment with low correlation and low inference risk. Through the above process, the classification of unstructured correlation segments based on cost correlation strength is realized, providing a basis for subsequent difference comparison and layered encryption processing.

[0026] In some embodiments, analyzing the cost correlation between the structured core cost data and the unstructured related segments to generate a first set of related segments with a correlation greater than or equal to a preset safety threshold and a second set of related segments with a correlation less than the preset safety threshold includes:

[0027] Define the relationship type and train the relationship classification model; for each unstructured relationship segment, construct the feature vector between it and each data item in the structured cost core data, and input it into the relationship classification model to output the relationship type and cracking relationship score corresponding to each segment; use the cracking relationship score as the relationship quantification index and compare it with the preset security threshold to divide the first relationship segment set and the second relationship segment set.

[0028] Specifically, firstly, based on the characteristics of construction project cost estimation, several types of correlation relationships are predefined to characterize cost estimation risk. These correlation relationship types include at least price basis relationships, quantity source relationships, technical specification limitation relationships, and engineering change relationships. For each correlation relationship type, a corresponding correlation relationship classification model is trained based on structured cost data and corresponding unstructured document samples already labeled in historical projects. This correlation relationship classification model can be constructed using support vector machines, random forests, gradient boosting trees, CNNs, BERT, etc. Taking BERT as an example, during the model training phase, a training sample set is constructed based on the labeled data corresponding to the correlation relationship type. Each training sample consists of an unstructured text fragment, its corresponding structured core cost data, and annotation information. The annotation information includes the cost correlation relationship type pre-labeled manually or semi-automatically, and the correlation strength label of the text fragment's inference risk to the core cost data.

[0029] Subsequently, the unstructured text fragments undergo input preprocessing. The system uses a word segmenter matched with the BERT model to segment the text, adding a classification marker [CLS] at the beginning and a separator marker [SEP] at the end. Simultaneously, the text sequence is uniformly padded or truncated to a preset maximum length to form a fixed-length word sequence. Word embedding vectors, position embedding vectors, and sentence embedding vectors are generated from the processed word sequence, and these three are summed as the input embedding representation for the BERT model. Structurally, the association classification model includes a text semantic encoding layer, a cost feature fusion layer, and a classification and scoring output layer. The text semantic encoding layer employs a multi-layer Transformer Encoder structure BERT model. This BERT model includes multi-layer self-attention mechanisms and a feedforward neural network. Each Transformer Encoder layer models the contextual dependencies between different words throughout the entire text sequence through a multi-head self-attention mechanism, thereby learning the implicit semantic association features in the text. During forward propagation, the input lexical sequence passes through each Transformer Encoder layer sequentially. Each layer calculates attention weights based on the query vector, key vector, and value vector, enabling the model to automatically focus on keywords, reminder descriptions, or numerical expressions related to cost semantics. After processing by all Transformer layers, the model outputs the contextual semantic representation corresponding to each lexical unit. The output vector corresponding to the [CLS] tag serves as the global semantic representation vector for the entire unstructured associated segment. After obtaining the semantic representation vector, a cost feature vector corresponding to the unstructured associated segment is constructed. This cost feature vector includes at least the numerical features of the engineering quantity, the numerical features of the unit price, the cost type encoding features, and the matching degree features between the text segment and the engineering item names from the structured cost core data.

[0030] Next, the text semantic representation vector and the cost feature vector are concatenated or weighted and fused along the feature dimension to form a joint feature representation. This joint feature representation is then input into the subsequent classification and scoring output layer. This classification and scoring output layer includes at least one fully connected layer. One output branch is used to classify the cost correlation types, using a Softmax activation function to output the probability distribution of each correlation type. The other output branch is used to generate the cracking correlation score, using a Sigmoid or linear activation function to output a continuous numerical value representing the inferred risk intensity. After the model's forward propagation is complete, the predicted correlation types output by the model are compared with the true correlation types labeled in the training samples to calculate the classification cross-entropy loss. Simultaneously, the cracking correlation score output by the model is compared with the true correlation strength label to calculate the regression loss or mean squared error loss. These two losses are then weighted and summed according to preset weights to obtain the model's comprehensive loss value on the current training samples.

[0031] Then, based on the backpropagation algorithm, starting from the comprehensive loss function, the gradient information of the parameters of each layer is calculated layer by layer, including the attention weight parameters of each Transformer Encoder layer, the feedforward network parameters, and the weight parameters of the fully connected output layer in the BERT model. The optimizer preferably uses the Adam optimization algorithm, combined with a preset learning rate, to adaptively update the model parameters, enabling the model to reduce the prediction error of the cost correlation and cracking correlation scores in the next training round. The above training process is executed repeatedly in batches until the preset number of training rounds is reached or the model loss value converges.

[0032] After model training is complete, the trained model is validated using a validation dataset to evaluate whether it meets the preset performance requirements in terms of cost correlation classification accuracy and correlation scoring error. When the model performance meets the requirements, it is stored as the final correlation classification model and used for subsequent analysis of cost correlations between non-structural correlation segments and structured core cost data in the target building project. When the model performance does not meet the requirements, it is retrained by adjusting hyperparameters such as learning rate, batch size, or training epochs.

[0033] In practical applications, for each unstructured related segment in the current target construction project, a multi-dimensional feature vector is constructed based on the same splicing process described above, relating it to various data points in the core structured cost data. This constructed multi-dimensional feature vector is then input into a correlation classification model for inference operations, outputting the correlation type corresponding to each unstructured related segment and its corresponding cracking correlation score. Finally, the cracking correlation score is used as a correlation quantification index and compared with a pre-set security threshold. When the cracking correlation score for a certain unstructured related segment is greater than or equal to the preset security threshold, the unstructured related segment is assigned to the first correlation segment set and determined to be a high-correlation, high-risk segment. When the cracking correlation score is less than the preset security threshold, the unstructured related segment is assigned to the second correlation segment set and determined to be a low-correlation, low-risk segment. Through this process, the correlation quantification and classification of unstructured related segments are achieved, providing a basis for subsequent difference comparison and multi-layer encryption strategies.

[0034] In some embodiments, the types of associations include price basis, source of work volume, technical specification limitation, and change association.

[0035] Specifically, to accurately characterize the inference capability of unstructured related segments on structured cost core data, cost relationships are divided into several specific types. Different types reflect different business dependency paths between unstructured information and cost data. Among them, the price basis relationship describes whether the unstructured related segment contains content related to the source of unit price formation or pricing basis, such as quota standards, market price references, material quotation instructions, or cost calculation rules. When this type of information corresponds to the unit price field in the structured cost core data, it is easy to directly infer the price value. The quantity source relationship describes whether the unstructured related segment involves quantity calculation methods, measurement basis, or... Drawing dimension descriptions, such as dimensions marked on drawings, quantity calculation formulas, or bill of quantities preparation instructions, can provide reverse verification or extrapolation paths for the quantity field in the core structured cost data. Technical specification constraints describe whether non-structural related segments specify limitations on material specifications, construction techniques, technical parameters, or quality grades. These constraints are often implicitly related to unit price levels and cost composition, narrowing the range of cost estimates. Change-related relationships describe whether non-structural related segments include reasons for engineering changes, scope of changes, increases or decreases in quantities, or cost adjustments. This information directly affects cost adjustment items in the core structured cost data. These relationship types provide a clear business semantic basis for subsequent correlation scoring calculations and risk classification.

[0036] Based on the second set of associated fragments, a historical cost data set is retrieved. In the historical cost data set, a difference comparison is performed between the structured cost core data and the first set of associated fragments to generate a key mapping database for the structured cost core data and the first set of associated fragments.

[0037] Specifically, after classifying the unstructured related fragments, a second set of related fragments with a correlation score below a preset security threshold is selected as the basis for historical cost data retrieval. Since the second set of related fragments has a low decryption correlation score, its content in plaintext makes it difficult to directly infer the core structured cost data. Therefore, the system does not encrypt the second set of related fragments itself. Instead, it uses the second set of related fragments as the search condition to retrieve historical project data similar to the current project from historical cost data resources, forming a historical cost data set. Subsequently, in the historical cost data set, each piece of historical cost data is aligned with the core structured cost data of the current target project and the first set of related fragments. Differences are compared across multiple dimensions, including quantity, unit price, and cost composition, to calculate the degree of difference at both the numerical and semantic levels. Subsequently, based on the obtained difference comparison results, historical data content with the greatest comprehensive difference from the current structured cost core data and the first associated fragment set is selected from the historical cost data set. This data is then used as the source of the obfuscation key to construct a key mapping relationship between the structured cost core data and the first associated fragment set, thereby generating a key mapping database. Through this method, even after encrypting the structured cost core data and the first associated fragment set, the overall data format, field structure, and contextual relationships remain consistent with normal cost data. At the same time, reasonable business relationships still exist between different data points, making it difficult to directly identify as abnormal data. Furthermore, after subsequent secondary encryption based on cost relationships, even if the secondary encryption is partially cracked, the exposed data content will approximate a reasonable distribution of historical cost data, further reducing the risk of continuous inference or further cracking, thus achieving high security and concealment protection for the overall cost data.

[0038] In some embodiments, retrieving a historical cost data set based on the second associated fragment set includes:

[0039] Keyword extraction and feature encoding are performed on the second set of associated fragments to generate a retrieval query vector; in the historical cost data set, historical engineering project retrieval based on project feature similarity and historical engineering project retrieval based on text content similarity are performed in parallel, and the results returned by the two retrievals are merged and deduplicated to generate the historical cost data set.

[0040] Specifically, for the second set of related fragments with a relevance lower than a preset safety threshold, the same keyword extraction process described above is first performed on each unstructured related fragment in the second set of related fragments. The extracted keywords are then standardized and vectorized to generate a unified retrieval query vector. This retrieval query vector is used to characterize the overall features of the current project at the business attribute and textual semantic levels. Subsequently, using the retrieval query vector as input, two different dimensions of historical project retrieval processing are performed in parallel within the historical cost data resources. One retrieval method is historical project retrieval based on project feature similarity. That is, the retrieval query vector is compared with the structured project feature information stored in historical projects. This project feature information includes at least the project category, construction scale, structural form, regional attributes, and pricing model. Through the aforementioned similarity calculation method, historical projects with similar features to the target project can be selected. The other retrieval method is historical project retrieval based on text content similarity. That is, the retrieval query vector is compared with the corresponding cost description text, project description text, and change description text in historical projects to select historical projects that are similar to the second set of related fragments at the textual semantic level. After obtaining the historical project results returned by the two retrieval methods, the results are fused. Specifically, the two retrieval results are merged, duplicate historical projects are removed, and the results are sorted according to similarity to generate a historical cost data set for subsequent difference comparison and key mapping. This process ensures that the generated historical cost data set has reasonable similarity to the target project in terms of business attributes and semantic content, providing a valid data foundation for subsequent encryption and obfuscation.

[0041] In some embodiments, in the historical cost data set, a difference comparison is performed between the structured cost core data and the first associated fragment set to generate a key mapping database for the structured cost core data and the first associated fragment set, including:

[0042] Each historical cost data point in the historical cost data set is aligned with the structured cost core data and the first associated fragment set, and a comparison matrix is ​​constructed according to data dimensions. Based on the comparison matrix, the comprehensive numerical difference rate and text description difference degree of each historical cost data point with the structured cost core data and the first associated fragment set in multiple dimensions are calculated. According to the numerical difference rate and text description difference degree, the obfuscation key with the largest comprehensive difference is selected from the historical cost data set for the structured cost core data and the first associated fragment set, forming the key mapping database.

[0043] Specifically, after obtaining the historical cost data set, the first step is to perform data alignment processing on each historical cost data entry and its corresponding structured cost core data and first associated fragment set for the current target construction project. This involves mapping and matching data from different sources according to a unified data dimension, which includes at least the project name, sub-items, quantity, unit price, total price, cost composition ratio, and corresponding textual descriptions. Through alignment processing, a one-to-one correspondence is established between each field in the historical cost data and its corresponding fields or semantic content in the structured cost core data and first associated fragment set. Based on this, a comparison matrix for difference analysis is constructed, where each row of the comparison matrix corresponds to one historical cost data entry, and each column corresponds to a data dimension. After constructing the comparison matrix, the degree of difference between each historical cost data entry and the structured cost core data and first associated fragment set across multiple dimensions is calculated based on the comparison matrix. At the numerical level, the numerical difference rate is calculated for fields such as quantity, unit price, total price, and cost ratio. This numerical difference rate is quantified by normalized differences or relative deviations. At the textual level, the textual description difference degree is calculated for the explanatory text in the historical cost data and the text content in the first associated fragment set. This textual description difference degree is evaluated based on the difference between 1 and the similarity score. Then, the numerical difference rate and textual description difference degree of each data dimension are weighted and summarized to obtain a comprehensive difference index for each historical cost data point relative to the structured cost core data and the first associated fragment set. Finally, based on these comprehensive difference indices, the historical cost data content with the largest comprehensive difference is selected from the historical cost data set. The numerical fields and text fragments corresponding to this historical cost data are used as the source of the obfuscation key, establishing a key mapping database between it and the structured cost core data and the first associated fragment set. Through this process, the selected obfuscation key remains within a reasonable range of cost data in terms of business semantics, but differs significantly from the real data in terms of numerical and descriptive aspects, effectively improving the obfuscation and anti-inference capabilities of subsequent encryption results.

[0044] The structured cost core data and the first associated fragment set are encrypted based on the key mapping database while keeping the format and non-sensitive content unchanged. The encryption result is then encrypted a second time based on the cost association relationship.

[0045] Specifically, after generating the key mapping database, the structured cost core data and the first associated fragment set are first encrypted using the key mapping database, preserving the format and non-sensitive content. During this process, field-level parsing is performed on the data content of the structured cost core data and the first associated fragment set to identify sensitive information fields, such as unit price, total price, and the monetary value corresponding to the quantity of work. Subsequently, based on the key mapping database, the actual values ​​or key text in the sensitive fields are replaced with the corresponding obfuscated key content. This generates the first layer of encryption without changing the original data format, field length, or contextual semantic structure, ensuring that the encrypted data is consistent with normal cost data in appearance and business logic. After completing the first layer of encryption, a second layer of encryption is performed on the first layer of encryption based on the cost correlation relationship analyzed between the structured cost core data and the first associated fragment set. This ensures that the correlation constraints must be satisfied simultaneously when decrypting or tampering with the encrypted data. This secondary encryption method makes it difficult for attackers to recover complete and accurate cost information even if they crack a single field or fragment. Furthermore, in extreme cases, even if the secondary encryption is partially removed, the exposed data still appears as reasonable but obfuscated cost data, making it difficult to further deduce the true structured cost core data. This significantly improves the security and anti-deduction capabilities of the overall data management solution.

[0046] In some embodiments, data encryption is performed on the structured cost core data and the first associated fragment set based on the key mapping database while maintaining the original format and non-sensitive content, including:

[0047] Identify the sensitive fields in the structured cost core data and the first associated fragment set. The sensitive fields include at least unit price, total price, and cost composition ratio. Read the key mapping database and use the corresponding obfuscation key to perform reversible replacement encryption on the content of the sensitive fields, while preserving the unit, format, and contextual description of the data, and generate an encryption result.

[0048] Specifically, the process begins by identifying sensitive fields in the core structured cost data and the first associated fragment set based on preset key names. These sensitive fields include at least unit price, total price, and cost composition ratio, directly related to the financial data of the construction project and thus considered highly sensitive information. Next, the obfuscation key corresponding to each sensitive field is read from the key mapping database, and these obfuscation keys are used to reversibly replace and encrypt the content of the sensitive fields, ensuring that the real data of the sensitive fields is replaced with the obfuscated data, thereby hiding the real data content. During the encryption process, the units, format, and contextual descriptions of the data remain unchanged. For example, "100 yuan / square meter" in the unit price field will be encrypted as "102 yuan / square meter," but the unit "yuan / square meter" and the data format (such as two decimal places) remain unchanged. Furthermore, the contextual descriptions related to the sensitive fields are retained to ensure that the encrypted data is structurally consistent with the original data. Through these steps, the generated encryption result not only effectively hides sensitive data but also maintains the data format, units, and descriptive text, ensuring that the encrypted data can still be used in business processing without leaking sensitive information.

[0049] In some embodiments, the encryption result is further encrypted based on the cost correlation, including:

[0050] Based on the cost correlation, a combination relationship key is generated for the fragment combination corresponding to the encryption result; the combination relationship key is used to perform block binding encryption on the corresponding data in the encryption result.

[0051] Specifically, based on the identified cost-related relationships, the types of relationships existing in the encrypted results are determined, and fragments corresponding to each relationship type are extracted from the encrypted results. Then, based on the cracking correlation score in the cost-related relationships, the required key strength and corresponding encryption strategy are matched. Next, according to the selected encryption strategy and the key strength requirements, the fragment combinations corresponding to each relationship type are encrypted. For example, a hash-based dynamic key algorithm can be used to encrypt the fragment combinations, generating a feature hash value. This feature hash value is then concatenated with the relationship type code to form a combined relationship key. Afterwards, this combined relationship key is used to perform block-binding encryption on the relevant data in the encrypted results. That is, the encrypted data is grouped according to a preset data structure to ensure that the data fields or text fragments within each group have close business relationships. For example, data related to price is divided into price blocks, and data related to the source of work volume is divided into work volume blocks. This combined relationship key is then used to perform secondary encryption on the data in the corresponding blocks, ensuring that the data within the same group maintains close relationships even after encryption, while data from different groups are encrypted independently using different keys. This block-binding encryption method not only ensures data security, but also effectively prevents data from being cracked at the level of a single field or fragment. Even if some data is decrypted by an attacker, other data that lacks correlation cannot be inferred, thus greatly enhancing the overall data's resistance to inference.

[0052] In summary, the construction project cost data management method provided by this invention has the following technical effects:

[0053] The system receives structured core cost data of a target building project and automatically identifies unstructured related segments in unstructured documents that have semantic connections with the structured core cost data. It analyzes the cost correlation between the structured core cost data and the unstructured related segments, generating a first set of related segments with a correlation greater than or equal to a preset security threshold and a second set of related segments with a correlation less than the preset security threshold. Based on the second set of related segments, it retrieves a historical cost data set. In this historical cost data set, it performs a difference comparison between the structured core cost data and the first set of related segments, generating a key mapping database for the structured core cost data and the first set of related segments. Based on the key mapping database, it encrypts the structured core cost data and the first set of related segments while maintaining the original format and non-sensitive content. Then, it performs a second encryption based on the cost correlation, thereby achieving the technical effect of blocking the semantic correlation derivation between peripheral data and core data through collaborative encryption protection of structured and unstructured data, and improving the overall cost data security protection capability.

[0054] Example 2, as Figure 2 This is a schematic diagram of the structure of a construction project cost data management system according to the present invention. For example, Figure 1 The flowchart of a construction project cost data management method of the present invention can be illustrated as follows: Figure 2 The structure shown is implemented.

[0055] Based on the same concept as the construction project cost data management method in the above embodiment, the present invention also provides a construction project cost data management system comprising:

[0056] Semantic association recognition module 11: Receives structured cost core data of the target building project and automatically identifies unstructured related segments in unstructured documents that have semantic association with the structured cost core data; Cost association analysis module 12: Analyzes the cost association relationship between the structured cost core data and the unstructured related segments, generating a first set of related segments with an association greater than or equal to a preset security threshold and a second set of related segments with an association less than the preset security threshold; Difference comparison module 13: Retrieves a historical cost data set based on the second set of related segments, performs a difference comparison on the historical cost data set in combination with the structured cost core data and the first set of related segments, and generates a key mapping database for the structured cost core data and the first set of related segments; Data encryption module 14: Encrypts the structured cost core data and the first set of related segments based on the key mapping database while keeping the format and non-sensitive content unchanged, and then performs a second encryption on the encryption result based on the cost association relationship.

[0057] In some embodiments, the semantic association recognition module 11 includes:

[0058] The system acquires drawings, contracts, and change orders related to the target construction project as unstructured document input sources. It then performs optical character recognition (OCR) and paragraph segmentation on the input unstructured documents to obtain an initial set of text fragments. The system extracts core entity keywords from the structured cost data and calculates the semantic similarity between each initial text fragment and the core entity keywords. Initial text fragments with a semantic similarity greater than or equal to a first threshold are marked as unstructured related fragments.

[0059] In some embodiments, the cost correlation analysis module 12 includes:

[0060] Define the relationship type and train the relationship classification model; for each unstructured relationship segment, construct the feature vector between it and each data item in the structured cost core data, and input it into the relationship classification model to output the relationship type and cracking relationship score corresponding to each segment; use the cracking relationship score as the relationship quantification index and compare it with the preset security threshold to divide the first relationship segment set and the second relationship segment set.

[0061] In some embodiments, the cost correlation analysis module 12 includes:

[0062] The types of relationships include price basis, source of project quantity, technical specification limitation, and change relationship.

[0063] In some embodiments, the difference comparison module 13 includes:

[0064] Keyword extraction and feature encoding are performed on the second set of associated fragments to generate a retrieval query vector; in the historical cost data set, historical engineering project retrieval based on project feature similarity and historical engineering project retrieval based on text content similarity are performed in parallel, and the results returned by the two retrievals are merged and deduplicated to generate the historical cost data set.

[0065] In some embodiments, the difference comparison module 13 includes:

[0066] Each historical cost data point in the historical cost data set is aligned with the structured cost core data and the first associated fragment set, and a comparison matrix is ​​constructed according to data dimensions. Based on the comparison matrix, the comprehensive numerical difference rate and text description difference degree of each historical cost data point with the structured cost core data and the first associated fragment set in multiple dimensions are calculated. According to the numerical difference rate and text description difference degree, the obfuscation key with the largest comprehensive difference is selected from the historical cost data set for the structured cost core data and the first associated fragment set, forming the key mapping database.

[0067] In some embodiments, the data encryption module 14 includes:

[0068] Identify the sensitive fields in the structured cost core data and the first associated fragment set. The sensitive fields include at least unit price, total price, and cost composition ratio. Read the key mapping database and use the corresponding obfuscation key to perform reversible replacement encryption on the content of the sensitive fields, while preserving the unit, format, and contextual description of the data, and generate an encryption result.

[0069] In some embodiments, the data encryption module 14 includes:

[0070] Based on the cost correlation, a combination relationship key is generated for the fragment combination corresponding to the encryption result; the combination relationship key is used to perform block binding encryption on the corresponding data in the encryption result.

[0071] In embodiment three, the present invention also provides a computer-readable storage medium that can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to a construction project cost data management method in an embodiment of the present invention, thereby realizing the aforementioned construction project cost data management method.

[0072] It should be understood that the embodiments disclosed in this invention and the above description enable those skilled in the art to implement this invention. However, this invention is not limited to the embodiments mentioned above. It should be understood that those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this invention, and should all be included within the protection scope of this invention.

Claims

1. A method for managing construction project cost data, characterized in that, include: Receive the structured cost core data of the target building project and automatically identify unstructured related segments in unstructured documents that have semantic association with the structured cost core data; Analyze the cost correlation between the structured core cost data and the unstructured related segments to generate a first set of related segments with a correlation greater than or equal to a preset safety threshold and a second set of related segments with a correlation less than the preset safety threshold; Based on the second associated fragment set, a historical cost data set is retrieved. Within this historical cost data set, a difference comparison is performed between the structured cost core data and the first associated fragment set to generate a key mapping database for the structured cost core data and the first associated fragment set. Specifically, this includes: Align each historical cost data point in the historical cost data set with the structured cost core data and the first associated fragment set, and construct a comparison matrix according to the data dimensions; Based on the comparison matrix, calculate the numerical difference rate and text description difference degree of each historical cost data with the structured cost core data and the first associated fragment set across multiple dimensions. Based on the numerical difference rate and the text description difference degree, the key mapping database is constructed by selecting the obfuscated key with the largest comprehensive difference from the structured cost core data and the first associated fragment set in the historical cost data set. The structured cost core data and the first associated fragment set are encrypted based on the key mapping database while keeping the format and non-sensitive content unchanged. The encryption result is then encrypted a second time based on the cost association relationship.

2. The method for managing construction project cost data as described in claim 1, characterized in that, Receives structured cost core data for the target building project and automatically identifies unstructured related segments in unstructured documents that have semantic connections with the structured cost core data, including: The system obtains drawings, contracts, and change orders related to the target construction project as unstructured document input sources. It then performs optical character recognition and paragraph segmentation on the input unstructured documents to obtain an initial set of text fragments. Extract the core entity keywords from the structured cost core data, and calculate the semantic similarity between each initial text fragment in the initial text fragment set and the core entity keywords; Initial text segments with semantic similarity greater than or equal to the first threshold are marked as unstructured related segments.

3. The method for managing construction project cost data as described in claim 1, characterized in that, Analyze the cost correlation between the structured core cost data and the unstructured related segments to generate a first set of related segments with a correlation greater than or equal to a preset safety threshold and a second set of related segments with a correlation less than the preset safety threshold, including: Define the types of relationships and train a relationship classification model; For each unstructured related segment, a feature vector is constructed between it and each data item in the structured cost core data, and then input into the relationship classification model to output the relationship type and the relationship score of each segment. The cracking correlation score is used as a correlation quantification indicator and compared with the preset security threshold to divide the first correlation fragment set and the second correlation fragment set.

4. The method for managing construction project cost data as described in claim 3, characterized in that, The types of relationships include price basis, source of project quantity, technical specification limitation, and change relationship.

5. The method for managing construction project cost data as described in claim 1, characterized in that, Based on the second set of associated fragments, a historical cost data set is retrieved, including: Keyword extraction and feature encoding are performed on the second set of related fragments to generate a retrieval query vector; In the historical cost data set, historical engineering project retrieval based on project feature similarity and historical engineering project retrieval based on text content similarity are performed in parallel. The results returned by the two retrievals are merged and deduplicated to generate the historical cost data set.

6. The method for managing construction project cost data as described in claim 1, characterized in that, Based on the key mapping database, data encryption is performed on the structured cost core data and the first associated fragment set while maintaining the format and non-sensitive content, including: Identify the sensitive fields in the structured cost core data and the first associated fragment set, wherein the sensitive fields include at least unit price, total price, and cost composition ratio; The key mapping database is read, and the content of sensitive fields is reversibly replaced and encrypted using the corresponding obfuscation key, while preserving the data's unit, format, and contextual description text, thus generating an encrypted result.

7. The method for managing construction project cost data as described in claim 1, characterized in that, The encryption result is then further encrypted based on the aforementioned cost correlation, including: Based on the cost correlation, generate a combination relationship key for the fragment combination corresponding to the encryption result; The combined relation key is used to perform block binding encryption on the corresponding data in the encryption result.

8. A construction project cost data management system, characterized in that, A method for managing construction project cost data as described in any one of claims 1-7 includes: Semantic association recognition module: Receives the structured cost core data of the target building project and automatically identifies unstructured related segments in unstructured documents that have semantic association with the structured cost core data; Cost correlation analysis module: Analyzes the cost correlation between the structured core cost data and the unstructured correlation segments, and generates a first set of correlation segments with a correlation greater than or equal to a preset safety threshold and a second set of correlation segments with a correlation less than the preset safety threshold; The difference comparison module retrieves a set of historical cost data based on the second associated fragment set. In the set of historical cost data, it performs a difference comparison by combining the structured cost core data and the first associated fragment set to generate a key mapping database about the structured cost core data and the first associated fragment set. Data encryption module: Based on the key mapping database, the structured cost core data and the first associated fragment set are encrypted without changing the format and non-sensitive content, and then the encryption result is encrypted a second time based on the cost association relationship.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a construction project cost data management method as described in any one of claims 1 to 7.