Construction project document classified storage and retrieval method and system based on BIM (Building Information Modeling)

By adopting BIM-based knowledge graph and natural language processing technology in the construction project document management system, automating document classification and providing diversified search methods, the problem of low document management and search efficiency in existing systems is solved, and more efficient, safe and traceable document management is achieved.

CN120144585APending Publication Date: 2025-06-13陈琛
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216907.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing construction project document management system is difficult to manage and retrieve documents in a unified manner due to different projects or enterprises using different document classification standards, and manual classification has subjectivity and limitations, which affects the search efficiency.

Method used

Using BIM-based construction project document classification storage and search methods, we automate document classification by building knowledge graphs, natural language processing and machine learning technology, and provide document version control and traceability functions in combination with blockchain to establish the relationship between documents and BIM models, and provide keyword search, semantic search and classification search methods.

Benefits of technology

It improves the accuracy and efficiency of document classification, reduces manual intervention, provides a diverse search method, facilitates users to find required documents from different angles, and quickly searches related documents through BIM model components, enhancing the security and traceability of document management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144585A_ABST
    Figure CN120144585A_ABST
Patent Text Reader

Abstract

The invention discloses a BIM-based construction project document classified storage and retrieval method and system, and relates to the technical field of construction industry information.The BIM-based construction project document classified storage and retrieval method comprises the following steps that 1, a knowledge graph containing BIM model information, entity nodes, relation edges, attribute information, project terms and professional knowledge is constructed; 2, analyzing document contents into semantic units by utilizing natural language processing, associating the semantic units with entities in the knowledge graph, and establishing semantic association between the documents; according to the method, the automatic classification result is optimized and adjusted through the predefined classification rule, the classification accuracy is ensured, manual intervention is reduced, and the working efficiency is improved. And various modes such as keyword retrieval, semantic retrieval and classification retrieval are provided, so that the user can conveniently search the required document from different angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology in the construction industry, and specifically to a method and system for classifying, storing, and retrieving construction project documents based on BIM. Background Art

[0002] Construction project documents refer to the general term for various document materials generated in each stage of construction project planning, design, construction, supervision, and acceptance. These documents play an important role in proving and tracing aspects such as the quality, progress, and investment control of construction projects, and are also an important part of project management.

[0003] The classified storage and retrieval of construction project documents is a very important link in project management, which helps to ensure the orderliness, efficiency, and security of documents.

[0004] However, different existing projects or enterprises often adopt different document classification standards, resulting in difficulties in unified management and retrieval of documents. Moreover, due to the subjectivity and limitations of manual classification, the classification results are biased, affecting the retrieval efficiency. At the same time, although many systems provide keyword retrieval, they cannot meet the diverse retrieval needs of users. Therefore, we propose a method and system for classifying, storing, and retrieving construction project documents based on BIM to solve the problems mentioned above. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for classifying, storing, and retrieving construction project documents based on BIM to solve one of the problems such as low document retrieval efficiency and insufficient document management security existing in the current market as mentioned in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A method for classifying, storing, and retrieving construction project documents based on BIM includes the following steps:

[0008] Step 1: Construct a knowledge graph containing BIM model information, entity nodes, relationship edges, attribute information, engineering terms, and professional knowledge;

[0009] Step 2: Use natural language processing to parse the document content into semantic units, associate them with the entities in the knowledge graph, and establish semantic associations between documents;

[0010] Step 3: Use natural language processing and machine learning to extract keywords, themes, and entity information from the documents, train a document classification model, automatically classify according to document features, and optimize and adjust in combination with predefined classification rules;

[0011] Step 4: Record the version information, modification records, and permission change records of the document on the blockchain, and provide document version control and traceability functions through the blockchain;

[0012] Step 5: Combine the document content with the BIM model to establish an association relationship between the document and the BIM model components;

[0013] Step 6: Store the document on the cloud platform, and establish an association relationship between the document and the BIM model information, providing keyword retrieval, semantic retrieval, BIM model retrieval, and classification retrieval methods.

[0014] As a further optimization solution of the present invention, in Step 2, the specific steps of parsing the document content into semantic units are as follows:

[0015] For text preprocessing, first perform word segmentation to cut the document text into a sequence of words, perform part-of-speech tagging to identify the part of speech of each word, filter out stop words, restore the words to their basic forms, decompose the document text into sentences and words, analyze the sentence structure, identify the semantic roles of each component in the sentence, link the named entities in the document with the entities in the knowledge graph, extract the relationships between entities from the document text, and perform semantic reasoning using the knowledge in the knowledge graph;

[0016] Finally, identify the new entities appearing in the document, add them to the knowledge graph, update the relationships between entities in the knowledge graph, and update the attribute information of entities in the knowledge graph.

[0017] As a further optimization solution of the present invention, in Step 3, use the bag-of-words model to represent the document as a word frequency vector, calculate the TF-IDF value of the words, map the words to a low-dimensional space, identify the named entities in the document through named entity recognition, analyze the syntactic structure of the document, extract sentence components and dependency relationships, and use the Latent Dirichlet Allocation model to represent the document as a topic distribution to extract the topic information of the document.

[0018] As a further optimization solution of the present invention, calculate the occurrence probability of each category in the training set and the occurrence probability of each word in each category, smooth the conditional probability, and for each document, calculate its posterior probability in each category, and select the category with the largest posterior probability as the classification result of the document.

[0019] As a further optimization solution of the present invention, in the bag-of-words model, the word frequency vector is represented as:

[0020] V(d)=(tf i,d ,tf 2,d ,...,tf n,d )

[0021] Among them, tf i,d is the number of occurrences of word i in document d;

[0022] TF-IDF calculation formula:

[0023] tfidf(t, d, D) = tf(t, d) × idf(t, D)

[0024] Among them, tf(t, d) is the term frequency of word t in document d, and idf(t, D) is the inverse document frequency of word t in document set D.

[0025] A BIM-based construction project document classification storage and retrieval system, including a knowledge graph engine, an artificial intelligence engine, a blockchain engine, a virtual reality engine, and a document management platform;

[0026] Among them, the knowledge graph engine is used to construct and manage the knowledge graph and provide semantic understanding functions; the artificial intelligence engine is used for automatic document classification and semantic analysis; the blockchain engine is used for document version control and traceability; the virtual reality engine is used for three-dimensional document display; the document management platform is used for document storage, retrieval, and management.

[0027] Compared with the prior art, the beneficial effects of the present invention are:

[0028] The present invention optimizes and adjusts the automatic classification results through predefined classification rules, ensures classification accuracy, reduces manual intervention, and improves work efficiency. And it provides keyword retrieval, semantic retrieval, and classification retrieval methods, facilitating users to find the required documents from different perspectives. Moreover, the system also establishes an association relationship between the documents and BIM model components, and users can quickly find relevant documents through the BIM model components.

[0029] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the above-described illustrative aspects, embodiments, and features, further aspects, embodiments, and features of the present invention will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flowchart of the BIM-based construction project document classification storage and retrieval method of the present invention;

[0031] Figure 2 It is a structural block diagram of the BIM-based construction project document classification storage and retrieval system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] Embodiment 1

[0034] As shown in the Figure 1 accompanying drawings, a method for classified storage and retrieval of construction project documents based on BIM includes the following steps:

[0035] Construct a knowledge graph including BIM model information, entity nodes, relationship edges, attribute information, engineering terms, and professional knowledge;

[0036] Among them, the knowledge graph includes entity nodes, relationship edges, attribute information, engineering terms, and professional knowledge;

[0037] The entity nodes include project entities, component entities, attribute entities, personnel entities, organizational entities, equipment entities, and material entities;

[0038] The relationship edges include association relationships, attribute relationships, and hierarchical relationships;

[0039] The association relationships include project-component, component-attribute, project-personnel, project-organization, component-equipment, and component-material; the attribute relationships include material type-material specification, equipment model-equipment parameters; the hierarchical relationships include project-phase, component-subcomponent, and organization-department;

[0040] The attribute information includes entity attributes and relationship attributes;

[0041] The document semantic understanding includes syntactic analysis, part-of-speech tagging, named entity recognition, semantic role labeling, semantic role labeling, and semantic association:

[0042] Syntactic analysis is used to decompose the document text into sentences and words; part-of-speech tagging is used to identify the part of speech of each word; named entity recognition is used to identify named entities in the document, such as person names, place names, and organization names; semantic role labeling is used to identify the semantic roles of each component in the sentence; semantic association is used to associate the document content with the entities in the knowledge graph.

[0043] The formula used by the conditional random field model is:

[0044] f k (y i-1 ,y i ,x,i)

[0045] Denote the given current and previous label y i-1 , y i and the specific feature k of the input sequence x at position i.

[0046] The formula used for the conditional probability is:

[0047]

[0048] That is, the conditional probability of the label sequence y given the input sequence x.

[0049] The formula used for the maximum likelihood estimation is:

[0050]

[0051] The formula used for the log-likelihood function is:

[0052]

[0053] where y * is the true label sequence in the training data, and x * is the corresponding input sequence.

[0054] The formula used for updating the weights by gradient descent is:

[0055]

[0056] The above formulas form the basis of the CRF model for the sequence labeling task.

[0057] The formal expression of the relational feature function is:

[0058] f k ((y i-1 , y i ), (y j-1 , y j ), x, i, j)

[0059] where (y i-1 , y i ) and (y j-1 , y j ) respectively represent the previous or next label of the entity pair at the i-th and j-th positions in the text, x is the input text or sequence, k is the feature index, and i and j are the position indices of the entity pair in the text.

[0060] The formal expression of the relational potential function is:

[0061]

[0062] where w k is the weight of feature k.

[0063] Using natural language processing technology, parse the document content into semantic units, associate them with entities in the knowledge graph, and establish semantic associations between documents.

[0064] Specifically, for text preprocessing: First, perform word segmentation to cut the document text into a sequence of words. For example, "Construction project document management system based on BIM" is cut into "based on", "BIM", "of", "construction", "project", "document", "management", "system".

[0065] Subsequently, perform part-of-speech tagging to identify the part of speech of each word. For example, label "based on" as a preposition, "BIM" as a proper noun, and "of" as a particle.

[0066] Filter out stop words, remove words with no practical meaning, restore the words to their basic forms, and identify personal names, place names, organization names, and component names in the document.

[0067] Decompose the document text into sentences and words, analyze the sentence structure, and identify the semantic roles of each component in the sentence.

[0068] Link the named entities in the document to the entities in the knowledge graph, extract the relationships between entities from the document text. For example, extract the "responsible for" relationship between "XX" and "BIM model design" from "XX is responsible for BIM model design", and perform semantic reasoning using the knowledge in the knowledge graph.

[0069] Identify new entities that appear in the document, add them to the knowledge graph, and update the relationships between entities and the attribute information of entities in the knowledge graph.

[0070] More specifically, in part-of-speech tagging, the conditional probability formula used by conditional random fields (CRF) and deep learning is:

[0071]

[0072] where ψ is the potential function and Z(x) is the normalization factor.

[0073] The Levenshtein distance formula used in word form reduction:

[0074] D(i,j) = min(D(i - 1,j) + 1, D(i,j - 1) + 1, D(i - 1,j - 1) + cost(i,j))

[0075] where cost(i,j) is the cost of converting character x i to character y j

[0076] ​Using natural language processing and machine learning techniques, extract information such as keywords, themes, and entities from documents, train a document classification model, automatically classify according to document features, and optimize and adjust in combination with predefined classification rules;

[0077] Specifically, for feature extraction: Use the bag-of-words model to represent the document as a word frequency vector, ignoring word order and grammatical structure; Calculate the TF-IDF value of the word, that is, the ratio of the frequency of the word in the document to the frequency of the word in the entire document collection, which is used to measure the importance of the word; Map the word to a low-dimensional space to retain the semantic information of the word;

[0078] Identify named entities in the document through named entity recognition, analyze the grammatical structure of the document, and extract sentence components and dependency relationships. Use the Latent Dirichlet Allocation model to represent the document as a topic distribution and extract the topic information of the document.

[0079] Subsequently, calculate the occurrence probability of each category in the training set, calculate the occurrence probability of each word in each category, and smooth the conditional probability to prevent the probability from being 0. For each document, calculate its posterior probability in each category, and select the category with the largest posterior probability as the classification result of the document.

[0080] However, when assuming that features are independent of each other, it often leads to a decrease in classification accuracy and is sensitive to the expression form of the input data. Based on this, part-of-speech tagging can also be used to group words according to their parts of speech, reducing the independence between features. Map the word to a low-dimensional space to retain the semantic information of the word and reduce the sensitivity to lemmatization and stop-word processing.

[0081] Use the trained document classification model to classify new documents. Combine predefined classification rules to optimize and adjust the automatic classification results, filter out documents that do not conform to the rules, and correct the documents that do not conform to the rules. Then conduct a manual review of the automatic classification results to ensure classification accuracy. Regularly update the document classification model with new document datasets, and adjust the classification model and classification rules to adapt to new classification requirements.

[0082] Specifically, in the bag-of-words model, the word frequency vector is represented as:

[0083] V(d)=(tf i,d ,tf 2,d ,...,tf n,d )

[0084] where tf i,d is the number of occurrences of word i in document d.

[0085] TF-IDF calculation formula:

[0086] tfidf(t, d, D) = tf(t, d) × idf(t, D)

[0087] Among them, tf(t, d) is the term frequency of word t in document d, and idf(t, D) is the inverse document frequency of word t in document set D.

[0088] IDF calculation formula in inverse document frequency:

[0089]

[0090] Among them, |D| is the total number of documents in the document set, and |d ∈ D: t ∈ d| is the number of documents containing word t.

[0091] Use the Naive Bayes classifier to calculate the posterior probability of a document:

[0092]

[0093] Among them, P(c|d) is the posterior probability that document d belongs to category c, P(d|c) is the probability of document d under category c, and P(c) is the prior probability of category c.

[0094] In document classification, calculate the posterior probability of document d in category c:

[0095]

[0096] Use maximum likelihood estimation to update model parameters:

[0097]

[0098] Among them, θ is the model parameter, and d is the document set.

[0099] Record the version information, modification records, and permission changes of the document on the blockchain, and provide document version control and traceability functions through the blockchain;

[0100] Specifically, DocumentRegistry is used to register and manage the basic information of the document, VersionControl is used to record the version information and modification records of the document, AccessControl is used to manage the permission change records of the document, use Truffle and Hardhat to compile the smart contract, deploy the compiled smart contract to the Ethereum network, and enable the BIM system to interact with the Ethereum blockchain through the API interface.

[0101] When the BIM document is uploaded, the registerDocument function of the smart contract is triggered to record the metadata of the document. And each time the document is updated, the updateVersion function is triggered to record the new version information and modification details. When the document permissions change, the changePermission function is triggered to record the permission change.

[0102] More specifically, upload the document through the BIM system and call the registerDocument function of the DocumentRegistry contract to record the initial state of the document on the blockchain. Each time the document is modified, a new version is generated. Call the updateVersion function of the VersionControl contract. Record the version information on the blockchain. When the access permissions of the document change, call the changePermission function of the AccessControl contract, passing in the document ID, user address, and new permission level, and record the permission change history on the blockchain.

[0103] Combine the document content with the BIM model to establish an association relationship between the document and the BIM model components;

[0104] Specifically, prepare the BIM model and assign a unique identifier to each component in the BIM model.

[0105] For complex and difficult-to-automatically-identify information, manually associate the document content with specific BIM components. For simple information, automatically identify the association between the document content and BIM components according to the existing model mapping rules. The algorithm traverses the document content, identifies the information related to the BIM components, and associates this information with the corresponding component identifiers.

[0106] Design the database structure to store the association information between the document and the BIM components, and store the established association relationship in the database.

[0107] Store the document on the cloud platform, establish an association relationship between the document and the BIM model information, and provide keyword retrieval, semantic retrieval, BIM model retrieval, and classification retrieval methods.

[0108] Specifically, establish an association between the document and the BIM model, and use the defined mapping rules and algorithms to associate the keywords, themes, and entity information in the document metadata with the identifiers of the BIM model components. Create an association table in the cloud database to store the association relationship between the document and the BIM model components.

[0109] Develop a keyword search interface that allows users to enter keywords for searching. Query the document metadata containing the specified keywords in the cloud database and return the search results. Use Elasticsearch to build a semantic retrieval engine and allow users to enter natural language queries. The retrieval engine parses the query intent and returns relevant documents.

[0110] Develop a BIM model retrieval function that allows users to search for documents through BIM model components. The user selects a component in the BIM model, and the system queries the association table and returns all documents associated with that component.

[0111] According to the document classification model, classify the documents into different categories, provide a classification navigation interface, and display all documents under the selected classification after the user makes a selection. Integrate the retrieval function into the document management system for system testing.

[0112] Embodiment 2

[0113] As shown in the Figure 2 accompanying figure, a BIM-based construction project document classification storage and retrieval system includes a knowledge graph engine, an artificial intelligence engine, a blockchain engine, a virtual reality engine, and a document management platform;

[0114] Among them, the knowledge graph engine is used to build and manage the knowledge graph and provide semantic understanding functions; the artificial intelligence engine is used for automatic document classification and semantic analysis; the blockchain engine is used for document version control and traceability; the virtual reality engine is used for three-dimensional document display; the document management platform is used for document storage, retrieval, and management.

[0115] The knowledge graph engine includes a knowledge graph construction module, a knowledge graph management module, and a semantic understanding module;

[0116] Among them, the knowledge graph construction module is used to build the knowledge graph; the knowledge graph management module is used to manage the knowledge graph data; the semantic understanding module is used to perform semantic understanding on the document content.

[0117] The artificial intelligence engine includes a document feature extraction module and a document classification module;

[0118] Among them, the document feature extraction module is used to extract document features; the document classification module is used to classify the documents.

[0119] The blockchain engine includes a blockchain node management module, a consensus algorithm module, and a smart contract module;

[0120] Among them, the blockchain node management module is used to manage blockchain nodes; the consensus algorithm module is used to implement the consensus algorithm; the smart contract module is used to implement smart contract functions.

[0121] The virtual reality engine includes a BIM model loading module, a document association module, and a three-dimensional display module;

[0122] Among them, the BIM model loading module is used to load the BIM model; the document association module is used to establish the association relationship between the document and the BIM model; the three-dimensional display module is used to realize the three-dimensional display of the document.

[0123] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0124] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0125] In addition, each functional unit in various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0126] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A BIM-based construction project document classification storage and retrieval method, characterized in that: The following steps are involved: Step 1: Build a knowledge graph containing BIM model information, entity nodes, relationship edges, attribute information, engineering terms, and professional knowledge; Step 2: Use natural language processing to parse the document content into semantic units, associate them with entities in the knowledge graph, and establish semantic associations between documents; Step 3: Use natural language processing and machine learning to extract keywords, topics, and entity information from documents, train document classification models, automatically classify documents based on their features, and optimize and adjust them based on predefined classification rules; Step 4: Record the document’s version information, modification records, and permission changes on the blockchain, and provide document version control and traceability functions through the blockchain; Step 5: Combine the document content with the BIM model and establish the relationship between the document and the BIM model components; Step 6: Store the document on the cloud platform and establish an association between the document and BIM model information, providing keyword search, semantic search, BIM model search and classification search methods.

2. The BIM-based construction project document classification storage and retrieval method according to claim 1 is characterized by: In step 2, the specific steps of parsing the document content into semantic units are: For text preprocessing, first perform word segmentation, split the document text into word sequences, perform part-of-speech tagging, identify the part of speech of each word, filter stop words, restore words to their basic form, decompose the document text into sentences and words, analyze the sentence structure, identify the semantic role of each component in the sentence, link the named entities in the document with the entities in the knowledge graph, extract the relationship between entities from the document text, and use the knowledge in the knowledge graph for semantic reasoning; Finally, new entities appearing in the document are identified and added to the knowledge graph, the relationships between entities in the knowledge graph are updated, and the attribute information of entities in the knowledge graph is updated.

3. The BIM-based construction project document classification storage and retrieval method according to claim 1 is characterized by: In step three, the bag-of-words model is used to represent the document as a word frequency vector, the TF-IDF value of the word is calculated, the word is mapped to a low-dimensional space, the named entities in the document are identified through named entity recognition, the grammatical structure of the document is analyzed, the sentence components and dependencies are extracted, and the Latent Dirichlet Allocation model is used to represent the document as a topic distribution to extract the topic information of the document.

4. The BIM-based construction project document classification storage and retrieval method according to claim 3 is characterized by: Calculate the probability of occurrence of each category in the training set and the probability of occurrence of each word in each category, smooth the conditional probabilities, calculate the posterior probability of each document in each category, and select the category with the largest posterior probability as the classification result of the document.

5. The BIM-based construction project document classification storage and retrieval method according to claim 3 is characterized by: In the bag-of-words model, the word frequency vector is represented as: V(d)=(tf i,d ,tf 2,d ,...,tf n,d ) Among them, tf i,d is the number of occurrences of word i in document d; TF-IDF calculation formula: tfidf(t,d,D)=tf(t,d)×idf(t,D) Among them, tf(t,d) is the term frequency of word t in document d, and idf(t,D) is the inverse document frequency of word t in document set D.

6. A BIM-based construction project document classification storage and retrieval system, characterized in that: Including knowledge graph engine, artificial intelligence engine, blockchain engine, virtual reality engine, and document management platform; Among them, the knowledge graph engine is used to build and manage knowledge graphs and provide semantic understanding functions; the artificial intelligence engine is used for automatic document classification and semantic analysis; the blockchain engine is used for document version control and traceability; the virtual reality engine is used for three-dimensional display of documents; and the document management platform is used for document storage, retrieval and management.