A BIM Model Review Method Based on Text Classification and Named Entity Recognition
By constructing building code knowledge graphs and using text classification and named entity recognition technologies, the problem of inefficiency of traditional BIM model review systems is solved, and a more efficient and accurate review process is achieved.
Patent Information
- Application Number
- CN202411407582.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-10-10
AI Technical Summary
Traditional BIM model review systems rely on manual labor, have low processing efficiency, and are difficult to keep up with rapidly changing industry standards and specifications, and the weight setting for similar articles in different contexts is unreasonable, resulting in inefficient reviews.
Using methods based on text classification and named entity recognition, building specification knowledge graphs are constructed, BIM review data is obtained through review data extraction strategies, matching rule sets are obtained using knowledge graph matching, and priority is adjusted according to context relevance, and reports are output.
It improves the efficiency and accuracy of BIM model review, can automatically and quickly find matching rules, consider the complex relationships between components, and make more reasonable and comprehensive review decisions.
Smart Images

Figure CN119337474B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model review, and particularly to a BIM model review method based on text classification and named entity recognition. Background Art
[0002] The construction industry is an essential basic industry in people's daily lives. Due to the large population in our country, various types of buildings are needed. Therefore, the development of the construction industry is of great significance to our country. Building Information Modeling (BIM) is a revolutionary technology in the construction field and an advanced development of CAD technology. It is based on the three-dimensional data attributes of a building, uses a three-dimensional geometric model as a carrier, and integrates all information files and material attributes during the construction process of a building project.
[0003] Traditional building compliance reviews are based on professional drawing reviewers comparing and checking two-dimensional CAD drawings of building designs. This method is time-consuming, laborious, and prone to errors. BIM compliance reviews automatically check building designs through BIM. This method greatly liberates human and material resources, can discover more problems during the design stage, and can express review results more intuitively. Currently, BIM compliance reviews are achieved by hard-coding the specifications, that is, manually interpreting the specification articles and translating them into computer-executable languages to review construction projects. This method has high feasibility and relatively high accuracy of review results, but it also has its limitations. First, hard-coding requires interpreting and programming each article one by one. Facing the current huge building compliance review system, translating and coding all articles will consume a lot of human resources, and most article review processes are similar but the review objects are different. Such a method wastes human resources to a certain extent. Second, the hard-coding method is not flexible enough. For articles with similar but inconsistent requirements in different regions, they need to be rewritten, and the articles are constantly revised. Each revision requires upgrading the source code, and the maintainability is poor. In actual projects, the specification standards issued by the state are only the minimum requirements for construction projects, and the different requirements of different projects may be more flexible. Hard-coding for such variable building code requirements will be very cumbersome.
[0004] For example, the patent with the publication number CN116992302A discloses a BIM model review method and device. The method includes: obtaining the codes and multiple attribute information of all components in the BIM model. The codes of the components are generated according to digital coding rules when the BIM model is established. The digital coding rules include multiple nodes and node description information. The key of each node includes the identifier of the node. The values of the root node and child nodes include the identifiers of the next-level nodes, and the value of the leaf node is empty; obtaining the node description information corresponding to each component according to the code of each component; sequentially comparing the node description information corresponding to each component with the multiple attribute information of the same component, and outputting the matching degree of each component; and outputting the review results of each component of the BIM model according to the matching degree of each component. The invention can improve the accuracy and efficiency of BIM model review.
[0005] The above patents all have the problems proposed in this background technology: traditional review systems rely on manual work, with low processing efficiency and difficulty in keeping up with rapidly changing industry standards and specifications; the weight settings of similar articles in different contexts are unreasonable, resulting in low review efficiency. To solve the above problems, this application designs a BIM model review method based on text classification and named entity recognition. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide, in view of the deficiencies of the prior art, a BIM model review method based on text classification and named entity recognition, constructing a building code knowledge graph, obtaining BIM review data through a review data extraction strategy, using the knowledge graph matching to obtain a matching rule set, and adjusting the priority for review and outputting a report. The knowledge graph construction adopts an ontology model and a knowledge extraction network. The review data extraction considers component attributes, spatial relationships, and time dimensions. The matching process combines multi-dimensional semantic matching and reasoning to ensure the accuracy and efficiency of the review.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A BIM model review method based on text classification and named entity recognition, the method includes:
[0009] Obtaining building domain specification texts, constructing a building code ontology according to the building domain specification texts, and constructing a building code knowledge graph according to the building code ontology;
[0010] Processing the BIM model to be reviewed through a review data extraction strategy according to the BIM model to be reviewed, and obtaining a review data file and a context relevance;
[0011] Matching in the building code knowledge graph according to the review data file to obtain a matching rule set;
[0012] Adjust the priority of the matching rule set according to the context relevance, and review the BIM model to be reviewed according to the adjusted matching rule set, and output a review report.
[0013] The construction of the building code knowledge graph includes:
[0014] Determine the building code ontology according to the building domain specification text;
[0015] Determine the keywords and key terms of the domain according to the building code ontology, and build a building code ontology knowledge model from top to bottom. The building code ontology knowledge model includes a top-level ontology, concept subtrees, and sub-layer instances;
[0016] Add the building domain specification text as a knowledge feature to the concept subtree, and logically verify the building code ontology knowledge model according to expert knowledge to determine whether it meets the building domain ontology construction principle. If not, continue to add concept subtrees;
[0017] Construct a knowledge extraction network, process the building domain specification text corresponding to the concept subtree through the knowledge extraction network, extract the building code sequence, and add the building code sequence to the sub-layer instance;
[0018] Construct a building code knowledge graph, instantiate the building code ontology knowledge model, and add it to the building code knowledge graph according to the input strategy.
[0019] The knowledge extraction network includes:
[0020] An input layer that maps the building domain specification text corresponding to the concept subtree through a pre-trained word vector model, converts it into a word vector matrix, and calculates a knowledge feature vector according to the word vector matrix;
[0021] A feature extraction layer that extracts semantic features in the knowledge feature vector through a bidirectional long short-term memory neural network;
[0022] An output layer that calculates a label score matrix by scoring the semantic features, extracts the optimal sequence according to the label score matrix, and decodes the optimal sequence to calculate the building code sequence.
[0023] The input strategy includes:
[0024] Obtain the number of concept branches, keywords, and keyword attributes of the input instantiated building code ontology knowledge model;
[0025] Determine the semantic tree where the instantiated building code ontology knowledge model is located in the building code knowledge graph, traverse all nodes of the semantic tree, extract the keyword pairs with the smallest semantic distance between the nodes and the instantiated building code ontology knowledge model as similar node pairs, and calculate the semantic similarity of the similar node pairs;
[0026] Compare the highest semantic similarity with the similarity threshold. If it is less than the similarity threshold, take the instantiated building code ontology knowledge model as the sibling node of the node with the highest semantic similarity. If it is greater than or equal to the similarity threshold, fuse the instantiated building code ontology knowledge model with the node with the highest semantic similarity.
[0027] The review data extraction strategy includes:
[0028] Perform hierarchical processing on the BIM model to be reviewed according to the component type;
[0029] Extract component attribute information, spatial relationship data, and time dimension data according to the hierarchical BIM model;
[0030] Integrate the component attribute information, spatial relationship data, and time dimension data into a review data file, where the review data file includes component geometric information, attribute information, spatial relationship table, and time node data;
[0031] Model the context information of components and their surrounding associated components through a graph neural network model. According to the constructed context dependency model, calculate the association degree between each component and its associated components. The association degree is dynamically updated through the message passing mechanism between nodes, and the association degrees are weighted and summed to output the context association degree.
[0032] The matching in the building code knowledge graph according to the review data file includes:
[0033] Extract features through a pre-trained model according to the review data file to generate multi-dimensional feature vectors;
[0034] Compare the multi-dimensional feature vectors with the specification clause vectors in the building code knowledge graph through a dynamic semantic matching algorithm, and perform preliminary matching according to the multi-dimensional semantic matching strategy to obtain a candidate answer set;
[0035] Perform semantic reasoning on the candidate answer set according to the relationships between nodes and edges in the building code knowledge graph, and obtain a matching rule set according to the matching parameters, where the matching parameters include semantic similarity, spatial similarity, and time similarity;
[0036] If the matching rule set does not meet the matching threshold, adjust the matching parameters through a feedback optimization mechanism until the matching rule set meets the matching threshold.
[0037] The multi-dimensional semantic matching strategy includes:
[0038] Matching the component geometric attributes in the BIM model and the component geometric specifications in the building code knowledge graph through geometric information features;
[0039] Matching the time node data with the construction time sequence requirements in the building code knowledge graph through time information features;
[0040] Matching the component attribute information with the specification clause terms in the building code knowledge graph through semantic information features.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] 1. By constructing a building code knowledge graph, the present invention transforms complex building code texts into a structured knowledge model, enabling the review process of the BIM model to automatically and quickly search for matching rules in the knowledge graph, improving the efficiency and accuracy of the review;
[0043] 2. By calculating the context relevance, the present invention adjusts the priority of the matching rule set, enabling the review process to not only rely on the direct matching of rules but also consider the complex relationships between components in the BIM model, thereby making more reasonable and comprehensive review decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:
[0045] Figure 1 It is a flowchart of a BIM model review method based on text classification and named entity recognition according to Embodiment 1 of the present invention;
[0046] Figure 2 It is a network structure diagram of knowledge extraction according to Embodiment 1 of the present invention;
[0047] Figure 3 It is a flowchart of the instantiation strategy according to Embodiment 1 of the present invention;
[0048] Figure 4 It is a module structure diagram of a BIM model review system based on text classification and named entity recognition according to Embodiment 2 of the present invention;
[0049] Figure 5 It is a flowchart of the building clause text classification according to Embodiment 2 of the present invention;
[0050] Figure 6 It is a flowchart of the named entity extraction of the building clause text according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0051] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0052] Embodiment 1
[0053] Please refer to Figure 1 , an embodiment provided by the present invention: a BIM model review method based on text classification and named entity recognition, and the specific steps of the method are as follows:
[0054] S1: Obtain the building domain specification text, construct a building specification ontology according to the building domain specification text, and construct a building specification knowledge graph according to the building specification ontology;
[0055] In this step, first obtain text data from multiple building domain specification texts. The specification text can include national building standards, local building regulations, and specification requirements for specific engineering projects. Parse and process these specification texts through natural language processing (NLP) technology to identify the key terms and rules therein;
[0056] S2: According to the BIM model to be reviewed, process the BIM model through a review data extraction strategy to obtain a review data file and a context relevance;
[0057] In this step, data extraction is performed according to the component type of the BIM model, the spatial position of the component in the building structure, the construction stage, etc., specifically including the following aspects:
[0058] Geometric information: Extract the three-dimensional dimensions, shapes, positions, etc. of each component.
[0059] Attribute information: Includes physical attributes such as the material, durability, and load capacity of the component.
[0060] Spatial relationship data: Analyze the relative positions of components with other surrounding components, and calculate the distances, angles, and geometric overlaps between components.
[0061] Time dimension data: According to the construction plan, extract the construction time nodes and installation sequences of each component;
[0062] Based on the information such as the component positions, functional dependencies, and physical connections in the BIM model, use graph neural network (GNN) technology to construct a context dependency model. Through the message passing mechanism, calculate the context relevance between each component and its surrounding components. Components with high context relevance mean they are more important in the entire building structure;
[0063] S3: Match in the building code knowledge graph according to the said review data file to obtain a set of matching rules;
[0064] S4: Adjust the priority of the set of matching rules according to the said context relevance, and review the BIM model to be reviewed according to the adjusted set of matching rules, and output a review report;
[0065] Dynamically adjust the priority of clause parsing according to the specific context of the BIM model, automatically identify specific building types, component types and structural features, assign different weights to clause parsing according to their importance in the building, and adjust the priority of review criteria;
[0066] The construction of the building code knowledge graph depends on the ontology model and natural language processing technology in the construction field. Through the parsing of the code text, keyword extraction and serialization processing, the text information is transformed into a visual knowledge graph. This process combines domain expert knowledge and automated information extraction technology to ensure that the graph not only meets the specific standards of the construction industry, but also can be dynamically updated and expanded;
[0067] The specific steps of the said S1 are as follows:
[0068] S1.1: According to the building code text in the construction field, first use the predefined ontology framework in the construction field to parse the text and identify the important entities and relationships in the text. For example, keywords such as "structural components", "safety standards", and "building materials" involved in the code text will be marked as important terms;
[0069] In this process, the entity recognition model (NER) in natural language processing is used to determine the ontology. Through the trained model, specific terms in the construction field are recognized, and a preliminary ontology structure is established according to the logical relationship between the terms;
[0070] S1.2: After determining the domain keywords and key terms, construct a building code ontology knowledge model in a top-down manner. The hierarchical structure of the model includes the top-level ontology, concept subtrees, and sub-level instances;
[0071] The top-level ontology represents the core knowledge of the construction field, such as the main functional modules of the building, building component types, and standard code clauses; the concept subtrees are the refinement of these core modules, including more specific building materials, component details, and related safety standards; the sub-level instances are the instantiated applications in each specific building project or scenario, such as the actual application cases of specific building materials or specific safety codes;
[0072] S1.3: After constructing a preliminary building ontology knowledge model, the logic of the model is verified through expert knowledge. The expert system checks the position of each node (keyword or term) in the concept subtree to ensure the hierarchical relationship and logical consistency between terms. If some concepts do not conform to the standards in the building field, the ontology structure will be automatically adjusted or new subtrees will continue to be added until the model meets the ontology construction principles in the building field;
[0073] The ontology construction principles in the building field are derived from the five basic criteria proposed by Gruber, including: accuracy, consistency, extensibility, minimum ontology commitment, and coding bias; the accuracy means that the ontology in the field should accurately convey the meaning of the terms it contains, the consistency means that the definition of the ontology is consistent, and the initial definition of the ontology must be logically consistent with the results deduced from it without any contradiction, the extensibility means that in the construction of the ontology, not only should the vocabulary agreed upon in the industry be adopted, but also the scope of use of the concept should be clarified to ensure that the ontology can be easily extended without external interference. As the knowledge of rotating equipment develops and improves, the constructed equipment ontology should maintain extensibility, the minimum ontology commitment means reducing the constraint conditions of the target object of constructing the ontology, and in the process of constructing the equipment ontology, the constraint conditions of terms should be reduced as much as possible to ensure the diversity of the constructed ontology, and the coding bias means that the concept needs to be introduced at the knowledge level rather than relying on the coding at the specific symbol level;
[0074] S1.4: Construct a knowledge extraction network to process the building field specification text in the concept subtree. The knowledge extraction network is a model based on the combination of bidirectional long short-term memory network (BiLSTM) and conditional random field (CRF), which can effectively identify the semantic relationships in the text and extract the key building specification sequences;
[0075] The network first performs word vector representation (word embedding) on the input building specification text, and maps each word to a high-dimensional vector through a pre-trained word embedding model to represent the semantic information of the word.
[0076] Then, the bidirectional long short-term memory network (BiLSTM) captures the dependencies between words in the text through forward and backward time steps, ensuring that it can identify long-distance dependencies across sentences, especially suitable for the complex semantic structures in building specifications.
[0077] Finally, the conditional random field (CRF) is used to optimize and decode the output label sequence to ensure that the extracted building specification sequence meets the expectations. For example, key provisions such as "the requirement for the thickness of the firewall ≥ 0.3m" can be accurately extracted from the specification text;
[0078] Through the knowledge extraction network, the core provisions can be automatically and efficiently extracted from a large amount of building code texts to construct a building code sequence. This not only reduces the workload of manual screening but also greatly improves the efficiency and accuracy of building code review. For different building projects, the parsing strategy of the knowledge extraction network can be quickly adjusted according to project requirements to ensure the flexibility of the review;
[0079] S1.5: After the extraction of the code sequence is completed, further through the instantiation strategy, the building code ontology knowledge model is converted into a building code knowledge graph, and during this process, according to different building scenarios, the structure and node content of the knowledge graph are dynamically adjusted;
[0080] The instantiation of the knowledge graph maps the concepts and code sequences of the building code ontology knowledge model to the nodes and edges in the knowledge graph. For example, concepts such as "wall" and "fire resistance rating" will be instantiated as nodes in the knowledge graph, while the "dependency relationship between the wall and the fire resistance rating" will be represented as an edge in the graph.
[0081] The instantiated knowledge graph can add new nodes and relationships according to the requirements of a specific building project to ensure the dynamic expandability of the graph. For example, for a high-rise building project, nodes for specific fire protection codes and structural stability requirements required for high-rise buildings can be automatically added;
[0082] Please refer to Figure 2 , the knowledge extraction network structure diagram of the embodiment of the present invention. This network performs knowledge mapping, feature extraction, scoring, and optimal sequence derivation through a multi-layer neural network structure, including an input layer, a feature extraction layer, and an output layer, and finally generates a building code sequence. The knowledge extraction network is specifically as follows:
[0083] The input layer is mainly responsible for mapping the building domain code text corresponding to the concept subtree into a high-dimensional vector space to provide a basis for subsequent feature extraction. The building domain code text corresponding to the concept subtree is mapped through a pre-trained word vector model and converted into a word vector matrix, which is then converted into query vector H, key vector K, and value vector V of the input parameters. They are represented by the word embeddings of the text. The attention matrix is calculated through the self-attention mechanism. The calculation formula of the self-attention mechanism is as follows:
[0084]
[0085] where M A represents the attention matrix, Softmax(·) represents the activation function, T represents the matrix transpose, and d represents the linear dimension;
[0086] Adjust the word vector matrix according to the attention matrix, and continuously update it through the network training process to optimize the knowledge explicit features of the word vector matrix. Calculate the knowledge feature vector according to the weight coefficient matrix. The calculation formula of the knowledge feature vector is:
[0087]
[0088] where, V c represents the knowledge feature vector, ω q , ω k and ω v represent the weight coefficient matrix;
[0089] The main function of the feature extraction layer is to extract deep semantic features from the knowledge feature vector and obtain the context information in the time series. This layer includes a bidirectional long short-term memory neural network, which is composed of a forward LSTM and a backward LSTM, and obtains the forward and backward information in the word sequence of the knowledge feature vector respectively. The forward LSTM processes the input sequence in the order from left to right and generates the forward semantic information at each time step. The backward LSTM processes the input sequence in the order from right to left and generates the backward semantic information at each time step, and merges the forward and backward information into the output at the same moment to obtain the semantic features;
[0090] The output of the bidirectional long short-term memory network contains the context information of each position in the input sequence, enabling the network to understand the complex relationships between the words in the sequence. Through the feature extraction layer, the time dependence relationship in the knowledge feature vector is extracted, providing semantic context support for subsequent sequence scoring and optimal sequence derivation;
[0091] The output layer first scores the semantic features to obtain the score sequence of the semantic feature sequence. The scoring process is implemented through a linear transformation and the Softmax function, generating the scores of the labels in each semantic feature and calculating the label score matrix;
[0092] Then, the global evaluation of the score sequence of the semantic feature sequence is performed through the conditional random field, considering the relationship between adjacent labels in the sequence, and judging whether the transition score matrix satisfies global optimality. If it does not satisfy global optimality, recalculate the transition score matrix. If it satisfies global optimality, adjust the label score matrix through the transition score matrix to make the label score matrix more in line with the law of sequence dependence;
[0093] The optimal sequence is obtained by using the Viterbi algorithm based on adjacent labels. The Viterbi algorithm calculates the scores of all possible paths from the beginning of the sequence to the current time step, records the best path of each step, that is, the label sequence with the highest score, and after calculating the best path score of the last step, backtracks and gradually determines the optimal label of each time step, generates the optimal sequence, and decodes the optimal sequence to obtain the maintenance knowledge sequence. This step mainly maps the optimal label sequence back to the original building specification knowledge. The calculation formula of the optimal sequence is:
[0094]
[0095] Among them, B q represents the optimal sequence, Score{·} represents the Viterbi algorithm, j represents the sequence number of a single label in the semantic feature sequence, M represents the total number of labels in the semantic feature sequence, D represents the transition score matrix, and y j represents the jth label in the semantic feature sequence, y j+1 represents the j+1th label in the semantic feature sequence, Q represents the label score matrix, represents matrix multiplication, p represents conditional probability, p(y j 丨y j+1 ) represents the conditional probability of transition between the j+1th label and the jth label in the semantic feature sequence;
[0096] Through global evaluation of random conditional fields and the Viterbi algorithm, the output layer not only selects the highest-scoring label at each time step, but also considers the dependencies of the entire sequence to ensure that the generated label sequence is globally optimal;
[0097] See also Figure 3 , the flowchart of the instantiation strategy disclosed in the present invention, firstly, according to the instantiated building specification ontology knowledge model, the instantiated building specification ontology knowledge model is converted into the number of concept branches, keywords and keyword attribute knowledge for addition, and then similarity detection is performed. The mechanism of similarity detection is to detect whether there is knowledge in the knowledge graph with a similarity greater than a set threshold to the knowledge concept to be entered. The detection process needs to calculate the semantic relevance between the keywords of the knowledge to be entered and the keywords of each record in the database. The expert sets the entry threshold. The knowledge entries above this threshold will be regarded as records with high similarity, indicating that such records already exist in the knowledge base, and this knowledge will not be allowed to be directly entered into the knowledge base. If it is lower than this threshold, it means that this knowledge entry does not exist in the knowledge base, and this knowledge can be entered into the knowledge graph after review;
[0098] The instantiation strategy includes:
[0099] S1.5.1: Obtain the instantiated building code ontology knowledge model, including building code articles, items, keywords, key terms, and their attributes (such as applicable building types, spatial layouts, component relationships, etc.);
[0100] Perform semantic analysis on the code text through natural language processing (NLP) technology, extract the core keywords and their attributes, and construct the concept branches and hierarchical structures. Each code article is regarded as a concept node, and the keywords and attributes below it are used as sub-layer instances of this node;
[0101] S1.5.2: The instantiated building code ontology knowledge model needs to find the best semantic position in the building code knowledge graph. In this process, first determine the semantic tree structure position of the instantiated building code ontology knowledge model in the graph. The semantic tree structure models the semantic and logical relationships of all building code articles to form a semantic tree with a hierarchical structure;
[0102] Once the pair of similar nodes with the smallest semantic distance is found, decide whether to fuse the instantiated building code ontology knowledge model with the target node or use it as the sibling node of the target node according to the similarity threshold;
[0103] If the calculated semantic similarity is less than the similarity threshold, it is considered that there are significant differences between the two articles, and the instantiated building code ontology knowledge model is inserted into the semantic tree as a new node, as the "sibling node" of the existing similar node.
[0104] If the calculated semantic similarity is greater than or equal to the similarity threshold, it is considered that the two nodes have high similarity or consistency, and the two nodes are merged and fused into a common node. The fusion mechanism maintains an efficient and compact knowledge graph structure, avoids redundant data, and through a reasonable node fusion and juxtaposition mechanism, can ensure the clear and definite semantic relationship between articles while maintaining an efficient graph structure;
[0105] Suppose there are the following two code articles:
[0106] Article 1: The thickness of the load-bearing wall of a building shall not be less than 30 cm.
[0107] Article 2: The thickness of the load-bearing wall in a residential building shall not be less than 25 cm.
[0108] When entering these two articles:
[0109] Through NLP parsing, the core keywords of the article are parsed as "load-bearing wall", "thickness", "30 cm", and "25 cm".
[0110] Embed keywords using the BERT model, vectorize the articles, and calculate the similarity. If the similarity between the new terms in Article 2 and Article 1 is greater than the similarity threshold, fuse the semantics of the two articles and mark them as the same node;
[0111] If the semantic similarity is less than the threshold, insert Article 2 as a sibling node of Article 1, and the differences between the articles (such as the special requirements for different building types) can be recorded in the graph;
[0112] The calculation formula for semantic similarity is:
[0113] Sem s = dist α * des β * deg γ ;
[0114] Among them, Sem s represents the semantic similarity between the node and the instantiated disease ontology, dist represents the semantic distance similarity of similar nodes, α represents the influence weight of semantic distance similarity, des represents the depth similarity of similar nodes, β represents the influence weight of depth similarity, deg represents the degree similarity of similar nodes, and γ represents the influence weight of degree similarity;
[0115] The calculation formula for the semantic distance similarity of similar nodes is:
[0116]
[0117] Among them, di(·) represents the semantic distance calculation function, A and B represent pairs of similar nodes, i represents a single keyword in the instantiated disease ontology, n represents the total number of keywords in the instantiated disease ontology, j represents a single keyword of the node, and m represents the total number of keywords of the node;
[0118] The calculation formula for the depth similarity of similar nodes is:
[0119]
[0120] Among them, dep(·) represents the depth calculation function, d 1 represents the average depth of the node, d 2 represents the average depth of the instantiated disease ontology;
[0121] The calculation formula for the degree similarity of similar nodes is:
[0122]
[0123] Among them, E{·} represents the Euclidean distance function, degree(·) represents the degree measurement function, A an represents the parent node of keyword A, O represents the ancestor node of the semantic tree, Ban Represents the parent node of keyword B;
[0124] The BIM model is hierarchically processed according to the component type, extracting attribute information, spatial relationship data, and time dimension data, and context modeling of components and their associated components is performed through a graph neural network (GNN) model. Finally, the context association degree is calculated and applied to generate a review data file. The specific steps of the review data extraction strategy are as follows:
[0125] S2.1: According to the BIM model to be reviewed, first, the model is hierarchically processed according to the type of components (such as walls, beams, columns, doors, and windows). The hierarchical strategy is based on the logical structure of the building, such as spatial location, functional category, and physical connection relationship. The classification of components is not limited to geometric components and can be further subdivided according to attributes such as function (load-bearing components, non-load-bearing components) and material (concrete, steel, wood, etc.);
[0126] The hierarchical processing of components is achieved by parsing the spatial data and structural information of the BIM model and using spatial indexing technology for rapid classification. In this process, the spatial coordinates, dimensions, and type information of each component are extracted and stored in a hierarchical structure, facilitating subsequent multi-dimensional data extraction and review;
[0127] S2.2: The geometric information (such as length, width, and height), material properties, and physical properties (such as durability, fire resistance, etc.) of each component are extracted from the BIM model and stored in a structured manner. The extraction of this information is completed through geometric parsing and attribute parsing algorithms, supporting matching with the requirements of building codes;
[0128] Using spatial indexing and geometric calculation algorithms, the spatial relationships between components (such as relative position, distance, overlap degree, etc.) are extracted. The spatial relationship data will form a spatial relationship table, which is used to describe the position relationship of components in three-dimensional space;
[0129] By synchronizing with the construction plan data, the construction stage, installation time, construction period, and other time dimension information of components are extracted from the BIM model. This data is matched with the construction schedule through time series analysis technology to evaluate the compliance of components during construction;
[0130] The extraction of spatial relationships is based on geometric analysis, using vector operations and spatial query algorithms (such as KD trees, etc.) to determine the spatial characteristics between components. The time dimension data is obtained by analyzing the construction logs and plan data of components and constructing a time series model to capture the construction progress;
[0131] S2.3: Use a graph neural network (GNN) to model the context information of components and their associated components in the BIM model. Components are regarded as nodes in the graph, and the edges between nodes are constructed according to the physical connection relationship, spatial distance, functional dependence, etc. of the components. The node features in the graph network include geometric information, functional attributes, and material properties;
[0132] S2.4: Based on the message passing mechanism in the graph neural network, nodes (components) obtain information from their associated neighbor nodes, and through a multi-layer neural network for information dissemination and aggregation, calculate the context association degree between each node (component) and its neighboring nodes. This association degree reflects the importance of a component in its surrounding structure;
[0133] S2.5: Integrate the geometric information, attribute information, spatial relationship table, and time node data of the components with the context association degree to form a complete review data file;
[0134] Extract features from the review data in the BIM model through a pre-trained model, generate multi-dimensional feature vectors, and perform semantic matching, reasoning, and optimization adjustment with the provisions in the building code knowledge graph. The core principle is to achieve precise matching and compliance review of BIM model components and building code provisions through a dynamic semantic matching and graph network reasoning mechanism based on multi-dimensional parameters such as semantic, spatial, and temporal similarity. The specific steps of S3 are as follows:
[0135] S3.1: Extract the geometric information, attribute information, spatial relationship data, and time dimension data of the components from the BIM model, and perform feature extraction through pre-trained natural language processing models and image processing models. The core of feature extraction is to uniformly convert different behavioral data types in the BIM model into vectorized multi-dimensional representations, and normalize the originally different source and format data (such as geometric data, attribute data, time series data) into unified multi-dimensional vectors, while retaining their semantic and context information. This feature extraction step greatly improves the efficiency and accuracy of subsequent matching and reasoning processes;
[0136] S3.2: According to the dynamic semantic matching algorithm, compare the generated multi-dimensional feature vectors with the provision vectors in the building code knowledge graph. The matching process is based on three similarity parameters - semantic similarity, spatial similarity, and temporal similarity. By calculating the cosine similarity between the feature vector and the code provision vector, obtain a preliminary candidate answer set;
[0137] Semantic similarity: Based on the matching of the attributes of the components and the terms in the code provisions, such as the matching of the material and type of the components with the standard requirements defined in the code;
[0138] Spatial similarity: Based on the matching of the geometric information of components with the spatial layout requirements in the specifications, such as the distance and arrangement between load-bearing walls and supporting beams;
[0139] Temporal similarity: Based on the matching of the construction stage and component time information with the specification time sequence, such as the usage requirements of building materials in specific construction stages;
[0140] S3.3: Based on the candidate answer set generated by the preliminary matching, further use the relationships between nodes and edges in the building code knowledge graph for reasoning. In this stage, a graph neural network (GNN) is used to reason about the context information (such as related components, construction conditions, spatial relationships) between components and code provisions to generate an optimal matching rule set;
[0141] The nodes of components in the building code knowledge graph represent code entities, and the edges represent the relationships between them (such as construction stages, spatial dependencies, etc.). GNN derives deeper associations between code provisions and components through information propagation and update of nodes and edges.
[0142] The introduction of semantic reasoning greatly enhances the intelligence and flexibility. By combining the relationships of different nodes in the graph, it is possible to match not only single code provisions but also composite rules involving multiple code provisions, thus generating a more comprehensive matching rule set;
[0143] S3.4: If the matching rule set fails to reach the preset matching threshold during the semantic reasoning process, a feedback optimization mechanism will be activated to adjust the matching parameters. The specific steps include:
[0144] By analyzing the calculation results of semantic, spatial, and temporal similarities, dynamically adjust their weights and re-perform the matching calculation;
[0145] Using the context correlation information, it is possible to identify the reasons for the matching failure, further optimize the matching weights, and ensure the generation of a more accurate rule set;
[0146] The feedback optimization mechanism ensures that the system can dynamically respond to the review requirements of complex scenarios. By gradually adjusting the matching parameters, the system can adapt to different building types and component environments, thus generating more accurate review results.
[0147] Embodiment 2
[0148] Please refer to Figure 4 , the present invention provides an embodiment: A BIM model review system based on text classification and named entity recognition, the system includes:
[0149] Data acquisition module: Used to acquire building code provisions and the BIM model to be reviewed, and preprocess the code provisions according to the review criteria of the code provisions to obtain the preprocessed code provisions;
[0150] Split and polish the specification clauses. For the clauses lacking a subject after splitting, add the corresponding subject and adjust the word order to facilitate subsequent clause classification and named entity extraction, obtaining the preprocessed specification clauses.
[0151] Clause classification module: used to classify the preprocessed specification clauses according to the review process of the specification clauses, and divide the preprocessed specification clauses into three types: clauses that stipulate the attribute values of BIM model components, clauses that determine whether a BIM model contains a certain component, and clauses that stipulate the distance between two components. Specifically, it includes:
[0152] Clause classification is implemented using a text classification model (RoBERTa-wwm + TextCNN) that combines the Robust optimize BERT approach-Whole Word Masking (RoBERTa-wwm) and Text Convolutional Neural Networks (TextCNN). First, classify and label the preprocessed specification clauses to create a clause classification dataset, and then use the dataset as the input of the RoBERTa-wwm + TextCNN model to enable the model to establish a mapping relationship between the clauses and classification labels, realizing that when a specification clause is input into the model, the model outputs the corresponding classification label.
[0153] Clause keyword extraction module: used to extract the keywords of the preprocessed specification clauses. The extraction basis is whether it describes the building type, whether it describes the component type, whether it describes the attribute type, whether it describes the building space, whether it describes the comparison relationship, and whether it is a number or a unit.
[0154] The extraction of specification clause keywords is implemented using named entity extraction in natural language processing, and a named entity extraction model (RoBERTa-wwm + BiLSTM + CRF) that combines a pre-trained model RoBERTa-wwm with a Bi-directional Long Short-Term Memory (BiLSTM) and a Conditional Random Field (CRF) is used.
[0155] Review module: used to extract the attributes of the BIM model components to be reviewed according to the clause classification results and the extracted keywords, and perform the review.
[0156] According to the clause classification results, different types of clauses correspond to different review processes, and the review process is a set of pre-set review procedures.
[0157] Prepare the BIM model to be reviewed and preprocess the articles to be reviewed in it into the form required by the system.
[0158] Input the article to be reviewed, "The width of the residential windows shall not be less than 0.9m", into the model.
[0159] Use the text classification model (RoBERTa-wwm+TextCNN) to classify the incoming article to be reviewed (such as the building article text). The process and results are as Figure 5 shown. The classification result of the article "not less than" is obtained as "attribute comparison type". The information required to perform the review of the "attribute comparison type" article is: the component to be inspected, the attribute, and the inspection rule. The classification result of the article "residential windows" is "element existence type", and the classification result of the article "0.9m" is "distance comparison type".
[0160] Use the named entity extraction model (RoBERTa-wwm+BiLSTM+CRF) to perform named entity extraction on the incoming article to be reviewed. The named entity extraction process is as Figure 6 shown. The named entity extraction can be performed by passing the article "The width of the residential windows shall not be less than 0.9m" to the model through the API. The extracted building type entity is "residential", the building component entity is "windows", the attribute entity is "width", the comparison relationship entity is "not less than", the numerical entity is "0.9", and the unit entity is "m".
[0161] The system executes the article review process according to the extracted entities:
[0162] 1) If a building type entity is extracted, it is necessary to extract the building type information in the BIM model to be reviewed and check whether it is the extracted building type entity. If not, directly jump to step 2);
[0163] 2) Extract the extracted component entity from the BIM model;
[0164] 3) Extract the component attributes according to the identified attribute entity;
[0165] 4) Convert the extracted attribute values according to the unit entity;
[0166] 5) Compare the numerical entity with the extracted attribute according to the comparison relationship.
[0167] 6) Record the results.
[0168] The system forms the final review list according to the above review results. The list includes the reviewed components, the reviewed attribute values, and the review results.
[0169] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A BIM model review method based on text classification and named entity recognition, characterized in that: The method comprises: Acquire a specification text in the field of construction, construct a construction specification ontology according to the specification text in the field of construction, and construct a construction specification knowledge graph according to the construction specification ontology; The construction of the building specification knowledge graph includes: Determine the building specification ontology according to the standard text in the building field; According to the keywords and key terms of the domain determined by the building specification ontology, a building specification ontology knowledge model is established from top to bottom, wherein the building specification ontology knowledge model includes a top-level ontology, a concept subtree and a sub-layer instance; Add the construction field specification text as a knowledge feature to the concept subtree, perform a logical check on the construction specification ontology knowledge model, and determine whether it meets the construction field ontology construction principle. If not, continue to add concept subtrees; Constructing a knowledge extraction network, processing the construction field specification text corresponding to the concept subtree through the knowledge extraction network, extracting the construction specification sequence, and adding the construction specification sequence to the sublayer instance; Constructing a building specification knowledge graph, instantiating a building specification ontology knowledge model, and adding it to the building specification knowledge graph according to an input strategy; According to the BIM model to be reviewed, the BIM model is processed by a review data extraction strategy to obtain a review data file and context relevance; The data extraction strategy for the review includes: The BIM model to be reviewed is processed in layers according to the component type; According to the layered BIM model, component attribute information, spatial relationship data and time dimension data are extracted; Integrate the component attribute information, spatial relationship data and time dimension data into a review data file, wherein the review data file includes component geometry information, attribute information, spatial relationship table and time node data; The context information of the component and its surrounding associated components is modeled through a graph neural network model. According to the constructed context dependency model, the correlation between each component and its associated components is calculated. The correlation is dynamically updated through a message passing mechanism between nodes. The correlation is weighted and summed to output the context correlation. Matching the review data file in the building specification knowledge graph to obtain a matching rule set; The priority of the matching rule set is adjusted according to the context relevance, and the BIM model to be reviewed is reviewed according to the adjusted matching rule set, and a review report is output.
2. A BIM model review method based on text classification and named entity recognition according to claim 1, characterized in that: The knowledge extraction network comprises: The input layer maps the architectural field specification text corresponding to the concept subtree through a pre-trained word vector model, converts it into a word vector matrix, and calculates the knowledge feature vector according to the word vector matrix; A feature extraction layer extracts semantic features from the knowledge feature vector through a bidirectional long short-term memory neural network; The output layer calculates the label score matrix by scoring the semantic features, extracts the optimal sequence according to the label score matrix, decodes the optimal sequence, and calculates the building specification sequence.
3. A BIM model review method based on text classification and named entity recognition according to claim 2, characterized in that: The entry strategy includes: Obtain the number of concept branches, keywords and keyword attributes of the input instantiated building specification ontology knowledge model; Determine the semantic tree where the instantiated building specification ontology knowledge model is located in the building specification knowledge graph, traverse all nodes of the semantic tree, extract the keyword pairs with the smallest semantic distance between the nodes and the instantiated building specification ontology knowledge model as similar node pairs, and calculate the semantic similarity of the similar node pairs; The highest semantic similarity is compared with the similarity threshold. If it is less than the similarity threshold, the instantiated building specification ontology knowledge model is used as a sibling node of the node with the highest semantic similarity. If it is greater than or equal to the similarity threshold, the instantiated building specification ontology knowledge model is merged with the node with the highest semantic similarity.
4. According to the BIM model review method based on text classification and named entity recognition according to claim 1, it is characterized in that: The matching in the building specification knowledge graph according to the review data file includes: According to the review data file, feature extraction is performed through a pre-trained model to generate a multi-dimensional feature vector; The multi-dimensional feature vector is compared with the specification clause vector in the building specification knowledge graph through a dynamic semantic matching algorithm, and a preliminary match is performed according to the multi-dimensional semantic matching strategy to obtain a candidate answer set; Performing semantic reasoning on the candidate answer set according to the relationship between nodes and edges in the building specification knowledge graph, and obtaining a matching rule set according to matching parameters, wherein the matching parameters include semantic similarity, spatial similarity, and temporal similarity; If the matching rule set does not satisfy the matching threshold, the matching parameters are adjusted through a feedback optimization mechanism until the matching rule set satisfies the matching threshold.
5. A BIM model review method based on text classification and named entity recognition according to claim 4, characterized in that: The multi-dimensional semantic matching strategy includes: Match the geometric properties of components in the BIM model with the geometric specifications of components in the building specification knowledge graph through geometric information features; Matching time node data with construction timing requirements in the building specification knowledge graph through time information features; The component attribute information is matched with the specification terms in the building specification knowledge graph through semantic information features.
Citation Information
Patent Citations
BIM model examination method and device
CN116992302A
Automatic building scheme compiling method and system based on NLP
CN117350166A