Knowledge extraction and knowledge graph construction method and system for construction drawing review specifications
By constructing a knowledge graph of construction drawing review specifications through Colabeler annotation and GlobalPointer model, the problems of entity nesting and non-nesting in construction drawing review are solved, intelligent review of construction drawings is realized, and the accuracy and efficiency of review are improved.
Patent Information
- Application Number
- CN202211263033.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-10-14
AI Technical Summary
Construction drawing review is still in the traditional manual review mode, lacking a regular knowledge system, resulting in review quality and efficiency that do not meet the requirements. The existing model has limited effectiveness when faced with situations in the construction drawing review specifications where multiple triples overlap, there are many relationship categories, and entities are nested or non-nested.
The Colabeler annotation tool is used to annotate the construction drawing review specifications. The entity relationship joint extraction model based on GlobalPointer is used, combined with the Neo4j graph database to build a knowledge graph, realize the extraction optimization of entity nesting, and combine it with the BIM model for intelligent drawing review.
It improves the accuracy and speed of construction drawing review, realizes intelligent review of construction drawings, breaks through the traditional manual review situation, and meets the industry's needs for intelligent review.
Smart Images

Figure CN115905553B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent drawing review technology for knowledge extraction and graph construction of construction drawing review specifications, and specifically relates to a method and system for knowledge extraction and knowledge graph construction for construction drawing review specifications. Background Art
[0002] Currently, construction drawing review remains a traditional manual review process. Construction drawing review standards lack a standardized knowledge base, and reviewers' understanding of the standards varies. Consequently, review quality and efficiency fail to meet current requirements. Given the complex structure and strong interdependencies of construction drawing review standards, this review method consumes significant manpower and material resources, even with the aid of review tools.
[0003] Extracting entities and relationships from raw text is a crucial step in building knowledge graphs. With the continuous advancement of the natural language processing (NLP) field in recent years, most neural network models for entity and relationship extraction assume that a sentence contains only one relationship and lacks nested entities. However, existing models have limitations when faced with overlapping triples, multiple relationship categories, and both nested and non-nested entities in construction drawing review specifications. Summary of the Invention
[0004] Purpose of the invention: In response to existing technical problems, the present invention provides a method and system for extracting knowledge and building a knowledge graph for construction drawing review specifications, which solves the problem of exposure deviation and optimizes the extraction effect of entity nesting; the review system that combines the review specification knowledge graph and the BIM model promotes the optimization and upgrading of intelligent review of construction drawings, while also ensuring the speed and accuracy of review.
[0005] Technical solution: This invention proposes a method for extracting knowledge and constructing a knowledge graph for construction drawing review specifications, which specifically includes the following steps:
[0006] S1. Preprocess the content of the construction drawing review specification and use the Colabeler annotation tool to generate labeled text data and obtain the labeled dataset Data; then divide Data into the training set train_data and the validation set dev_data;
[0007] S2. Use the training set train_data to train a model for entity relationship joint extraction based on GlobalPointer using the pre-trained model BERT to obtain the construction drawing review specification entity relationship joint extraction training model Model;
[0008] S3. Input the single sentences in the validation set into the Model model, obtain the entity relationship attributes of each single sentence through the scoring function; use sparse multi-label cross entropy decoding to identify and extract entity attribute relationships, and predict relationship triples; obtain the construction drawing review specification entity attribute relationship joint extraction model;
[0009] S4. Use the knowledge storage mapping algorithm to convert its triples into the Neo4j graph database to complete the storage of construction specification knowledge and build the construction drawing review specification knowledge graph;
[0010] S5. Extract the review model data and match it with the knowledge graph, extract and analyze the BIM construction drawing file data to be reviewed, and convert the matching content into a three-dimensional visual intelligent review result.
[0011] Furthermore, the step S1 specifically includes:
[0012] S1.1. Establish the knowledge system to be extracted: Use the terminology in the specification as the basis for extracting entities; use the meta-attributes that express the relationship between entity objects, design specification knowledge, design specification articles, and design specification documents, as well as the attributes that express inclusion and spatial relationships, as the basis for extracting attributes; use numerical attributes, measures taken, and spatial distance as the basis for extracting attribute values; use orientation, combination, modification, constraint, attribute definition, attribute setting, operation, inclusion, and peer level as the basis for extracting relationships; that is, there are entities, attributes, and attribute value elements to be extracted, and the relationships between these elements must be extracted;
[0013] S1.2. Preprocess the standard construction drawing review specifications, convert long sentences into single sentences, use the Colabeler annotation tool to annotate entities, relationships, and attributes based on the knowledge system, and convert the output labeled text data into experimental data; the structure of the experimental data is:
[0014] {“text”:”Original sentence”,”spo_list”:[{“subject”:”Entity text”,”predicate”:”Relationship type”,”object”:”Entity text”,”subject_type”:”Entity type”,”object_type”:”Entity type”}]}.
[0015] Furthermore, the step S2 specifically includes:
[0016] S2.1. Establish a model for joint extraction of entity relationships based on GlobalPointer, wherein the input of the model for joint extraction of entity relationships based on GlobalPointer is a single sentence and the output is a triple;
[0017] The model of entity relationship joint extraction based on GlobalPointer first performs word segmentation encoding on the text content of each valid input data to obtain token_ids and segment_ids; the token_ids list is:
[0018] token_ids=[X1,X2,....,X n ]
[0019] Perform word segmentation encoding on the entity text subject and entity text object in each triple in the triple list spo_list to obtain token_ids after removing the first and last columns;
[0020] Search the token_ids of subject and object with the token_ids of text to find the first position of the head entity sh, the last position of the head entity st, the first position of the tail entity oh, and the last position of the tail entity ot;
[0021] According to (sh,st), (oh,ot), (sh,oh), (st,ot), we can form subject label, object label, relation head label, and relation tail label respectively;
[0022] Pass token_ids and segment_ids as data to the BERT model to get the vector sequence (h1,h2,...h n )=BERT(x1,x2,...x n );
[0023] The output of BERT is [batch_size, maxlength, hidden_size] which serves as the input of GlobalPointer;
[0024] The first step of GlobalPointer is to convert the BERT output vector into [batch_size, maxlength, head_size*2*heads] through the fully connected layer, where heads represents the number of entity types; head_size represents the output dimension of the linear transformation required by the pointer for each head;
[0025] Through two feed-forward layers, the span representation is calculated depending on the start and end index of the span: q i,ɑ =W q,ɑ h i +b q,ɑ ;k i,ɑ =W k,ɑ h i +bk,ɑ , get the sequence vector sequence [q 1,ɑ ,q 2,ɑ ,...,q n,ɑ ] and [k 1,ɑ , k 2,ɑ ,...,k n,ɑ ]; for the span S[i:j] of type ɑ, the start and end positions are q i,ɑ and k i,ɑ , i and j are the head index and tail index respectively;
[0026] The relative position information is explicitly injected into the model, and the ROPE position encoding is applied to the entity representation to meet The scoring function for the span S[i:j] of type ɑ is:
[0027]
[0028] The output of BERT is fed into GlobalPointer, with heads = 2, S(sh,st) and S(oh,ot) being the head and tail scores of the subject and object respectively. All subjects and objects are identified when S(sh,st)>0 and S(oh,ot)>0, completing the NER task.
[0029] The output of BERT is fed into GlobalPointer, so that heads equals the number of relation categories, and the head relation of the entity is obtained based on (sh,oh);
[0030] The output of BERT is fed into GlobalPointer, so that heads equals the number of relation categories, and the tail relation of the entity is obtained based on (st,ot);
[0031] S2.2. Divide the training set into multiple batches, use each batch to train the model parameters of the entity relationship joint extraction based on GlobalPointer, and obtain the training model Model; optimize by reducing the loss function, the loss function is:
[0032] Where N is the set of negative categories of training samples.
[0033] Furthermore, the step S3 specifically includes:
[0034] S3.1. Read the text content in the valid data of the validation set dev_data and perform word segmentation. Then add the characters "[CLS]" to the beginning of the word segmentation result and add the characters "[SEP]" to the end to obtain the tokens list:
[0035] tokens=[[CLS],X1,X2,...,X j ,...,X maxLength-2 ,[SEP]]
[0036] Among them, X j is the j+1th word element of tokens, j=1,2,…,maxlength, maxlength is the maximum length of the word list;
[0037] The tokens are then mapped with the original text to obtain a mapping list:
[0038] mapping=[[],[0],[1],...,[j+1],...,[maxLength-1],[]];
[0039] S3.2. Perform word segmentation encoding on the text content in the valid data of the validation set dev_data to obtain token_ids and segment_ids, which are then fed into the model for prediction.
[0040] S3.3, satisfying S(sh,st)>0, S(oh,ot)>0, S(sh,oh)>0 and S(st,ot)>0; using sparse multi-label cross entropy decoding: only the subscript corresponding to the positive class is transmitted each time; combined with mapping, a triple list of the construction drawing review specification constraint text is obtained:
[0041] spo_list_pred:[[subject,predicate,object], [subject,predicate,object],
[0042] ...,[subject,predicate,object]
[0043] S3.4. Calculate the F1 value; select the model with the largest F1 value as the joint extraction model for entity attribute relationships in the construction drawing review specification.
[0044] Furthermore, the step S4 specifically includes:
[0045] S4.1. Read the triple list spo_list_pred, get all triples R, and then parse triple A i Triple = {S, P, O};
[0046] S4.2. Design and encapsulate a mode based on the publicly available REST_API, and use this as an interface to connect to the Neo4j graph database address. For database transactions, use the begin_Transaction and commit_transaction modules to start and confirm transactions. At the same time, create database indexes RestNode and RestRelationship for 'entities' and'relationships'.
[0047] S4.3. Obtain the corresponding nodes V of triple.S and triple.O from the entity index s and V o , and check whether V s and V o already exist in the database. If they have been stored, proceed to the next step; otherwise, recreate the nodes and add them to the entity index.
[0048] S4.4. Obtain the corresponding edge E of triple.P from the relationship index P , and check whether E P has been stored in the database. If it already exists, proceed to the next operation; otherwise, create a new directed edge V s →V o , and add it to the relationship index.
[0049] S4.5. Check whether all triples A i have completed the traversal task. If i < n, the task is not completed, and go to step S4.1; otherwise, it means the task has been completed, and proceed to the next step.
[0050] S4.6. Store the construction drawing review specification knowledge in the Neo4j graph database and perform knowledge graph visualization display.
[0051] 6. According to the method for extracting construction drawing review specification knowledge and constructing a knowledge graph described in claim 1, wherein the step S5 specifically includes:
[0052] S5.1. Parse and convert the BIM data into an IFC format file, and at the same time, combine the already constructed knowledge graph to obtain triples (E, R t , S) of the entity, attribute, and relationship sets of the construction drawing specification.
[0053] S5.2. Based on the attribute set A t therein, obtain the required type set T for drawing review; according to the relationship between entities and attributes, obtain the required module set C for drawing review.
[0054] S5.3. Extract the specific data information of each module in the module set C and divide it into a basic information dataset B, a state information dataset G, and an attribute information dataset A to obtain the final drawing review dataset P;
[0055] S5.4. Based on the constraints in the constructed knowledge graph, all data elements in the model data vector group are analyzed and compared with the standard elements in the drawing review specification vector group to determine whether they meet the specification requirements;
[0056] S5.5. Build an HTML5 framework to implement the system's web-based architecture; build a Three.js rendering scene to create scenes, cameras, and light sources, and implement 3D display of building models on the web.
[0057] S5.6. Import the BIM construction drawing file data to be reviewed into the system, and read and analyze, reconstruct component information, and construct geometric space. Click on a component of the model, and the system will automatically generate a review report.
[0058] Based on the same inventive concept, the present invention also provides a system for extracting knowledge and building a knowledge graph for construction drawing review specifications, including:
[0059] The dataset annotation acquisition module is used to pre-process the text data in the standard construction drawing review specification, establish and base the knowledge system, annotate it using the Colabeler annotation tool, and convert the experimental data;
[0060] The construction drawing review specification entity relationship joint extraction model establishment module is used to train the GlobalPointer-based entity relationship joint extraction model using the pre-trained model BERT on the training set to obtain the construction drawing review specification entity relationship joint extraction training model Model;
[0061] The validation set prediction triple extraction module is used to input single sentences from the validation set of text data in the specification constraint clauses of the construction drawing to be reviewed into the model model, and obtain the entity relationship attributes of each single sentence through the scoring function; sparse multi-label cross entropy decoding is used to identify and extract entity attribute relationships, and predict and extract relationship triplets; thus, a joint extraction model of entity attribute relationships in the construction drawing review specification is obtained;
[0062] The construction drawing review specification knowledge storage and graph building module is used to convert the extracted triples into the Neo4j graph database using a knowledge storage mapping algorithm, complete the storage of construction specification knowledge, and build a construction drawing review specification knowledge graph;
[0063] The review result acquisition and display module is used to extract the review model data and match it with the knowledge graph, extract and analyze the BIM construction drawing file data to be reviewed, and convert the matching content into a three-dimensional visual intelligent review result.
[0064] Furthermore, the system adopts a web-based client to receive the BIM construction drawing file input by the user.
[0065] Furthermore, the system also includes a visualization interface for visually displaying the review results.
[0066] Furthermore, the system also includes an audit result file generation module for exporting the audit result in the form of a file.
[0067] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention adopts the Colabeler annotation tool to establish a GlobalPointer entity relationship joint extraction model for construction drawing review specifications to predict and extract nested and non-nested entities, attributes and relationships in the specification text, and then construct triples, and adopts the knowledge storage mapping algorithm to construct the construction drawing review specification knowledge graph with the help of the Neo4j library, and then matches the construction drawing drawing data with the review specification knowledge graph to achieve the intelligent review effect of construction drawings, breaking the traditional manual review situation of construction drawings, breaking through the tedious steps of separate extraction of entity relationships, and having high accuracy and completeness in intelligent review, meeting the industry's intelligent review needs; the present invention is suitable for associating review specification text data containing nested and non-nested entities into structured knowledge, and then realizing the automation of construction drawing review according to knowledge graph matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 This is a flowchart of the method for extracting knowledge and building a knowledge graph for construction drawing review specifications;
[0069] Figure 2 This is a structural diagram of the GlobalPointer-based entity-relationship joint extraction model for construction drawing review specifications;
[0070] Figure 3 A flow chart of the knowledge storage mapping algorithm for construction drawing review specification knowledge;
[0071] Figure 4 Schematic diagram of the composition of the system for knowledge extraction and knowledge graph construction for construction drawing review specifications. DETAILED DESCRIPTION
[0072] The present invention will be described in further detail below with reference to the accompanying drawings.
[0073] The present invention discloses a method for extracting knowledge and constructing a knowledge graph for construction drawing review specifications. Figure 1 As shown, the following steps are included:
[0074] S1. Preprocess the content of the construction drawing review specification and use the Colabeler annotation tool to generate labeled text data and obtain the labeled dataset Data; specifically,
[0075] S1.1. Establish the knowledge system that needs to be extracted: use the terminology part in the specification as the basis for extracting entities; use the meta-attributes that express the relationship between entity objects, design specification knowledge, design specification articles and design specification documents, and the attribute part that expresses the inclusion relationship and spatial relationship as the basis for extracting attributes; use numerical attributes, measures taken and spatial distance as the basis for extracting attribute values; use orientation, combination, modification, constraint, attribute definition, attribute setting, operation, inclusion, same level, etc. as the basis for extracting relationships; that is, there are entities, attributes and attribute value elements that need to be extracted, and the relationship between the elements must be extracted.
[0076] The knowledge of construction drawing review specifications is divided into object attributes, inclusion relationships, measures taken, spatial relationships, specification compliance, terminology clarification and others according to the type of constraint content.
[0077] Triples usually describe the relationship between concepts and entities, between entities, between entities and attributes, and between attributes and attribute values. Specifically, there are the following situations:
[0078] Concept and entity: The relationship between civil buildings and first and second-level civil buildings is that of concept and entity.
[0079] Between entities: usually the same-level relationship and the inclusion relationship. For example, a building includes civil buildings and industrial buildings, which are of the same level. The relationship between structural walls and walls is an inclusion relationship.
[0080] Entity and attribute: For example, the height and fire resistance level of a firewall are the relationship between entities and attributes.
[0081] Attributes and attribute values: For example, the building area of the underground equipment room is 150m 2 .
[0082] S1.2. Preprocess the standard construction drawing review specifications, convert long sentences into single sentences, use the Colabeler annotation tool to annotate entities, relationships, and attributes based on the knowledge system, and convert the output labeled text data into experimental data; the structure of the experimental data is:
[0083] {“text”:”Original sentence”,”spo_list”:[{“subject”:”Entity text”,”predicate”:”Relationship type”,”object”:”Entity text”,”subject_type”:”Entity type”,”object_type”:”Entity type”}]}.
[0084] In the present invention, the standard construction drawing review specification clauses are downloaded from the construction standard website and exported in TXT files as the original data set.
[0085] Use the split function to split the data in the original dataset according to punctuation marks, convert the long sentences into single sentences, define the lines list to store the single sentence data, and use the write function to write the lines line breaks into the standard preprocessed text.
[0086] S2. Use the training set to train a model for entity relationship joint extraction based on GlobalPointer using the pre-trained model BERT to obtain the construction drawing review specification entity relationship joint extraction training model Model; specifically, the following steps are included:
[0087] S2.1. Divide the data in the dataset Data into a training set train_data and a validation set dev_data.
[0088] In this embodiment, in order to test the accuracy of the model, a part of the data set Data is used as a validation set. Specifically, Data is divided into a training data set train_data and a validation data set dev_data in a ratio of 7:3;
[0089] S2.2. Establish a model for entity relationship joint extraction based on GlobalPointer. The input of the model for entity relationship joint extraction based on GlobalPointer is a single sentence, and the output is a triple. Its structure is as follows Figure 2 shown.
[0090] The model of entity relationship joint extraction based on GlobalPointer first performs word segmentation encoding on the text content of each valid input data to obtain token_ids and segment_ids; the token_ids list is:
[0091] token_ids=[X1,X2,....,X n ]; In this embodiment, Google's open source Chinese vocabulary data is used to obtain the word segmentation encoding of character elements.
[0092] Perform word segmentation encoding on the entity text subject and entity text object in each triple in the triple list spo_list to obtain token_ids after removing the first and last columns.
[0093] Search the token_ids of subject and object with the token_ids of text to find the first position sh of the head entity, the last position st of the head entity, the first position oh of the tail entity, and the last position ot of the tail entity.
[0094] According to (sh,st), (oh,ot), (sh,oh), (st,ot), the subject label, object label, relationship head label, and relationship tail label are formed respectively. In this embodiment, at least one label is required, and if there is no label, it is filled with 0.
[0095] Pass token_ids and segment_ids as data to the BERT model to get the vector sequence (h1,h2,...h n )=BERT(x1,x2,...x n ).
[0096] The output of BERT is [batch_size, maxlength, hidden_size], which serves as the input of GlobalPointer.
[0097] The first step of GlobalPointer is to convert the BERT output vector into [batch_size, maxlength, head_size*2*heads] through the fully connected layer, where heads represents the number of entity types; head_size represents the output dimension of the linear transformation required by the pointer for each head; the initialization w weight of the fully connected layer here is normalized lecun initialization, which is a flattening to complete the linear transformation in one go.
[0098] Through two feed-forward layers, the span representation is calculated depending on the start and end index of the span: q i,ɑ =W q,ɑ h i +b q,ɑ ;k i,ɑ =W k,ɑ h i +b k,ɑ , get the sequence vector sequence [q 1,ɑ ,q 2,ɑ ,...,q n,ɑ ] and [k 1,ɑ , k 2,ɑ ,...,k n,ɑ]; for the span S[i:j] of type ɑ, the start and end positions are q i,ɑ and k i,ɑ , i and j are the head index and tail index respectively; in this example, the entity is determined by the index position of the row and column.
[0099] The relative position information is explicitly injected into the model, and the ROPE position encoding is applied to the entity representation to meet The scoring function for the span S[i:j] of type ɑ is:
[0100]
[0101] The output of BERT is fed into GlobalPointer, with heads = 2, S(sh,st) and S(oh,ot) being the head and tail scores of the subject and object respectively. All subjects and objects are identified when S(sh,st)>0 and S(oh,ot)>0, completing the NER task.
[0102] The output of BERT is fed into GlobalPointer, so that heads equals the number of relation categories, and the head relation of the entity is obtained based on (sh,oh).
[0103] The output of BERT is fed into GlobalPointer, so that heads equals the number of relation categories, and the tail relation of the entity is obtained according to (st,ot).
[0104] S2.3. Divide the training set into multiple batches, and use each batch to train the model parameters of the entity relationship joint extraction based on GlobalPointer to obtain the training model Model; the training is optimized by reducing the loss function, and the loss function is:
[0105]
[0106] Where N is the set of negative categories of training samples.
[0107] In this embodiment, when training the GlobalPointer entity relationship joint extraction model for the construction drawing review specification, the library train_generator is used to build a data iterator, and the input parameters include the batch_size of single training samples and the training cycle epochs.
[0108] S3. Input the single sentences in the validation set into the Model model, obtain the entity relationship attributes of each single sentence through the scoring function; use sparse multi-label cross entropy decoding to identify and extract entity attribute relationships, and predict relationship triples; obtain the construction drawing review specification entity attribute relationship joint extraction model Model; specifically including:
[0109] S3.1. Read the text content in the valid data of the validation set dev_data and perform word segmentation. Then add the characters "[CLS]" to the beginning of the word segmentation result and add the characters "[SEP]" to the end to obtain the tokens list:
[0110] tokens=[[CLS],X1,X2,...,X j ,...,X maxLength-2 ,[SEP]]
[0111] Among them, X j It is the j+1th word element of tokens, j=1,2,…,maxlength, maxlength is the maximum length of the word list.
[0112] The tokens are then mapped with the original text to obtain a mapping list:
[0113] mapping=[[],[0],[1],...,[j+1],...,[maxLength-1],[]].
[0114] S3.2. Perform word segmentation encoding on the text content in the valid data of the validation set dev_data to obtain token_ids and segment_ids, which are then fed into the model for prediction.
[0115] S3.3, satisfying S(sh,st)>0, S(oh,ot)>0, S(sh,oh)>0 and S(st,ot)>0; using sparse multi-label cross entropy decoding: only the subscript corresponding to the positive class is transmitted each time; combined with mapping, a triple list of the construction drawing review specification constraint text is obtained:
[0116] spo_list_pred:[[subject,predicate,object],[subject,predicate,object],...,[subject,predicate,object];
[0117] S3.4. Calculate the F1 value; select the model with the largest F1 value as the joint extraction model for entity attribute relationships in the construction drawing review specification.
[0118] In this embodiment, the experimental evaluation criteria are analyzed in terms of Precision, Recall, F1 value, and training loss value. The training loss value includes the loss value of the extracted entity, the loss value of the predicted relationship head, and the loss value of the predicted relationship tail.
[0119] S4. Convert its triples into a Neo4j graph database using the knowledge storage mapping algorithm to complete the storage of construction specification knowledge and construct a knowledge graph for construction drawing review specifications; The flowchart of the knowledge storage mapping algorithm for construction drawing review specifications is as Figure 3 shown.
[0120] Specifically, it includes:
[0121] S4.1. Read the triple list spo_list_pred to obtain all triples R, and then parse the triple R i as Triple = {S, P, O}.
[0122] S4.2. Design and encapsulate a mode based on the public REST_API, and use this as an interface to connect to the Neo4j graph database address. For database transactions, use the begin_Transaction and commit_transaction modules to start and confirm the transactions. At the same time, create database indexes RestNode and RestRelationship for 'entities' and'relationships'.
[0123] S4.3. Obtain the corresponding nodes V s and V o of triple.S and triple.O from the entity index, and check whether V s and V o already exist in the database. If they have been stored, proceed to the next step; otherwise, recreate the nodes and add them to the entity index.
[0124] S4.4. Obtain the corresponding edge E P of triple.P from the relationship index, and check whether E P has been stored in the database. If it already exists, proceed to the next operation; otherwise, create a new directed edge V s →V o and add it to the relationship index.
[0125] S4.5. Check whether all triples R i have completed the traversal task. If i < n, the task is not completed, and go to step S4.1; otherwise, it means the task is completed, and proceed to the next step.
[0126] S4.6. The knowledge of construction drawing review specifications is stored in the Neo4j graph database and visualized as a knowledge graph.
[0127] S5. Extract the review model data and match it with the knowledge graph, extract and analyze the BIM construction drawing file data to be reviewed, and convert the matching content into a 3D visual intelligent review result. Specifically including:
[0128] S5.1. Analyze and convert BIM data into IFC format files, and combine them with the constructed knowledge graph to obtain the triples of entities, attributes and relationship sets (E, R t ,S).
[0129] S5.2, according to the attribute set R t , we get the type set T required for drawing review; based on the relationship between entities and attributes, we get the module set C required for drawing review.
[0130] S5.3. Extract the specific data information of each module in the module set C and divide it into a basic information dataset B, a status information dataset G, and an attribute information dataset A to obtain the final drawing review dataset P.
[0131] S5.4. Based on the constraints in the constructed knowledge graph, all data elements in the model data vector group are analyzed and compared with the standard elements in the drawing review specification vector group to determine whether they meet the specification requirements.
[0132] In this example, the drawing review specification vector group includes a geometric information drawing review specification vector group and a module / spatial information drawing review specification vector group.
[0133] S5.5. Build the HTML5 framework to complete the implementation of the system's web-based architecture; build a Three.js rendering scene to create scenes, cameras, and light sources to achieve three-dimensional display of building models on the web.
[0134] S5.6. Import the BIM construction drawing file data to be reviewed into the system, and read and analyze, reconstruct component information, and construct geometric space. Click on a component of the model, and the system will automatically generate a review report.
[0135] This example preprocesses and annotates 17,730 construction drawing review specifications, then uses the aforementioned method to extract knowledge and construct a graph of these specifications. The GlobalPointer entity-relationship joint extraction model effectively extracts triples from review specification sentences containing both nested and non-nested entities, as shown in Table 1.
[0136] Table 1 is a sample of triples extracted from some sentences.
[0137]
[0138] The GlobalPointer entity-relationship joint extraction model for construction drawing review specifications has an identification accuracy of 89.36%, a recall rate of 91.80%, and an F1 value of 90.58% on the validation set.
[0139] Based on the same inventive concept, the present invention also provides a knowledge extraction and knowledge graph construction system for construction drawing review specifications, such as Figure 4 Shown, including:
[0140] The dataset annotation acquisition module is used to perform standard preprocessing on the text data in the standard construction drawing review specification clauses according to step S1, establish and base the knowledge system, annotate using the Colabeler annotation tool, and transform the experimental data.
[0141] The construction drawing review specification entity relationship joint extraction model establishment module is used to train the GlobalPointer-based entity relationship joint extraction model using the pre-trained model BERT on the training set according to step S2 to obtain the construction drawing review specification entity relationship joint extraction training model Model.
[0142] The validation set prediction triplet extraction module is used to input the single sentence in the validation set of the text data in the specification constraint clauses of the construction drawing to be reviewed in step S3 into the Model model, and obtain the entity relationship attributes of each single sentence through the scoring function; use sparse multi-label cross entropy decoding to perform entity attribute relationship identification and extraction, predict and extract the relationship triples; and obtain the construction drawing review specification entity attribute relationship joint extraction model Model.
[0143] The construction drawing review specification knowledge storage and graph establishment module is used to convert the extracted triples into the Neo4j graph database using the knowledge storage mapping algorithm according to step S4, complete the storage of construction specification knowledge, and build the construction drawing review specification knowledge graph.
[0144] The review result acquisition and display module is used to extract the review model data and match it with the knowledge graph according to step S5, extract and analyze the BIM construction drawing file data to be reviewed, and complete the conversion of the matching content into a three-dimensional visual intelligent review result.
[0145] In this embodiment, a web-based client is used to receive user-entered BIM construction drawing files. Furthermore, users can view model information and review content directly in the browser, which is convenient and quick, without the need to install specialized model generation software. The system is also scalable, facilitating the subsequent addition and modification of content. A visual interface is also included for visually displaying review results; review results can also be exported as files using the review result file generation module.
[0146] The above description is merely an example of the present invention and is not intended to limit the present invention. Any equivalent substitutions made within the principles of the present invention are intended to be included within the scope of protection of the present invention. Any content not elaborated in detail herein is already known to those skilled in the art.
Claims
1. A method for extracting knowledge and building a knowledge graph for construction drawing review specifications, characterized by: The steps include: S1. Preprocess the content of the construction drawing review specification and use the Colabeler annotation tool to generate labeled text data and obtain the labeled dataset Data; then divide Data into the training set train_data and the validation set dev_data; S2. Use the training set train_data to train a model for entity relationship joint extraction based on GlobalPointer using the pre-trained model BERT to obtain the construction drawing review specification entity relationship joint extraction training model Model; S3. Input the single sentence in the validation set into the Model model and obtain the entity relationship attributes of each single sentence through the scoring function; Sparse multi-label cross entropy decoding is used to identify and extract entity attribute relationships and predict relationship triples; a joint extraction model of entity attribute relationships in the construction drawing review specification is obtained; S4. Use the knowledge storage mapping algorithm to convert its triples into the Neo4j graph database to complete the storage of construction specification knowledge and build the construction drawing review specification knowledge graph; S5. Extract the review model data and match it with the knowledge graph. Extract and analyze the BIM construction drawing file data to be reviewed, and convert the matched content into a 3D visual intelligent review result. The step S2 specifically includes: S2.
1. Establish a model for joint extraction of entity relationships based on GlobalPointer, wherein the input of the model for joint extraction of entity relationships based on GlobalPointer is a single sentence and the output is a triple; The model of entity relationship joint extraction based on GlobalPointer first performs word segmentation encoding on the text content of each valid input data to obtain token_ids and segment_ids; the token_ids list is: token_ids=[X'1,X'2,X'3,...,X' n-1 ,X' n ] Perform word segmentation encoding on the entity text subject and entity text object in each triple in the triple list spo_list to obtain token_ids after removing the first and last columns; Search the token_ids of subject and object with the token_ids of text to find the first position of the head entity sh, the last position of the head entity st, the first position of the tail entity oh, and the last position of the tail entity ot; According to (sh,st), (oh,ot), (sh,oh), (st,ot), we can form subject label, object label, relation head label, and relation tail label respectively; Pass token_ids and segment_ids as data to the BERT model to get the vector sequence (h1,h2,...h n )=BERT(x1,x2,...x n ); The output of BERT is [batch_size, maxlength, hidden_size] and is used as the input of GlobalPointer; The first step of GlobalPointer is to convert the BERT output vector into [batch_size, maxlength, head_size*2*heads] through the fully connected layer, where heads represents the number of entity types; head_size represents the output dimension of the linear transformation required by the pointer for each head; Through two feed-forward layers, the span representation is calculated depending on the start and end index of the span: q i,ɑ =W q,ɑ h i +b q,ɑ ;k i,ɑ =W k,ɑ h i +b k,ɑ , get the sequence vector sequence [q 1,ɑ ,q 2,ɑ ,...,q n,ɑ ] and [k 1,ɑ , k 2,ɑ ,...,k n,ɑ ]; for span S[i:j] of type ɑ, the start and end positions are q i,ɑ and k i,ɑ , i and j are the head index and tail index respectively; The relative position information is explicitly injected into the model, and the ROPE position encoding is applied to the entity representation to meet For the span S[i:j] of type ɑ, the scoring function S ɑ (i,j) is: The output of BERT is fed into GlobalPointer, with heads = 2, S(sh,st) and S(oh,ot) being the head and tail scores of the subject and object respectively. All subjects and objects are identified when S(sh,st)>0 and S(oh,ot)>0, completing the NER task. The output of BERT is fed into GlobalPointer, so that heads equals the number of relation categories, and the head relation of the entity is obtained based on (sh,oh); The output of BERT is fed into GlobalPointer, so that heads equals the number of relation categories, and the tail relation of the entity is obtained based on (st,ot); S2.
2. Divide the training set into multiple batches, use each batch to train the model parameters of the entity relationship joint extraction based on GlobalPointer, and obtain the training model Model; optimize by reducing the loss function, the loss function is: Where N is the set of negative categories of training samples.
2. The method for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 1 is characterized in that: The step S1 specifically includes: S1.
1. Establish the knowledge system to be extracted: Use the terminology in the specification as the basis for extracting entities; use the meta-attributes that express the relationship between entity objects, design specification knowledge, design specification articles, and design specification documents, as well as the attributes that express inclusion and spatial relationships, as the basis for extracting attributes; use numerical attributes, measures taken, and spatial distance as the basis for extracting attribute values; use orientation, combination, modification, constraint, attribute definition, attribute setting, operation, inclusion, and same level as the basis for extracting relationships; that is, there are entities, attributes, and attribute value elements to be extracted, and the relationships between these elements must be extracted; S1.
2. Preprocess the standard construction drawing review specifications, convert long sentences into single sentences, use the Colabeler annotation tool to annotate entities, relationships, and attributes based on the knowledge system, and convert the output labeled text data into experimental data; the structure of the experimental data is: {"text":"Original sentence","spo_list":[{"subject":"Entity text","predicate":"Relationship type","object":"Entity text","subject_type":"Entity type","object_type":"Entity type"}]}.
3. The method for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 1 is characterized in that: The step S3 specifically includes: S3.
1. Read the text content in the valid data of the validation set dev_data and perform word segmentation. Then add the characters "[CLS]" to the beginning of the word segmentation result and add the characters "[SEP]" to the end to obtain the tokens list: tokens=[[CLS],X1,X2,...,X j ,...,X maxLength-2 ,[SEP]] Among them, X j is the j+1th word element of tokens, j=1,2,…,maxlength, maxlength is the maximum length of the word list; The tokens are then mapped to the original text to obtain a mapping list: mapping = [[], [0], [1],..., [j + 1],..., [maxLength - 1], []]; S3.
2. Tokenize and encode the text content text in the valid data of the validation set dev_data to obtain token_ids and segment_ids, and send them into the model Model for prediction; S3.
3. Satisfy S(sh, st)>0, S(oh, ot)>0, S(sh, oh)>0, and S(st, ot)>0; decoded by sparse multi-label cross-entropy: only transmit the subscripts corresponding to the positive classes each time; combined with mapping, obtain the triple list of the construction drawing review specification constraint text: spo_list_pred: [[subject, predicate, object], [subject, predicate, object],..., [subject, predicate, object]] S3.
4. Calculate the F1 value; select the model with the largest F1 value as the construction drawing review specification entity attribute relationship joint extraction model Model.
4. The method for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 1 is characterized in that: The specific steps of step S4 include: S4.
1. Read the triple list spo_list_pred, get all triples R, and then parse the triples R i ; S4.
2. Design and encapsulate a mode based on the public REST_API, and use this as an interface to connect to the Neo4j graph database address. For database transactions, use the begin_Transaction and commit_transaction modules to start and determine the transactions; at the same time, create database indexes RestNode and RestRelationship for 'entities' and'relationships'; S4.
3. Get the corresponding nodes V of triple.S and triple.O from the entity index s and V o , check whether V already exists in the database s and V o If it has been stored, proceed to the next step; otherwise, recreate the node and add it to the entity index; S4.
4. Get the corresponding edge E of triple.P from the relation index P , check whether E is stored in the database P If it already exists, proceed to the next step, otherwise create a new V s →V o and add it to the relationship index; S4.
5. Check whether all triples R have completed the traversal task. If i < n, the task is not completed, and go to step S4.1; otherwise, it means the task is completed, and proceed to the next step; S4.
6. Store the construction drawing review specification knowledge in the Neo4j graph database and perform knowledge graph visualization display.
5. The method for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 1 is characterized in that: The specific steps of step S5 include: S5.
1. Analyze and convert BIM data into IFC format files, and combine them with the constructed knowledge graph to obtain the triples of entities, attributes and relationship sets (E, R t ,S); S5.2, according to the attribute set R t , get the type set T required for drawing review; according to the relationship between entities and attributes, get the module set C required for drawing review; S5.
3. Extract the specific data information of each module in the module set C, and divide it into a basic information data set B, a status information data set G, and an attribute information data set A to obtain the final drawing review data set P; S5.
4. Based on the limited conditions in the constructed knowledge graph, by analyzing and comparing all data elements in the model data vector group with the standard elements in the drawing review specification vector group, determine whether it meets the specification requirements; S5.
5. Build an HTML5 framework to complete the implementation of the system web page architecture; build a Three.js rendering scene to create a scene, a camera, and a light source to achieve three-dimensional display of the building model on the web page; S5.
6. Import the BIM construction drawing file data to be reviewed into the system, execute the steps of reading and parsing, component information reconstruction, and geometric space construction. Click on a component of the model, and the system automatically generates a drawing review report.
6. A system for extracting knowledge from construction drawing review specifications and building a knowledge graph using the method according to any one of claims 1 to 5, characterized in that: The system includes: The dataset annotation acquisition module is used to pre-process the text data in the standard construction drawing review specification, establish a knowledge system, annotate it using the Colabeler annotation tool, and convert experimental data; The construction drawing review specification entity relationship joint extraction model establishment module is used to train the GlobalPointer-based entity relationship joint extraction model using the pre-trained model BERT on the training set to obtain the construction drawing review specification entity relationship joint extraction training model Model; The validation set prediction triple extraction module is used to input single sentences from the validation set of text data in the specification constraint clauses of the construction drawing to be reviewed into the model model, and obtain the entity relationship attributes of each single sentence through the scoring function; sparse multi-label cross entropy decoding is used to identify and extract entity attribute relationships, and predict and extract relationship triplets; thus, a joint extraction model of entity attribute relationships in the construction drawing review specification is obtained; The construction drawing review specification knowledge storage and graph building module is used to convert the extracted triples into the Neo4j graph database using a knowledge storage mapping algorithm, complete the storage of construction specification knowledge, and build a construction drawing review specification knowledge graph; The review result acquisition and display module is used to extract the review model data and match it with the knowledge graph, extract and analyze the BIM construction drawing file data to be reviewed, and convert the matching content into a three-dimensional visual intelligent review result.
7. The system for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 6 is characterized in that: The system adopts a client based on a web page to receive the BIM construction drawing file input by the user.
8. The system for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 6 is characterized in that: The system also includes a visualization interface for visually displaying the review results.
9. The system for extracting knowledge and building a knowledge graph for construction drawing review specifications according to claim 6 is characterized in that: The system also includes an audit result file generation module for exporting the audit result in the form of a file.
Citation Information
Patent Citations
Innovative knowledge map construction method and device based on concept element extraction
CN114328958A
Building specification examination method and system based on BiLSTM and knowledge graph
CN114880468A