Building specification logic rule extraction method based on large language model
Through the building code logic rule extraction method based on large language models, the data in building code documents is transformed into a computer-processable logical rule form, solving the problem of handling complex and diverse data forms and semantic requirements in the prior art, and achieving efficient and accurate logical extraction and automated processing capabilities.
Patent Information
- Application Number
- CN202510014348.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology is difficult to effectively deal with complex and diverse data forms and semantic requirements in building code documents, resulting in inefficient automatic compliance reviews.
The building code logic rule extraction method based on large language models is adopted, and the data preprocessing, model fine-tuning and rule extraction modules are transformed into a computer-processable logical rule form to build a logical rule library.
It significantly improves the operability and automated processing capabilities of standardized content, reduces the workload of manual data labeling, improves processing efficiency and accuracy, and provides a reliable logical foundation to support compliance review and automated design verification.
Smart Images

Figure CN120012900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of large language models, and specifically to a method for extracting logical rules of building specifications based on a large language model. Background Art
[0002] The entire life cycle of a construction project is strictly constrained by various specifications, standards and regulations. Compliance review is an indispensable part of the project implementation process. However, manual compliance review is time-consuming, labor-intensive, costly and error-prone, which limits the improvement of the efficiency of construction project design and approval. To solve this problem, automated compliance checking (ACC) came into being. ACC can significantly accelerate the design and verification process of construction projects through intelligent means. Among them, rule interpretation is the most critical and complex part of compliance review, that is, converting the text-based specification content into computer-processable logical rules. With the rapid development of the construction industry, the complexity and diversity of specification documents are increasing. These documents contain a large number of implicit logical rules, including conditional relationships, requirement standards and restriction rules. However, the diversity and implicitness of rule expression make it difficult for existing automation systems to directly understand and process them. Summary of the invention
[0003] In view of the defects of the prior art that a large number of specific domain data sets need to be constructed to train the model and the processing capacity is limited to natural language text generation, which makes it difficult to adapt to the complex and diverse data forms and semantic requirements in building specification documents, the present invention proposes a logical rule extraction method for building specifications based on a large language model. With the help of the rich semantic knowledge contained in the large language model, different forms of data in the building specifications are extracted into different forms of computer-processable rule forms, and a rich logical rule library is constructed. The operability and automatic processing capability of the specification content are greatly improved, while significantly reducing the workload of manual data annotation and improving processing efficiency and accuracy.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a method for extracting logical rules of building codes based on a large language model. After converting, cleaning and segmenting building code documents, the tables, formula data and text data therein are converted into a form that is easy for computers to process. Then, a TextToRule prompt word is written according to a template, and the content in the prompt word is marked in a form similar to an HTML tag. The large model technology is combined to extract logical knowledge from building code provisions.
[0006] The present invention relates to a system for realizing the above method, comprising: a data preprocessing module, a model fine-tuning module and a rule extraction module, wherein: the data preprocessing module first converts the building specification document PDF into a markdown document with a latex format and performs data cleaning and data segmentation operations to generate specification clauses; the model fine-tuning module fine-tunes the large model through a small fine-tuning data set; the rule extraction module extracts rules from the specification clauses according to prompt words for different data forms in combination with the fine-tuned large model to form a computer-processable logical rule base. Technical Effects
[0007] The present invention realizes efficient and accurate logic extraction of multiple data forms through the automatic parsing and extraction framework of building specifications and the conversion forms of different data types, through prompt words that adapt to the needs, and combined with the semantic capabilities of large language models. Compared with the prior art, the present invention can cover all core information in building specifications, enhance the availability of data, realize the direct calculation, verification and engineering call of specification information, and realize the processing of complex logical relationships in the specifications, providing a reliable logical basis for compliance review and automated design verification. The SQL representation of tabular data improves the dynamic query and management efficiency of data. The Python function processing of formulas supports instant calculation and simulation, providing support for engineering design and optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a flow chart of the present invention;
[0009] Figure 2 It is a schematic diagram of an example of the data in the table of embodiments;
[0010] Figure 3 It is a schematic diagram of an example of formula data of an embodiment;
[0011] Figure 4 This is a schematic diagram of the prompt word template of the embodiment TextToRule;
[0012] Figure 5 This is a schematic diagram of an example of text data in an embodiment. DETAILED DESCRIPTION
[0013] like Figure 1 As shown, this embodiment relates to a method for extracting logical rules of building specifications based on a large language model, including:
[0014] Step 1: Data preprocessing: Convert, clean, and segment building code documents, including:
[0015] 1.1 Collect building specification documents for electricity, water, heating and general use in pdf format.
[0016] 1.2 Use mathpix to convert the pdf document into a Mathpix markdown file to facilitate data processing in subsequent steps.
[0017] 1.3 The content of the building code is organized into chapters, sections, and articles. The building code is divided into articles, and the text, table, and formula data are identified according to the chapter and article numbers. The data in three formats are separated and stored as csv files in the format of <serial number, data content>.
[0018] 1.4 Perform data cleaning on plain text data, remove useless characters, etc. At the same time, manually review the converted files and fix errors in the conversion.
[0019] Step 2: Generate table and formula rules: Convert table and formula data into Figure 3 The computer-processable rule forms shown include:
[0020] 2.1 There are tags such as \begin, \tag and \right in the syntax of Latex. When writing programs, you need to prevent \b, \t and \r from being recognized as escape characters. Anti-escape processing is required. In Python, use repr() to process latex strings.
[0021] 2.2 Write prompt words for TableToSQL and EquationToPython according to requirements, including:
[0022] TableToSQL prompt: You are an expert in converting latex tables to sql tables. Given a latex table, please understand the content of the table and generate a sql statement for the corresponding sql table. The table name is the table number. Please output the sql statement directly without explanation.
[0023] EquationToPython prompt: You are an expert in converting latex formulas into python functions. Given your latex formula and the explanation of related variables, please understand the content of the formula and generate the corresponding python function. The annotation is the explanation of the variables, and the function name is the serial number of the formula. Please give the python function directly without explanation.
[0024] 2.3 Use LangChain to construct the written prompt words and gpt4 model into a chain that can execute the corresponding functions, and then call it to extract rules from table and formula data. The extracted rule information is stored in the form of <serial number, rule information>.
[0025] Step 3: Generating text rules: Converting text data into a rule form that can be processed by a computer, including:
[0026] 3.1 Based on the content of the building code, define the elements of the extracted first-order logic rules, including:
[0027] Entity: description of nouns, the form of unary predicates, e.g. transformer room (x), distribution equipment room (y), capacitor room (z)
[0028] Entity attribute: description of noun class. x is a door entity, which means x is a Class A fire door: type (x, a) ∧ a = "Class A fire door"
[0029] Joint attribute: represents the attribute between two entities, for example: distance(x,y,z) means the distance between x and y is z
[0030] Entity state: Use a unary predicate form. egx is a door entity, and open(x) means the door is in the open state.
[0031] Relationship: a verb-like description. The connection between entities, using at most ternary predicates to represent ternary relationships. For example, adjacent(x,y): indicates that x and y are adjacent; connected(x,y,z): indicates that x is connected to y and z is connected to y.
[0032] 3.2 Extract about 100 data from the building code and mark them, construct a small fine-tuning dataset, and fine-tune the gpt4 model.
[0033] 3.3 According to Figure 4 The template shown writes the TextToRule prompt word and uses HTML-like tags to mark the content in the prompt word.
[0034] The prompt words include: work to be completed, the processing process and the 3-shot technology used.
[0035] The work to be done is: you are an expert in the field of architecture and also an expert in first-order logic rule extraction. Now you need to extract logic rules from the building specification text in latex format. If there are no logic rules, please return Null. Then use <definitions>< / definitions> Tags describe the definition of logic rule elements.
[0036] The processing process is: please extract entities, entity attributes, states and other elements in the text, and then construct logical rules based on the meaning expressed by the text. Please note that the recognition of these elements should be fine-grained, and the variables in the predicate are represented by lowercase letters. Please use them logically and correctly.
[0037] The 3-shot technique used in the above example gives 3 examples for the model to learn. <sample>< / sample> Description. Give the specific processing steps in the sample and use <thinking>< / thinking> Description, including:
[0038] <thinking>
[0039] Step 1: Understand the meaning of the sentence and construct the elements needed to express the logical rules.
[0040] Step 2: Based on the semantic information, identify the predicates belonging to the premise and conclusion parts and determine the number of logical rules that can be extracted.
[0041] Step 3: Based on the results of Step 2, construct logical rules. Note that each lowercase letter has a meaning. Please use them correctly.
[0042] < / thinking>
[0043] 3.4 Use LangChain to construct the written prompt words and the fine-tuned gpt4 model into a chain for extracting logical rules, and call it to extract logical rules from the standard text. The results are also stored in the form of <serial number, rule information>.
[0044] Example results are as follows Figure 5 shown.
[0045] Using the designed table prompts and formula prompts, combined with the GPT-4 model, the LangChain framework was used to conduct an extraction experiment on 120 table data and 45 formula data. The accuracy evaluation index was used to evaluate the extraction results, that is, whether the content expressed in the converted form is correct and consistent with the original. The results are shown in the following table:
[0046] For text data, this paper uses untuned and fine-tuned GPT-3.5 and GPT-4 models, combined with simple 3-shot prompt technology and designed TextToRule prompt, and uses the LangChain framework to conduct experiments on a test set with 50 data. The evaluation indicators are Precision, Recall and F1 score. The results are shown in the following table:
[0047] Compared with the prior art, the present invention uses a fine-tuned large model and a designed TextToRule prompt, which performs well in the task of extracting first-order logic rules from text. Compared with the method using 3-shot combined with a large model, the performance effect is improved by 2 times, significantly improving the extraction efficiency and accuracy. Through the precise design of the TextToRule prompt words and the fine-tuning of the large model, its processing capabilities for specific tasks are improved, making its logical analysis of complex sentences in building specifications more accurate, and greatly improving the Recall index of the extraction results, ensuring the consistency and availability of the extraction results.
[0048] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.
Claims
1. A method for extracting logical rules from building specifications based on a large language model, characterized in that: After converting, cleaning and segmenting the building code documents, the tables, formula data and text data are converted into a form that is easy for computers to process. Then, TextToRule prompts are written according to the template, and the contents of the prompts are marked in the form of HTML tags. The large model technology is used to extract logical knowledge from the building code provisions.
2. The method for extracting logical rules of building specifications based on a large language model according to claim 1 is characterized in that: The operations of converting, cleaning and segmenting the building specification documents specifically include: 1.1 Collect building specification documents on electricity, water, heating and general use in PDF format; 1.2 Use mathpix to convert the pdf document into a Mathpix markdown file to facilitate data processing in subsequent steps; 1.3 The content of the building code is organized into chapters, sections, and articles. The building code is divided into articles. At the same time, the text, table, and formula data are marked according to the chapter and article numbers. The data in three formats are separated and stored as csv files in the format of <serial number, data content>. 1.4 Perform data cleaning on plain text data, remove useless characters, etc., and manually review the converted files to fix errors in the conversion.
3. The method for extracting logical rules of building specifications based on a large language model according to claim 1 is characterized in that: The conversion refers to converting table and formula data into a regular form that can be processed by a computer, specifically including anti-escape processing and using repr() to process latex strings.
4. The method for extracting logical rules of building specifications based on a large language model according to claim 1 is characterized in that: The prompt words include the prompt words of TableToSQL and EquationToPython, where: TableToSQL prompt: You are an expert in converting latex tables into sql tables. Given a latex table, please understand the content expressed in the table and generate a sql statement corresponding to the sql table. The table name is the table number. Please directly output the sql statement without explanation. EquationToPython prompt: You are an expert in converting latex formulas into python functions. Given your latex formula and the explanation of related variables, please understand the content of the formula and generate the corresponding python function. The annotation is the explanation of the variables, and the function name is the serial number of the formula. Please give the python function directly without explanation.
5. The method for extracting logical rules of building specifications based on a large language model according to claim 4 is characterized in that: The written prompt words and gpt4 model are constructed into a chain that can execute the corresponding functions using LangChain, and then it is called to extract rules from table and formula data. The extracted rule information is stored in the form of <serial number, rule information>.
6. The method for extracting logical rules of building specifications based on a large language model according to claim 1 is characterized in that: The step of writing the TextToRule prompt word according to the template specifically includes: 3.1 Define the elements of the extracted first-order logic rules according to the content of the building code; 3.2 Extract several pieces of data from the building code and mark them, construct a small fine-tuning dataset, and fine-tune the gpt4 model; 3.3 Write the TextToRule prompt word according to the template, and use HTML-like tags to mark the content in the prompt word. The prompt word includes: the work to be completed, the processing process, and the 3-shot technology used.
7. The method for extracting logical rules of building specifications based on a large language model according to claim 1 is characterized in that: The logical knowledge extraction refers to: using LangChain to construct the written prompt words and the fine-tuned gpt4 model into a chain for extracting logical rules, calling it to extract logical rules from the standard text, and the results are also stored in the form of <serial number, rule information>.
8. A system for extracting logical rules from building specifications based on a large language model, which implements the method described in any one of claims 1 to 7, characterized in that: include: Data preprocessing module, model fine-tuning module and rule extraction module, among which: the data preprocessing module first converts the building specification document PDF into a markdown document with latex format and performs data cleaning and data segmentation operations to generate specification clauses; the model fine-tuning module fine-tunes the large model through a small fine-tuning data set; the rule extraction module extracts rules from the specification clauses based on prompt words for different data forms and the fine-tuned large model to form a computer-processable logical rule base.
Citation Information
Cited By
Method and system for extracting ecological environment access list rule of intelligent agent
CN121615756A