Automatic interpretation method for security technical specification articles
Through the large language model and the small sample thinking chain prompt method, the security technical specification articles are automatically interpreted and the article parsing tree is generated, which solves the problems of low accuracy and low efficiency of digital transformation of security technical specification articles in the existing technology, and realizes the efficient and precise structured representation of standard articles.
Patent Information
- Application Number
- CN202510439382.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-22
AI Technical Summary
The digital conversion accuracy and efficiency of security technical specifications in the prior art are low. The existing software tools rely on manual writing and entry and review rules, which are costly and have limited results when dealing with complex and multi-meaning clauses, and lack in-depth understanding of context information.
The large language model is used to decompose normative clauses into single-constrained clauses based on the small sample thinking chain prompt, and the article dependency graph structure is automatically converted into a article parsing tree through the article parsing tree generation algorithm, including standardized article hierarchical decomposition, information extraction and article parsing tree generation, and gradual prompt example selection and retrieval enhancement generation technology is used.
It realizes the precise structured representation of the safety technical specifications, improves the interpretation efficiency and accuracy of the specifications, reduces the difficulty of information extraction, and promotes the efficient application of safety technical specifications.
Smart Images

Figure CN120354842A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital technology, and in particular to a method for automatically interpreting safety technical specification articles. Background Art
[0002] In current large language models (LLMs), such as Zhipu Qingyan, ChatGPT, etc., with their advanced semantic understanding and logical reasoning capabilities in the field of natural language processing (NLP), they have quickly become research hotspots; LLMs master language patterns in a large amount of text data through pre-training, especially showing excellent natural language understanding and generation capabilities in the latest models such as GPT-4, providing new possibilities for the automatic interpretation of building codes; safety technical specifications are the core guidelines to ensure the safety of building construction, and the key lies in preventing and reducing the safety threats of safety accidents to the building structure and construction workers; with the accelerating digital pace of the construction industry, the complexity of the construction site is increasing day by day, and the core position of safety technical specifications is becoming more prominent.
[0003] However, there are still the following problems in the process of digital transformation: 1. The digital level of building codes is still in machine-readable files, open digital formats or traditional text formats. Although a few automatic drawing review systems and software have built a large number of drawing review rules according to engineering specifications, they still mainly focus on mandatory articles, and there is still room for improvement in aspects such as user-defined requirements and complex specification article processing; 2. In terms of specification representation forms and reasoning, different software adopts different implementation methods according to their respective characteristics, and a unified technical solution and method system have not been formed; 3. In terms of specification interpretation, existing software tools still mainly rely on manually writing and inputting drawing review rules, with low efficiency and high costs; 4. The methods based on traditional natural language processing have limited effects in dealing with complex and polysemous articles, lack in-depth understanding of the context information in the specifications, and at the same time rely on a large number of manually labeled datasets, with high training costs. Summary of the Invention
[0004] The purpose of the present invention is to overcome the defects of the prior art and provide a method for automatically interpreting safety technical specification articles, so as to solve the problems of low accuracy and low efficiency in the digital transformation of safety technical specification articles in the prior art.
[0005] To achieve the above purpose, the present invention provides a method for automatically interpreting safety technical specification articles and its preparation method. Using a large language model, based on the syntactic structure of the specification article, the specification article is decomposed into single constraint pair articles based on few-shot chain of thought prompting, and then information extraction is performed on the specification article based on prompt engineering to form an article dependency graph as an intermediate representation structure. Finally, the article dependency graph structure is automatically converted into an article parse tree through an article parse tree generation algorithm. The automatic interpretation method includes the following steps:
[0006] S1. Obtain the regulatory provisions and hierarchically decompose the regulatory provisions based on few-shot chain-of-thought prompting;
[0007] S2. Extract information from the regulatory provisions based on prompt engineering;
[0008] S3. Generate a clause parsing tree.
[0009] Furthermore, the few-shot in step S1 is obtained through an incremental prompt example selection method, specifically by manually creating initial examples and then gradually optimizing them into an example set through manual review.
[0010] Furthermore, the hierarchical decomposition of the regulatory provisions in step S1 includes four steps: constraint rule analysis, constraint condition analysis, association analysis of constraint rules and constraint conditions, and semantic orientation analysis. Through these four steps, the input regulatory provisions are decomposed into single-constraint pair clauses;
[0011] Among them, constraint rule analysis is to identify the mutually independent constraint rules in the clause and list the constraint rules in the form of an ordered list;
[0012] Constraint condition analysis is to identify the mutually independent constraint conditions in the clause and list the constraint conditions in the form of an ordered list;
[0013] Association analysis of constraint conditions and constraint rules is to associate the identified constraint conditions and constraint rules to form a clear logical relationship;
[0014] Semantic orientation analysis is a detailed analysis of the semantic orientation in the constraint rules and constraint conditions in combination with the original clause to eliminate cross-constraint reference problems in the clause and ensure the semantic clarity of the clause components.
[0015] Furthermore, step S2 is based on few-shot prompting and chain-of-thought prompting, and further uses few-shot chain-of-thought for information extraction into a clause dependency graph, and uses retrieval-enhanced generation technology to enhance the few-shot chain-of-thought method to improve the effect of the few-shot chain-of-thought.
[0016] Furthermore, the generation of the clause parsing tree in step S3 is an algorithm for automatically converting the clause dependency graph structure of regulatory provision information extraction into a clause parsing tree. After defining an adjacency matrix with an antisymmetric structure, it includes two steps: constructing the adjacency matrix and converting the adjacency matrix into a clause parsing tree.
[0017] Furthermore, the definition of the adjacency matrix with an antisymmetric structure includes that if there is an edge from node i to node j, then A[i][j] = w, representing the out-edge of node i, and at the same time A[i][j] = -w, representing the in-edge of node j, where w is the weight of the edge;
[0018] When w = 1, it represents a non - constraint relationship; when w = 2, it represents a constraint relationship applied to a constraint rule; when w = 3, it represents a constraint relationship applied to a constraint condition.
[0019] If there is no edge between node i and node j, then w = 0. Therefore, the adjacency matrix A is an anti - symmetric matrix, that is, A[i][j] = -A[j][i].
[0020] Furthermore, define the edge set in the clause dependency graph structure as E. The construction of the adjacency matrix includes identifying the largest node identifier in the edge set E to determine the order n of the adjacency matrix A; then, traversing the edge set E, determining the subscripts i and j of the nodes corresponding to the edge according to the identifier of the edge, assigning the weight of the edge to the corresponding element A[i][j], and at the same time assigning the element A[j][i] the opposite of the weight.
[0021] Furthermore, the adjacency - matrix - converted clause parsing tree includes determining the starting node of the clause dependency graph, determining the level of each node, and analyzing the directed edges of the clause dependency graph, converting the multiple relationships of the nodes into a tree structure to form a clause parsing tree.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] 1. Aiming at the syntactic characteristics of safety - technology - specification clauses, using a large - language model to realize the automatic translation of safety - technology - specification clauses, generating a clause parsing tree as a structured representation form of safety - technology - specification clauses, and promoting the accurate transformation and efficient application of safety - technology - specification clauses.
[0024] 2. A method for hierarchical decomposition of specification clauses based on few - shot chain - of - thought prompting. Through few - shot chain - of - thought prompting, complex specification clauses are decomposed into simple clauses, reducing the difficulty of information extraction.
[0025] 3. A retrieval - enhanced information - extraction method based on few - shot chain - of - thought. Through retrieval - enhanced information extraction based on few - shot chain - of - thought, combining chain - of - thought prompting and retrieval enhancement, improving the accuracy and efficiency of specification information extraction.
[0026] 4. It includes a clause - parsing - tree generation algorithm for clause parsing trees, constructs an adjacency matrix, and uses an adjacency - matrix conversion algorithm to automatically convert the clause dependency graph into a clause parsing tree, realizing an accurate structured representation of specification clauses. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic flowchart of the method for automatic translation of safety - technology - specification clauses in the present invention;
[0028] Figure 2 It is a schematic flowchart of the method for selecting sample examples in a progressive manner in the present invention;
[0029] Figure 3 Schematic diagram of the steps for hierarchical decomposition of regulatory provisions based on few-shot chain of thought prompting in the present invention;
[0030] Figure 4 Schematic diagram of the prompting words for hierarchical decomposition of regulatory provisions based on few-shot chain of thought prompting in the present invention;
[0031] Figure 5 Flowchart of each step for information extraction of regulatory provisions based on prompt engineering in the present invention;
[0032] Figure 6 Schematic diagram of few-shot chain of thought examples in the vector database of the present invention;
[0033] Figure 7 Schematic diagram of the steps for generating a clause parse tree in the present invention;
[0034] Figure 8 Schematic diagram of constructing an adjacency matrix according to an edge set in the present invention;
[0035] Figure 9 Schematic diagram of the clause parse tree after conversion of the adjacency matrix in the present invention. Detailed implementation manners
[0036] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0037] Terminology standards: large language model (LLM); clause parse tree (CPT); clause dependency graph (CDG); retrieval-augmented generation (RAG).
[0038] Please refer to the attached Figure 1 , the present invention provides a method for automatic interpretation of safety technical specification clauses. By using a large language model, based on the syntactic structure of the specification clauses and few-shot chain of thought prompting, the specification clauses are decomposed into single-constraint pair clauses, and then based on prompt engineering, information extraction is performed on the specification clauses to form a clause dependency graph as an intermediate representation structure. Finally, through a clause parse tree generation algorithm, the clause dependency graph structure is automatically converted into a clause parse tree. It is characterized in that the automatic interpretation method includes the following steps: S1. Obtain the specification clauses and hierarchically decompose the specification clauses based on few-shot chain of thought prompting; S2. Perform information extraction on the specification clauses based on prompt engineering; S3. Generate a clause parse tree.
[0039] Aiming at the syntactic characteristics of safety technical specification clauses, a large language model is used to realize the automatic interpretation of safety technical specifications, generate a clause parse tree as a structured representation form of safety technical specification clauses, and promote the accurate transformation and efficient application of safety technical specification clauses.
[0040] Please refer to the attached Figure 2, a flowchart showing the selection of samples in a progressive manner. The few-shot samples in step S1 are obtained through a progressive prompt example selection method. Specifically, initial examples are manually created and then gradually optimized into an example set through manual review. Based on the initial examples, according to the performance of the large language model on the target task, examples are dynamically added or replaced. For the provisions that the large language model is prone to confusion or incorrect handling, more relevant examples are added to enable the large language model to better learn this prior knowledge. For the provisions that the large language model can already handle well, relevant examples are reduced or replaced to ensure the balance of the example distribution, avoid overfitting, and reduce the generalization ability of the model.
[0041] Please refer to the appendix Figure 3 , showing the steps of hierarchical decomposition of regulatory provisions under few-shot chain of thought prompting. The hierarchical decomposition of regulatory provisions in step S1 includes four steps: constraint rule analysis, constraint condition analysis, association analysis of constraint rules and constraint conditions, and semantic pointing analysis. Through these four steps, the input regulatory provisions are decomposed into single-constraint pair provisions. Exemplarily, it can be referred to the appendix Figure 4 , showing the prompt template for hierarchical decomposition of regulatory provisions under few-shot chain of thought prompting; Constraint rule analysis is to identify the mutually independent constraint rules in the provision and list the constraint rules in an ordered list; Constraint condition analysis is to identify the mutually independent constraint conditions in the provision and list the constraint conditions in an ordered list; Association analysis of constraint conditions and constraint rules is to associate the identified constraint conditions and constraint rules to form a clear logical relationship; Semantic pointing analysis is to conduct a detailed analysis of the semantic pointing in the constraint rules and constraint conditions in combination with the original provision to eliminate cross-constraint reference problems in the provision and ensure the semantic clarity of the provision components;
[0042] Furthermore, in the step of association analysis of constraint rules and constraint conditions, the constraint conditions and constraint rules are paired to ensure that each pair is logically self-consistent and all parts of the provision are processed; In the step of semantic pointing analysis, each component in the provision is analyzed to determine whether there are reference problems. If so, the object it refers to is clarified to ensure that all reference problems in the constraint are resolved.
[0043] The information extraction of regulatory provisions based on prompt engineering in step S2 is to extract information from regulatory provisions, including entity extraction and relationship extraction; Among them, the nodes of the provision dependency graph are category-independent, so only entity recognition is required without entity naming; Relationship extraction requires the ability to identify relationships and determine the category of the relationships at the same time;
[0044] Further, step S2 is based on few-shot prompting and chain-of-thought prompting, and further uses few-shot chain-of-thought for information extraction into a clause dependency graph. The retrieval-augmented generation technology is used to enhance the few-shot chain-of-thought method to improve the effect of the few-shot chain-of-thought, and a retrieval-augmented information extraction method based on the few-shot chain-of-thought is proposed. Please refer to the appendix Figure 5 , showing the flowcharts of each step of the specification clause information extraction based on prompt engineering. The specific process is to construct the chain of thought for the specification clause information extraction through the chain-of-thought prompting method, including four steps: semantic analysis, pointing analysis, entity extraction, and relationship extraction. After manual review and correction, it forms a structure with the input clause and the output clause dependency graph as Figure 6 Example;
[0045] Further, Figure 5 In the vector database shown in, each example needs to be included in the " <example>< / example> " tag, and the example is in the form of " <example>Slice with "tag" as the delimiter to ensure the integrity of each example; further calculate the semantic relevance score between the input specification clause and the problem part in the sliced examples in the vector database, and select the relevant examples with a relevance score greater than the threshold and merge them into the context of the system prompt in descending order of the relevance score to form a few-shot chain of thought; finally, the large language model generates a new chain of thought and a clause dependency graph structure based on the few-shot chain of thought.
[0046] The clause parsing tree generation in step S3 is an algorithm for automatically converting the clause dependency graph structure extracted from the specification clause information into a clause parsing tree. After defining the adjacency matrix of the antisymmetric structure, it includes two steps: constructing the adjacency matrix and converting the adjacency matrix into a clause parsing tree. Please refer to the appendix Figure 7 , showing the schematic diagrams of each step of the clause parsing tree generation algorithm;
[0047] Defining the adjacency matrix of the antisymmetric structure includes that if there is an edge from node i to node j, then A[i][j]=w, representing the out-edge of node i, and at the same time A[i][j]=-w, representing the in-edge of node j, where w is the weight of the edge; when w = 1, it represents a non-constraint relationship, w = 2 represents a constraint relationship applied to a constraint rule, and w = 3 represents a constraint relationship applied to a constraint condition; if there is no edge between node i and node j, then w = 0. Therefore, the adjacency matrix A is an antisymmetric matrix, that is, A[i][j]=-A[j][i];
[0048] Furthermore, please refer to the appendix Figure 8 , showing the schematic diagram of the initial stage of the clause parsing tree generation algorithm. Define the edge set in the clause dependency graph structure as E. Constructing the adjacency matrix includes identifying the maximum node identifier in the edge set E to determine the order n of the adjacency matrix A; then, traverse the edge set E, determine the subscripts i and j of the nodes corresponding to the edge according to the identifier of the edge, assign the weight of the edge to the corresponding element A[i][j], and at the same time assign the opposite number of the weight to the element A[j][i].
[0049] Please refer to the appendix Figure 9 , showing the schematic diagram of the subsequent stage of the clause parsing tree generation algorithm. The conversion of the adjacency matrix to the clause parsing tree includes determining the starting node of the clause dependency graph, determining the level of each node, and analyzing the directed edges of the clause dependency graph, and converting the multiple relationships of the nodes into a tree structure to form the clause parsing tree; in the clause dependency graph structure, a node may have multiple parent nodes or child nodes, while in the clause parsing tree, each node has only one parent node. Therefore, it is necessary to solve the problem of converting multiple relationships into a tree structure; in addition, the clause dependency graph structure may contain loops, while the clause parsing tree is acyclic. Handling circular dependencies is an important issue in the conversion process, and it is necessary to break the loop to ensure that the clause parsing tree can correctly reflect the structure of the normative clauses while maintaining the integrity of the clause parsing tree; in the conversion process, it is crucial to maintain semantic integrity, and it is necessary to ensure that the clause parsing tree can completely retain the semantic information in the original normative clauses without losing or distorting the semantics of the original text due to the structural conversion;
[0050] Specifically, define a list L to store the levels of each node in the clause dependency graph and initialize it to 0. The list follows the "rank-order access" method. Define stacks S1 and S2 to store the edges to be processed and the processed edges in the form of tuples (i, j) respectively. i and j are the row index and column index of the adjacency matrix A, representing the starting node and the ending node of the edge. The stacks follow the "last in, first out" rule;
[0051] Determine the root node. The root node is the node without a parent node. The root node in the clause parsing tree is unique, while there may be multiple nodes without incoming edges in the clause dependency graph structure; the selection of the root node has an impact on the clause parsing tree but has no impact on the semantics it expresses; since the central object mainly described in the normative clauses generally appears in a more forward position and is also more forward in the node set obtained from the extraction of the normative clause information, traverse in the order of the nodes, and set the first node without a parent node as the root node, that is, all elements A[i][j] of the edges of node i are greater than 0. Set the root node as the "current node", and at the same time write "[root node]" in the first line;
[0052] Start the loop and traverse all nodes in pre-order. In the adjacency matrix A, traverse the elements A[n][j] of the "current node" n. When A[n][j]>0, it indicates that there is an edge between node n and node j. When both (n, j) and (j, n) do not exist in the stack S2, it means that the edge between node n and node j has not been processed. When both of the above conditions are met, push (n, j) into the stack S1. If (j, n) exists in the stack S1, it means that there is a cycle in the graph, and the edge between node n and node j has already been put into the stack when processing the edges of node j. To ensure that each edge is not processed repeatedly, pop (j, n) from the stack S1. At the same time, the level of node j is one level lower than that of node n, that is, L[j]=L[n]+1. In addition, record the number of child nodes of node n. If node n has multiple child nodes or no child nodes, the clause parsing tree should be line-wrapped to ensure the hierarchy of the tree, and at the same time write "|" at the beginning of the new line.
[0053] Pop the top element (n, m) of the stack S1 and push it into the stack S2 at the same time. Set node m as the "current node", and the weight w of the edge e between node n and node m is the value of A[n][m]. First, process the edge e. If the current node is at the beginning of the line, that is, there is only "|" in the current line, write the corresponding number of "-" to represent the level of the node at the beginning of the line. If |w| = 1, it means that the edge e corresponds to a non-constraint relationship, and append "-" to the current line. If |w| = 2, it means that the edge e corresponds to a constraint relationship applied to the constraint rule, and append "-(: constraint relationship)" to the current line. If |w| = 3, it means that the edge e corresponds to a constraint relationship applied to the constraint condition, and append "-(? constraint relationship)" to the current line. Then process node m. If w>0, it means that the edge e is an out-edge of node n, that is, from node n to node m, and append "[node m]" to the current line. Otherwise, it means that the edge e is an in-edge of node n, that is, from node m to node n, and append "[&node m]" to the current line.
[0054] Continue the loop until the stack S1 is empty. Since the stack follows the "last in, first out" rule, the edges of the graph are pushed into and popped out of the stack in the order of depth-first search of the graph, and this order is also the pre-order traversal order of the clause parsing tree. By performing a depth-first search on the graph, a pre-order traversal sequence of the clause parsing tree is constructed, and at the same time, the levels of the tree nodes are recorded. From the pre-order traversal and level traversal, the structure of the tree can be uniquely determined.
[0055] The above embodiments of the present invention have been described in detail in combination with the accompanying drawings. Those of ordinary skill in the art can make various variations of the present invention according to the above description. Therefore, some details in the embodiments should not constitute a limitation to the present invention, and the protection scope of the present invention will be defined by the scope of the appended claims.< / example>
Claims
1. A method for automatically interpreting the provisions of a security technical specification. By using a large language model, the specification provisions are decomposed into single-constraint pair provisions based on the syntactic structure of the provisions and few-shot chain-of-thought prompting. Then, information extraction is performed on the specification provisions based on prompt engineering to form a provision dependency graph as an intermediate representation structure. Finally, the provision dependency graph structure is automatically converted into a provision parsing tree through a provision parsing tree generation algorithm, characterized in that The automatic interpretation method includes the following steps: S1. Obtain the regulatory provisions and hierarchically decompose the regulatory provisions based on few-shot chain-of-thought prompting; S2. Extract information from the regulatory provisions based on prompt engineering; S3. Generate a clause parsing tree.
2. The automatic interpretation method of safety technical specification articles according to claim 1, characterized in that: The few-shot in step S1 is obtained through the progressive prompt example selection method, specifically by manually creating initial examples and then gradually optimizing them into an example set through manual review.
3. The automatic interpretation method of the safety technical specification articles according to claim 1, characterized in that: The hierarchical decomposition of the regulatory provisions in step S1 includes four steps: constraint rule analysis, constraint condition analysis, association analysis of constraint rules and constraint conditions, and semantic reference analysis. Through these four steps, the input regulatory provisions are decomposed into single-constraint pair clauses; Among them, the constraint rule analysis is to identify the mutually independent constraint rules in the clause and list the constraint rules in the form of an ordered list; The constraint condition analysis is to identify the mutually independent constraint conditions in the clause and list the constraint conditions in the form of an ordered list; The association analysis of constraint conditions and constraint rules is to associate the identified constraint conditions and constraint rules to form a clear logical relationship; The semantic reference analysis is a detailed analysis of the semantic references in the constraint rules and constraint conditions in combination with the original clause to eliminate the cross-constraint reference problems in the clause and ensure the semantic clarity of the clause components.
4. The automatic interpretation method of safety technical specification articles according to claim 1, characterized in that: Step S2 is based on few-shot prompting and chain-of-thought prompting, and further uses few-shot chain-of-thought for information extraction into a clause dependency graph, and uses retrieval-augmented generation technology to enhance the few-shot chain-of-thought method to improve the effect of the few-shot chain-of-thought.
5. The automatic interpretation method for safety technical specification articles according to claim 4, characterized in that: The generation of the clause parsing tree in step S3 is an algorithm for automatically converting the clause dependency graph structure extracted from the regulatory provisions information into a clause parsing tree. After defining the adjacency matrix of the anti-symmetric structure, it includes two steps: constructing the adjacency matrix and converting the adjacency matrix into a clause parsing tree.
6. The automatic interpretation method of safety technical specification articles according to claim 5, characterized in that: The definition of the adjacency matrix of the anti-symmetric structure includes that if there is an edge from node i to node j, then A[i][j]=w, indicating the out-edge of node i, and at the same time A[i][j]=-w, indicating the in-edge of node j, where w is the weight of the edge; When w = 1, it represents a non-constraint relationship, when w = 2, it represents a constraint relationship applied to the constraint rule, and when w = 3, it represents a constraint relationship applied to the constraint condition; If there is no edge between node i and node j, then w = 0. Therefore, the adjacency matrix A is an anti-symmetric matrix, that is, A[i][j]=-A[j][i].
7. The automatic interpretation method of the safety technical specification provisions according to claim 6, characterized in that: Define the edge set in the clause dependency graph structure as E. The construction of the adjacency matrix includes identifying the largest node identifier in the edge set E to determine the order n of the adjacency matrix A; then, traversing the edge set E, determining the subscripts i and j of the nodes corresponding to the edge according to the identifier of the edge, assigning the weight of the edge to the corresponding element A[i][j], and at the same time assigning the opposite number of the weight to the element A[j][i].
8. The method for automatically interpreting safety technical specification clauses according to claim 7, characterized in that: The conversion of the adjacency matrix into a clause parsing tree includes determining the starting node of the clause dependency graph, determining the level of each node, and analyzing the directed edges of the clause dependency graph, and converting the multiple relationships of the nodes into a tree structure to form a clause parsing tree.