Knowledge graph-fused large-scale language model power grid rule and regulation semantic retrieval method

By constructing a combination of dual-structure knowledge graph and language model, the modeling problem of responsibility distribution and semantic inheritance relationship in power grid regulations is solved, realizing highly reliable structured semantic retrieval and interpretive retrieval in the power system, and supporting compliance query and fault location of power grid regulations.

CN120873208AActive Publication Date: 2025-10-31GUANGDONG POWER GRID CO LTD

Patent Information

Application Number
CN202511393861.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-10-31
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing power grid regulations retrieval systems struggle to comprehensively model complex responsibility distributions, semantic inheritance relationships, and semantic paths, resulting in fragmented retrieval results, lack of context, and weak explanatory power, failing to meet the practical application needs of the power industry.

Method used

A dual-structure knowledge graph is constructed with clauses as nodes and responsibility collaboration and hierarchical inheritance as edges. Structured semantic retrieval is performed by combining a language model, and fast path location is supported by building an index table. Furthermore, a structure-aware language model is introduced to parse user queries and generate structured query triples as the entry point for graph path retrieval.

Benefits of technology

It realizes the reconstruction of the responsibility chain in power grid regulations, the unification of hierarchical regulations, and the semantic output of closed-loop paths. It provides a reproducible, highly reliable, and strongly interpretable semantic retrieval scheme, supporting compliance queries and fault review in the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873208A_ABST
    Figure CN120873208A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge graph fused large language model power grid rule and regulation semantic retrieval method, which comprises the following steps of: calling and obtaining a rule document through a company rule document storage system API (Application Program Interface), performing structure recognition processing on the rule document, inputting the rule document into a semantic recognition model, outputting a term semantic unit set, and storing the term semantic unit set into a database; a double-structure knowledge graph suitable for a power grid rule and regulation retrieval scene is constructed based on a clause semantic unit set, a natural language question input by a user is mapped into structured semantic representation, and multi-hop retrieval of a clause set related to user question semantics and a path structure of the clause set is completed. All path structure representations are matched with structured semantics, and the first several paths with the highest scores according to cosine similarity serve as retrieval results. A structure-driven retrieval system capable of realizing responsibility chain reconstruction and path closed-loop semantic output is constructed, and a reliable semantic retrieval scheme is provided for key businesses such as compliance query, fault redisk and responsibility positioning of a power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of semantic retrieval, and in particular relates to a semantic retrieval method for power grid regulations and systems based on a large-scale language model that integrates knowledge graphs. Background Technology

[0002] As the complexity of power system operations continues to rise, power grid companies have accumulated a large number of highly regulatory rules and regulations in business scenarios such as equipment operation and maintenance, dispatch control, and fault handling. These documents come from diverse sources, including national laws and regulations, industry technical standards, and operational details and management specifications compiled by the group headquarters and its subsidiaries. These documents are generally expressed in natural language, and their content typically possesses structured features such as clause numbers, chapter levels, responsible parties, and applicable scenarios. However, existing retrieval systems generally ignore these structural dimensions, relying solely on keywords, word frequency, or shallow semantics for matching, which fails to meet the needs of semantic reasoning and responsibility chain tracing. Especially in the power industry, the handling of the same event often involves the collaboration of multiple positions. Different regulations specify the behavioral obligations of roles such as dispatching, maintenance, and safety supervision, but existing systems struggle to reconstruct the collaborative relationships between these clauses, resulting in fragmented search results, lack of context, and weak interpretability. Furthermore, due to the significant hierarchical differences in regulations, the same operational process may be described to different degrees in national standards and corporate rules, and existing methods lack the ability to model these cross-level semantic inheritance relationships. Furthermore, while some large-model-based text question-answering systems can improve language understanding capabilities, they lack encoding methods for the structure of regulations and cannot embed structured knowledge from existing regulations into practical systems, such as job categories, event types, and responsibility chain paths. Therefore, they cannot solve the problems of structural accuracy and responsibility consistency in retrieval. Consequently, existing technologies struggle to achieve comprehensive modeling and interpretable retrieval capabilities for the complex distribution of responsibilities, semantic inheritance relationships, and semantic path closures in power grid regulations and documents, limiting their feasibility for practical application in the power industry. Summary of the Invention

[0003] This invention addresses the challenges of complex semantics, multidimensional responsibilities, and strong hierarchical structure in power grid regulations by proposing a structured semantic retrieval method that integrates knowledge graphs and language models. This method constructs a dual-structure knowledge graph with clauses as nodes and responsibility collaboration and hierarchical inheritance as edges. It uniformly extracts and structurally models elements such as roles, behaviors, objects, and regulation levels from the regulatory text, and uses object fields to build an index table to support rapid path location.

[0004] To achieve the above objectives, a first aspect of the present invention provides a semantic retrieval method for power grid regulations based on a large-scale language model that integrates knowledge graphs, comprising: The company's regulations document storage system API is used to obtain regulations documents. Each regulation document contains metadata fields such as source number, internal chapter number, and regulation level identifier. The regulations documents are then processed for structural recognition, and all regulations documents are aggregated to obtain a regulations document set. The regulations document set is then input into a semantic recognition model to output a set of clause semantic units. Based on the set of semantic units of the clauses, a dual-structure knowledge graph suitable for the retrieval scenario of power grid rules and regulations is constructed, and a structural semantic index table is generated simultaneously. By mapping each semantic unit of the clauses in the set of semantic units of the clauses to a graph node, the edge set of the dual-structure knowledge graph includes rule hierarchy inheritance edges and responsibility collaboration chain edges. The system maps user-input natural language questions into structured semantic representations and, based on the constructed dual-structure knowledge graph of power grid regulations and the structured semantic index table, performs multi-hop retrieval of the set of articles and their path structures related to the semantics of the user's question. Jump links are extracted from the dual-structure knowledge graph, including paths from the starting article node to another article node. By setting a maximum number of hops of 4, the top k paths with the highest path scores are selected, and the final matching path set is obtained. Based on the final matched path set and the set of hit clauses, structured semantic modeling is performed to generate a structured aggregate representation, all node codes, and edge codes for each path. The node codes and edge codes are combined into a path structure representation. All path structure representations are matched against the structured semantics, and sorted according to cosine similarity. The path with the highest cosine similarity score is selected. The path was used as the search result.

[0005] Preferably, the semantic recognition model has the following structure: The input is a text sequence with a length of less than or equal to 256 characters. Contextual dependency information is extracted through a bidirectional LSTM text encoding layer, followed by a convolutional feature extractor to identify keyword regions. The final output consists of three fields: responsibility role, key behavior, and event object. The training corpus of the semantic recognition model comes from a set of manually annotated power grid regulations sentences. Roles and objects are labeled using BIO tags, and the training loss function is multi-label cross-entropy.

[0006] As a preferred embodiment, each semantic unit in the semantic unit set includes a responsible role, an action verb, an operation object, a clause number, and a rule level.

[0007] Preferably, the structural semantic index table uses the operation object as the primary key, with each key indexing a group of textual semantic units and recording information such as its node ID, role, regulation level, and graph adjacency edges, as shown in the following structure: Primary key: Object name; Field 1: Set of hit clauses; Field 2: Responsible role; Field 3: Graph edges and their types, weights, and pointers to adjacent nodes; The structural semantic index table is stored in a key-value database as key-value pairs and is used to construct an inverted index tree.

[0008] Preferably, after the user inputs a natural language question, the natural language question is converted into a semantic triple for graph path starting point retrieval, and then a structured query intent is output through a structure-aware language model.

[0009] More preferably, the structure-aware language model structure includes a base BERT for pre-training and three independent Transformer-MLP substructures, which perform classification predictions for three types of fields: responsibility role, key behavior, and operation object, respectively. Each substructure contains two Transformer layers and one fully connected classification head, with an input dimension of 256 and an output dimension determined by the size of the corresponding category dictionary. The training data comes from the enterprise power grid rules and regulations Q&A database and accident handling resolution texts, and semantic field samples were collected and labeled. The output is a structured query intent, specifically including: the responsibility role field, the behavior target field, and the operation object field. More preferably, after the structure-aware language model extracts the intent field, it locates the matching object entry in the structure semantic index table using the operation object field, obtains the relevant node set, and uses it as the path starting candidate set. For each node in the relevant node set, multi-hop path expansion is performed simultaneously based on the rule hierarchy inheritance edge and the responsibility collaboration chain edge in the dual-structure knowledge graph. The path expansion is prioritized using a scoring function based on edge type and semantic consistency.

[0010] More preferably, the scoring function is used to calculate the score of the selected path, which is obtained by using the rule hierarchy inheritance edge weight, responsibility collaboration chain edge weight, regularization adjustment factor, regularization reward item, the number of times the role or rule source is repeated in the path, and the penalty factor in the dual-structure knowledge graph. The score is calculated within a range where the maximum number of jumps is set to 4. Each path is used as the final set of matching paths, and the nodes hit by the endpoints of the paths are aggregated into a set of matching clauses.

[0011] Preferably, the structured aggregation representation includes semantically encoding each text node in the path, converting each text node into a structured vector, and assigning aggregation weights to each text node in the path.

[0012] More preferably, the aggregation weight is calculated using a role priority set, a rule hierarchy priority set, and a normalization factor. The role priority set includes dispatchers and duty officers, and the rule hierarchy priority set includes groups and countries. The aggregation weight ensures that rule-related nodes in the path, including dispatchers and groups, occupy a higher proportion in the final representation.

[0013] The beneficial technical effects of the present invention are at least as follows: This invention addresses the challenges of complex semantics, multi-dimensional responsibility distribution, and strong hierarchical structure in power grid regulations by proposing a structured semantic retrieval method that integrates knowledge graphs and language models. The method constructs a dual-structure knowledge graph with clauses as nodes and responsibility collaboration and hierarchical inheritance as edges. It uniformly extracts and structurally models elements such as roles, behaviors, objects, and regulation levels from the regulatory text, and uses object fields to build an index table to support rapid path location. Based on this, a structure-aware language model is introduced to parse the intent of user natural language queries, generating structured query triples as entry points for graph path retrieval. The path search process employs a multi-hop diffusion mechanism, integrating responsibility chain integrity, cross-role / cross-level reward items, and structural redundancy penalty factors in path scoring, effectively constructing semantically coherent and responsibility chain-closed-loop clause path sequences. Furthermore, during path modeling, a role priority mechanism and an edge type-guided attention mechanism are introduced to achieve weighted aggregation of clause nodes and path relationships, thereby generating path representations with structural expressive capabilities to support the generation of downstream structured question answering and explanatory retrieval results. Through a continuous processing mechanism of structured parsing, graph modeling, path retrieval, and representation aggregation, this invention constructs a structure-driven retrieval system capable of reconstructing the chain of responsibility, unifying hierarchical regulations, and outputting semantics in a closed-loop path. This provides a reproducible, highly reliable, and strongly interpretable semantic retrieval solution for key business operations such as compliance queries, fault review, and responsibility identification in the power system. Attached Figure Description

[0014] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0015] Figure 1 This is a flowchart of a semantic retrieval method for power grid regulations and rules based on a large-scale language model that integrates knowledge graphs, according to the present invention. Detailed Implementation

[0016] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0017] In one or more embodiments, such as Figure 1 As shown, a semantic retrieval method for power grid regulations using a large-scale language model that integrates knowledge graphs is disclosed. The method includes the following: S1. Obtain the regulations documents through the API call of the company's regulations document storage system. The regulations documents all have metadata fields such as source number, document internal chapter number and regulation level identifier. Then, perform structural recognition processing on the regulations documents and summarize all regulations documents to obtain a regulations document set. Input the regulations document set into the semantic recognition model and output a set of clause semantic units. This step focuses on the information infrastructure construction of power grid regulations and institutional documents. The goal is to extract unstructured regulatory texts into semantically structured units, serving as the basic nodes for subsequent graph modeling and structured semantic retrieval. Power grid regulations and institutional documents typically consist of standard procedures issued by the State Grid Corporation of China, implementation manuals compiled by subsidiaries, operating regulations, and emergency response plans. While highly structured, they lack uniformity in surface document format, often stored in PDF, Word, and HTML formats. Content may include long natural language sentences, nested numbering, and unannotated roles. Therefore, this step primarily addresses the following problem: how to accurately identify the boundaries, responsible roles, core actions, and event objects of each regulation, and structure them into semantic units. This is for use in subsequent map building and semantic computation.

[0018] The input for this step is a collection of regulatory documents. Each document The documents were obtained through the company's internal regulations management platform, and were stored in the following formats: original PDF (approximately 70%), Office Word (approximately 20%), and HTML webpage format (approximately 10%). All documents contained metadata fields such as source number (e.g., "State Grid Dispatch Regulations 2020 Edition"), internal chapter number (e.g., "3.2.1"), and regulation level identifier (e.g., "Group Regulations", "City-level Implementation Rules"). This data was obtained by calling the company's regulations document storage system API to retrieve the original content and metadata.

[0019] First, regarding the document Structural recognition processing is performed. Due to the heterogeneity of document formats, differentiated parsing strategies are required based on document type. For PDF documents, page-level recognition is performed using OCR-based text image parsing tools (such as PaddleOCR or ABBYYAPI) to extract information such as text content, paragraph position, text block coordinates, and number positions for each page. For Word and HTML documents, a combination of style tag parsing and regular expression rules is used to extract heading levels (such as Heading1-3), body text fields, and numbered areas. All documents are ultimately mapped to a unified document paragraph stream, with each paragraph corresponding to a paragraph number, hierarchy position, and text fields, forming the following field set: Article number, obtained by matching the number format using a regular expression (e.g., "3.2.4"); : Rule level, directly marked by document metadata; Chapter number; Paragraph text.

[0020] For example, the group's dispatching regulations document contains the following paragraph: "4.3.2 After receiving a line tripping message, the dispatcher should immediately notify the maintenance team to initiate the power restoration process." This paragraph is interpreted as: , , .

[0021] Next Perform semantic event extraction. Input the extracted semantic events into the semantic recognition model. Its structure is as follows: the input is a text sequence (length ≤ 256 characters), the context dependency information is extracted through a bidirectional LSTM text encoding layer, then a convolutional feature extractor is connected to identify keyword regions, and finally the responsibility role is output. Key behaviors Event object Three types of fields. The model training corpus comes from a set of manually annotated power grid regulations sentences (approximately 3200 sentences). Roles and objects are labeled using BIO tags (such as "B-role" and "I-object"). The training loss function is multi-label cross-entropy.

[0022] The model output is parsed into textual structure units:

[0023] in: It is a structured text; It is a responsible role, by Subject extraction module; These are action verbs; the convolutional layer identifies the core verbs and matches them in the action vocabulary. It is the object of operation, obtained through dependency parsing and cross-validation with a terminology database; and The results of the identification of the fields before the steps.

[0024] In the example above, the system output is: , , .

[0025] To ensure scale consistency in edge weight and similarity calculations in subsequent graphs, the above fields are further normalized. Since the frequency of the same semantic field may vary significantly across different regulations, inconsistencies will affect semantic alignment judgments. Therefore, the following normalization function is introduced:

[0026] in: This represents a score indicating the word frequency or semantic similarity of a particular character, behavior, or object within the corpus. After normalization The representation result within the range; It is a small constant to prevent division by zero, and its value is generally 1. .

[0027] This normalized value is only used to assist in the calculation of edge weights during semantic graph construction and does not affect the original representation of the text field.

[0028] Finally, a set of semantic units of the clauses is obtained. Each of them Include The field indicates that the regulation number is The level of regulations is In one of the system regulations, roles Execute behavior in a specific scenario Acting on objects .

[0029] S2. Based on the set of semantic units of the clauses, construct a dual-structure knowledge graph suitable for the retrieval scenario of power grid rules and regulations, and generate a structure semantic index table simultaneously. By mapping each semantic unit of the clauses of the semantic unit set to a graph node, the edge set of the dual-structure knowledge graph includes rule hierarchy inheritance edges and responsibility collaboration chain edges. The goal of this step is to build upon the set of semantic units of the text obtained in step S1. Construct a dual-structure knowledge graph suitable for power grid regulations retrieval scenarios. Simultaneously generate a structural semantic index table. This graph system describes the "hierarchical inheritance relationship" and "job responsibility collaboration link" between regulations and provisions, which is used to support subsequent structured semantic retrieval and path reasoning tasks.

[0030] In practice, power grid regulations and systems exhibit the following typical characteristics: First, regulations exist at multiple levels, including national standards, group regulations, and unit procedures, with potentially duplicated clause numbers but semantically guiding and subordinate relationships. Second, the same equipment event or accident type (such as "tripping," "fire," or "voltage anomaly") is typically handled by multiple roles in a "sequential + division of labor" manner, resulting in fragmented and semantically coupled regulations for each role. Third, different regulatory levels exhibit semantic style differences for the same object (e.g., "should be handled immediately" versus "reported as soon as possible"). These differences must be explicitly identified and standardized by the graph to avoid semantic retrieval confusion. Therefore, this graph modeling scheme emphasizes structural rationality, semantic integration, and application orientation oriented towards responsibility chains.

[0031] This step first constructs the graph node set. Each Mapped to graph nodes The node attributes retain the original 5-tuple fields. The unique identifier for each node uses... The system is assembled and combined to ensure that clauses with the same number but different regulatory sources remain structurally distinct.

[0032] Next, construct the edge set. It contains two subsets: rule hierarchy inheritance edge and the chain of responsibilities .in, Modeling the "semantic inheritance relationship" between rules and regulations. Model the "relationship of handling the same event by different roles" among multiple roles.

[0033] The method for constructing the rule hierarchy is as follows: In order to identify semantic inheritance clauses at different regulatory levels (such as group and city), this invention first identifies each Constructing joint semantic vectors by its behavior and object The fields are obtained by concatenating the encodings from a fine-tuned BERT model. The BERT model has been fine-tuned in the domain based on the State Grid regulations data. Its structure is a 12-layer Transformer, with an input of "action + object" sequence and an output vector dimension of 128.

[0034] The semantic inheritance weights are defined as follows:

[0035] in: For the clauses and Semantic similarity weights; and The embedding vector constructed in step one; Represents the cosine similarity of vectors; This is an indicator function for whether the regulatory levels are consistent. Penalty coefficient (recommended value) This is used to suppress the phenomenon of excessively high semantic similarity in regulations at the same level.

[0036] This innovative formula introduces a special negative term for "cross-level clause matching," ensuring that the graph prioritizes establishing inheritance relationships between cross-level clauses, rather than simply identifying similar text as synonyms. Only when... (Actual measurement) )and Only then should a hierarchical inheritance boundary be established. .

[0037] The method for constructing the chain of responsibility edges is as follows: The chain of responsibility structure aims to depict the logical connections between the clauses involved in different roles within an event handling chain. Its implementation involves three steps: ① By operation object Create an inverted index This is used to find all clause pairs that involve the same object; ② Use a predefined behavior sequence list (e.g., "Discover" < "Report" < "Process" < "Record") Determine the action of the clause. and Does the timing match? ③If satisfied and If the timing is met, then a responsibility collaboration edge is generated.

[0038] To achieve more refined edge weight differentiation, the following responsibility chain edge weight function is proposed:

[0039] in: For word vector similarity between actions (trained on power grid regulations data using Word2Vec); For semantic similarity between objects, the Jaccard similarity metric is used. Instructions for different job positions; To indicate whether the actions meet the sequence requirements, set the value to 1 or 0; These are weighted parameters with initial values ​​of (0.3, 0.3, 0.2, 0.2), which can be adjusted according to the task precision.

[0040] This formula innovatively integrates four dimensions: word vector similarity, object consistency judgment, job similarity and difference judgment, and action sequence structure, enabling the responsibility chain graph to truly recreate the rule chain in a cross-role event processing flow and support edge weight sorting.

[0041] The structural semantic index table is constructed as follows: To avoid performance bottlenecks caused by graph traversal retrieval, this step simultaneously establishes a structural semantic index table. By operating on the object Primary key, each key indexes a group It records information such as node ID, role, rule level, and graph adjacency edges, with the following structure: Primary key: object name Field 1: Set of Hit Clauses Field 2: Corresponding role Field 3: Graph Edges Its type, weight, and neighboring node pointers.

[0042] This table is stored in a key-value database as key-value pairs. An inverted index tree can be built as needed, supporting fast local path queries in the graph triggered by event objects or roles in the next step.

[0043] This step yields a structural knowledge graph. , where the set of nodes From edge set Includes hierarchical edges Collaboration with Responsibilities ;Structure semantic index table A structured index table organized with object fields as the primary key, storing graph nodes and their structural connection information.

[0044] S3. Map the natural language question input by the user into a structured semantic representation, and based on the constructed dual-structure knowledge graph of power grid regulations and the structured semantic index table, complete the multi-hop retrieval of the set of articles related to the semantics of the user question and its path structure; and extract the jump links in the dual-structure knowledge graph, the jump links include the path between the starting article node and another article node, by setting the maximum number of jumps to 4, select the top k paths with the highest path scores, and summarize to obtain the final matching path set; This step, based on the output of the first two steps, maps the user's input natural language question into a structured semantic representation, and then uses the already constructed dual-structure knowledge graph of power grid regulations. With structural semantic index table Complete the set of clauses related to the semantics of user questions. Its path structure Multi-hop retrieval. Its main goal is to solve two core challenges in semantic queries of power grid regulations: First, user expressions often contain issues such as missing job titles, ambiguous events, or non-standard behavioral verbs, directly affecting the accuracy of retrieval; second, regulations are presented in a multi-level hierarchy with a wide distribution of roles and dense redundancy in the graph, making it impossible to form accurate clause matching and responsibility path closure without structural guidance. In real-world scenarios, queries about power grid regulations such as "accident handling," "violation accountability," and "operation scheduling" are often not presented in standardized terms. For example, expressions like "who is responsible for reporting a fire" or "who should be notified when a cable trips" are neither standardized nor do they explicitly indicate the object category and job title. The structural semantic guidance mechanism designed in this step is specifically designed for such inputs. Through a structure-aware language model and path selection regularization design, it maximizes the restoration of the actual clause mapping path of the user query in the regulation graph structure, realizing the structured projection of semantic intent in the regulation graph.

[0045] spectral structure ,in For each clause node, Include Five-element structure; These are the rule hierarchy inheritance edges and responsibility chain edges, respectively. Index Table To operate on objects Using the primary key, it organizes structured inverted information containing relevant article IDs, role sets, and graph structure relationships.

[0046] In addition, user input issues It is a Chinese natural language statement, such as "Who should report a line trip?" or "How should a cable fire be handled?". This input is not yet structured and needs to be parsed into semantic metadata for graph path initiation.

[0047] The first phase of this step involves addressing natural language issues. Transform into structural semantic triples This is used for retrieving the starting point of a path in a graph. For this purpose, a structure-aware language model is used. The model structure consists of BERT (pre-trained base) plus three independent Transformer-MLP substructures, which perform classification predictions for three categories of fields: responsibility role, key behavior, and operation object. Each substructure contains two Transformer layers and one fully connected classification head layer. The input dimension is 256, and the output dimension depends on the size of the corresponding category dictionary.

[0048] The training data comes from the enterprise power grid regulations Q&A database and accident handling resolution texts, with a total of 3700 semantic field samples collected and labeled. The output is a structured query intent.

[0049] in: The field representing the responsibility role is identified from the syntactic subject or role noun, such as "who", "duty officer", "dispatcher"; The target field for the behavior is extracted from predicate phrases, such as "report", "notify", and "isolate". This indicates the field being operated on, using semantic similarity in the index table. Obtained by matching the primary key; This represents a structured query intent triple.

[0050] After the model extracts the intent field, that is... In the index table Locate the matching object entry in the middle and obtain its related node set. , as the candidate set for the starting point of the path.

[0051] Subsequently, in response to Each node in the graph is based on the graph. Chain of Responsibility With hierarchical chains Simultaneously, multi-hop path expansion is performed. Path expansion employs a scoring function based on edge type and semantic consistency for priority ranking. Specifically addressing the fragmentation problem of responsibility links in power grid regulations, the following path scoring function is designed:

[0052] in: For path The score; The edge weights (responsibility edges or inherited edges) obtained from step two. This is a new regularization term introduced in this step, representing the edge. Whether it crosses levels or positions, if it does, an additional reward value (such as 0.1) will be added to improve the diversity of the path; Regularization factor (recommended value) ); This indicates the number of times a role or regulation source is repeated in the path, serving as a path redundancy penalty item. As a penalty factor, a recommended value is... .

[0053] This scoring function is designed for the scenario of retrieving power grid regulations and rules, specifically addressing the problem of path selection tending towards "local fragment repetition" and "excessive role consistency" when responsibilities are distributed across multiple positions and regulations overlap. This is achieved by incorporating structural reward regularization. It encourages the priority activation of regulatory linkage chains that cross roles and regulations, making the path structure more aligned with accident accountability or handling procedures.

[0054] Path expansion employs a combination of breadth-first search (BFS) and heuristic scoring sorting, with a set maximum number of hops. Before selecting a path within the range of scores 10 paths, as the final set of matching paths Meanwhile, the set of node clauses hit by the path endpoint is denoted as .

[0055] For example, in the problem "Who is responsible for handling a cable fire?", model recognition , , Missing. The system starts with the object "cable fire" in the graph, prioritizes the path with the role of "dispatcher" or "duty officer", and jumps down to the "maintenance team" or "emergency repair team" node through the responsibility chain, forming a complete chain of "dispatch discovery → maintenance handling" that conforms to the accident scenario.

[0056] This step yields the set of clauses. : This is the set of endpoint clauses for the hit path, each clause being structured; Path set : is a set of structural paths, each path is Additional points Related to path explanation fields (such as annotation information for role jumps, rule level crossings, etc.).

[0057] The path set This represents the system's representation within the graph structure, indicating its relevance to the user's question. A semantically related set of "inter-article jump links". Each path Indicates starting from a clause node Initially, through a series of responsibilities or hierarchical inheritance relationships... Finally, it jumps to a clause node. This path reflects the process of linking clauses between multiple roles or levels in regulations and systems, that is, a chain of responsibilities such as "who is responsible → who handles → who reports". Each side of the path... They all represent a structural jump or logical connection. All paths together constitute the system's structured semantic understanding process of user questions.

[0058] S4. Based on the final matched path set and the hit text set, perform structured semantic modeling to generate a structured aggregate representation for each path, as well as all node and edge codes. Combine the node and edge codes into a path structure representation. Match all path structure representations with the structured semantics and sort them according to cosine similarity. The path with the highest cosine similarity score is selected. The path was used as the search result; The task of this step is to obtain the path set in step S3. and the collection of hit clauses Based on this, structured semantic modeling is performed to generate a structural aggregate representation for each path. This supports subsequent structured retrieval and interpretation generation. The core purpose of path modeling is to uniformly represent clause information from different sources, roles, and regulatory levels within the path, enabling the subsequent system to consistently model and parse the "responsibility chain jointly defined by multiple regulations." In power grid regulations, the same responsibility chain is often composed of multiple clauses across roles and regulatory levels. For example, a handling path involving a "cable fire" might consist of two clauses: "The dispatcher should notify the maintenance team" and "Maintenance personnel are responsible for handling the fire," from the group and unit regulations respectively. If these paths cannot be structurally aggregated, subsequent semantic retrieval and responsibility chain reconstruction will be broken. This step is the key modeling step in solving this problem.

[0059] To construct a structured representation from the set of paths, this invention first requires semantic encoding of each text node in the path. Each node It will be converted into a structured vector To maintain semantic consistency, this invention uses the text encoder model trained in step one. The model structure is a BiLSTM (2 layers, with a hidden layer size of 128, and the input is a concatenated sequence of roles, behaviors, and objects). The model has been pre-trained on a set of power grid regulations and rules (approximately 8,000 articles), giving it a strong ability to express industry terminology.

[0060] For example, the clauses The structured information for "The dispatcher should immediately notify the maintenance team upon receiving a fire report" is as follows: , , The concatenated sequence was input as "Dispatcher reports fire". This sequence was then input. The output is a context semantic vector of length 128. .

[0061] Next, this invention assigns an aggregation weight to each node in the path. The weight design considers two factors: first, whether the node belongs to a core role in the path; and second, whether the node comes from a high-priority regulatory level (such as group regulations). This invention constructs a role priority set. and rule hierarchy priority set The weights are calculated as follows:

[0062] in , , It is a normalization factor. This design makes the "dispatcher-group regulation" node in the path occupy a higher proportion in the final representation, which is more in line with the structure of the power grid event responsibility dominance chain.

[0063] Edge modeling employs the same encoding strategy as edge weight calculation in step two. This invention inherits the edge values ​​for each path from step three. weight In this step, the present invention calculates the aggregate weight for each edge. The calculation logic is as follows:

[0064] in These are the edge weights already present in step two. It's about adjusting parameters. It is a cross-level / cross-role reward item: if the side A reward (set to 0.1 and 0.15) is given for edges connecting different roles or levels of regulations. This allows parts of the path with frequent structural jumps and large responsibility spans to receive higher weight in the aggregated representation.

[0065] Obtain all node codes With edge encoding Subsequently, the present invention combines them into a path structure representation. The calculation formula is as follows:

[0066] After aggregation, each path Mapped to a vector This represents the comprehensive structural semantics. These vectors will then be input into the structured semantic retrieval module as the basis for sorting and filtering.

[0067] To generate the final search results, the system represents all paths. With structured query intent Matching is performed, and the results are sorted based on the cosine similarity between the path aggregation vector and the query intent. The top-scoring vectors are ranked first. Each path is returned as a search result. The article number, role chain, and path structure information attached to each search result will be output in a formatted manner to support power business personnel in a structured understanding of the scope of responsibility and source of the article.

[0068] For path This represents the path "Dispatcher Notification → Operations and Maintenance Processing → Unit Reporting," which includes two nodes: "Group Regulations Dispatcher Notification" and "Unit Regulations Operations and Maintenance Processing," one responsibility edge, and one hierarchical edge, ultimately forming a path aggregation vector. By combining role priority, hierarchical rewards, and contextual text representation, the system ultimately selects this path as the high-confidence retrieval result and outputs structural fields such as the clause number (e.g., "Group 3.2.1", "Unit 7.1.4"), role responsibility chain ("Dispatcher → Maintenance Personnel"), and path parsing description ("Cross-level + Cross-role double-hop path"), for users to quickly understand and apply.

[0069] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A semantic retrieval method for power grid regulations and rules based on a large-scale language model integrating knowledge graphs, characterized in that: include: The company's regulations document storage system API is used to obtain regulations documents. Each regulation document contains metadata fields such as source number, internal chapter number, and regulation level identifier. The regulations documents are then processed for structural recognition, and all regulations documents are aggregated to obtain a regulations document set. The regulations document set is then input into a semantic recognition model to output a set of clause semantic units. Based on the set of semantic units of the clauses, a dual-structure knowledge graph suitable for the retrieval scenario of power grid rules and regulations is constructed, and a structural semantic index table is generated simultaneously. By mapping each semantic unit of the clauses in the set of semantic units of the clauses to a graph node, the edge set of the dual-structure knowledge graph includes rule hierarchy inheritance edges and responsibility collaboration chain edges. The system maps user-input natural language questions into structured semantic representations and, based on the constructed dual-structure knowledge graph of power grid regulations and the structured semantic index table, performs multi-hop retrieval of the set of articles and their path structures related to the semantics of the user's question. Jump links are extracted from the dual-structure knowledge graph, including paths from the starting article node to another article node. By setting a maximum number of hops of 4, the top k paths with the highest path scores are selected, and the final matching path set is obtained. Based on the final matched path set and the set of hit clauses, structured semantic modeling is performed to generate a structured aggregate representation, all node codes, and edge codes for each path. The node codes and edge codes are combined into a path structure representation. All path structure representations are matched against the structured semantics, and sorted according to cosine similarity. The path with the highest cosine similarity score is selected. The path was used as the search result.

2. The semantic retrieval method for power grid regulations and rules based on a large-scale language model incorporating knowledge graphs as described in claim 1, characterized in that, The structure of the semantic recognition model is as follows: The input is a text sequence with a length of less than or equal to 256 characters. Contextual dependency information is extracted through a bidirectional LSTM text encoding layer, followed by a convolutional feature extractor to identify keyword regions. The final output consists of three fields: responsibility role, key behavior, and event object. The training corpus of the semantic recognition model comes from a set of manually annotated power grid regulations sentences. Roles and objects are labeled using BIO tags, and the training loss function is multi-label cross-entropy.

3. The semantic retrieval method for power grid regulations and rules based on a large-scale language model incorporating knowledge graphs as described in claim 1, characterized in that... Each semantic unit in the set of clause semantic units includes the responsible role, action verb, operation object, clause number, and rule level.

4. The semantic retrieval method for power grid regulations and rules based on a large-scale language model integrating knowledge graphs as described in claim 1, characterized in that, The structural semantic index table uses the operation object as the primary key. Each key indexes a group of textual semantic units and records information such as its node ID, role, regulation level, and graph adjacency edges. The structure is as follows: Primary key: Object name; Field 1: Set of hit clauses; Field 2: Responsible role; Field 3: Graph edges and their types, weights, and pointers to adjacent nodes; The structural semantic index table is stored in a key-value database as key-value pairs and is used to construct an inverted index tree.

5. The semantic retrieval method for power grid regulations and rules based on a large-scale language model integrating knowledge graphs according to claim 1, characterized in that, After the user inputs a natural language question, the natural language question is transformed into a semantic triple for graph path starting point retrieval, and then a structured query intent is output through a structure-aware language model.

6. The semantic retrieval method for power grid regulations and rules based on a large-scale language model incorporating knowledge graphs as described in claim 5, characterized in that, The structure-aware language model structure includes a base BERT for pre-training and three independent Transformer-MLP substructures, which perform classification predictions for three types of fields: responsibility role, key behavior, and operation object, respectively. Each substructure contains two Transformer layers and one fully connected classification head layer. The input dimension is 256, and the output dimension is determined according to the size of the corresponding category dictionary. The training data comes from the enterprise power grid rules and regulations Q&A database and accident handling resolution texts, and semantic field samples were collected and labeled. The output is a structured query intent, which includes: the responsibility role field, the behavior target field, and the operation object field.

7. The semantic retrieval method for power grid regulations and rules based on a large-scale language model integrating knowledge graphs as described in claim 5, characterized in that, After the structure-aware language model extracts the intent field, it locates the matching object entry in the structure semantic index table using the operation object field, obtains the relevant node set, and uses it as the path starting candidate set. For each node in the relevant node set, multi-hop path expansion is performed simultaneously based on the rule hierarchy inheritance edge and the responsibility collaboration chain edge in the dual-structure knowledge graph. The path expansion is prioritized using a scoring function based on edge type and semantic consistency.

8. The semantic retrieval method for power grid regulations based on a large-scale language model integrating knowledge graphs according to claim 7, characterized in that, The scoring function is used to calculate the score of the selected path. It is calculated using the rule hierarchy inheritance edge weights, responsibility collaboration chain edge weights, regularization adjustment factors, regularization reward items, the number of times the role or rule source is repeated in the path, and penalty factors in the dual-structure knowledge graph. The score is calculated within the range of a maximum number of jumps of 4. Each path is used as the final set of matching paths, and the nodes hit by the endpoints of the paths are aggregated into a set of matching clauses.

9. The semantic retrieval method for power grid regulations and rules based on a large-scale language model integrating knowledge graphs according to claim 1, characterized in that, The structured aggregation representation includes semantically encoding each text node in the path, converting each text node into a structured vector, and assigning aggregation weights to each text node in the path.

10. The semantic retrieval method for power grid regulations and rules based on a large-scale language model integrating knowledge graphs according to claim 9, characterized in that, The aggregation weight is calculated using a role priority set, a rule hierarchy priority set, and a normalization factor. The role priority set includes dispatchers and duty officers, and the rule hierarchy priority set includes groups and countries. The aggregation weight ensures that rule-related nodes in the path, including dispatchers and groups, have a higher weight in the final representation.

Citation Information

Patent Citations

  • Regulation retrieval method and device, computer equipment and storage medium

    CN117743511A

  • Modularized knowledge graph and retrieval enhanced large model fusion interaction method and system oriented to financial branch mechanism

    CN120448510A

Cited By

  • Inspection knowledge inference analysis method based on combination of large language model and knowledge graph

    CN121480667A