Safety specification intelligent question answering method and system based on double-layer knowledge graph and retrieval enhanced generation large model
Patent Information
- Application Number
- CN202610810569.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-18
AI Technical Summary
若仅采用向量检索,可能返回语义相近但适用条件不同的条款;若仅采用知识图谱,又难以处理自然语言问题中的省略表达、同义表达和复杂上下文
本发明通过构建规范知识图谱与场景知识图谱并建立跨层映射,将安全规范条款解析为可计算的对象、条件、行为、指标和后果单元,同时将企业现场检查记录、风险辨识和整改信息组织为可追踪的场景风险单元,使问答系统在生成答案前能够完成条款适用性判断和现场事实匹配,有助于减少条款误配、场景脱离和大模型幻觉问题;通过协同检索中的图谱路径检索与语义向量检索相结合,并引入跨层匹配分数和冲突校验机制,能够提高所引用条款与用户问题之间的相关性和有效性,避免因规范层级冲突或条件缺失而输出不可靠结论;通过证据压缩和答案可信度评估,可使系统在确定性不足时输出补充核验建议而非强制结论,从而提升安全规范问答结果的可追溯性和现场适用性。
Smart Images

Figure CN122594438A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of safety production informatization, knowledge graphs, and artificial intelligence question answering technologies, and in particular to a method and system for intelligent question answering of safety regulations based on a two-layer knowledge graph and a retrieval-enhanced generation model. Background Technology
[0002] Safety production management involves multiple sources of knowledge, including laws, administrative regulations, departmental rules, local regulations, national standards, industry standards, local standards, normative documents, and internal enterprise systems. This knowledge typically exhibits characteristics such as multiple levels, complex applicable conditions, inconsistent terminology, frequent revisions, and significant cross-industry differences. Taking on-site safety management in industrial and commercial enterprises as an example, the same issue may simultaneously involve multiple dimensions such as risk classification and control, hazard investigation and management, inherent equipment safety, warning signs, power outage labeling, education and training, fire protection facilities, electrical safety, and special equipment management. Traditional keyword search systems usually only return clauses or documents containing the same words, failing to determine whether a clause applies to a specific industry, equipment, location, or operational state, easily leading to clause mismatches.
[0003] Existing safety specification question-answering methods based on large models can understand natural language questions, but without a controlled knowledge structure and evidence constraints, they are prone to problems such as incomplete supporting clauses, missing scenario conditions, confusion of specification levels, and generalization of rectification suggestions. For example, when a user asks "Does the conveyor need to be tagged and powered off during cleaning?", the system not only needs to understand the semantics of "conveyor," "cleaning operation," and "tagged and powered off," but also needs to consider on-site conditions such as equipment length, control panel settings, emergency stop devices, protective covers, power distribution cabinet markings, "do not close" signs, and personnel working status to form an actionable safety management answer. If only vector retrieval is used, it may return clauses with similar semantics but different applicable conditions; if only knowledge graphs are used, it is difficult to handle ellipsis, synonyms, and complex context in natural language questions.
[0004] Therefore, existing technologies urgently need a smart question-and-answer method for safety regulations that can perform two-layer modeling of regulatory clauses and enterprise-level scenarios. This method would enable the system to simultaneously possess the ability to trace the source of regulatory clauses, determine the applicability of scenarios, and express generative ideas when answering questions. This method should not merely be a simple concatenation of knowledge graphs and retrieval-enhanced generative models. Instead, it should utilize cross-layer obligation mapping, applicability condition verification, risk trigger propagation, and evidence compression mechanisms to ensure that regulatory knowledge and on-site knowledge are closed within the same reasoning chain, thereby improving the accuracy and enforceability of safety regulation question-and-answer results. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, this invention provides a secure and intelligent question-answering method and system based on a two-layer knowledge graph and retrieval enhancement generation model, in order to solve the problems existing in the background art.
[0006] This invention provides the following technical solution: Firstly, this application provides a secure and standardized intelligent question answering method based on a two-layer knowledge graph and retrieval enhancement generation model, comprising the following steps: S100: Obtain safety specification texts, basic enterprise information, work scenario information, equipment and facility information, risk identification information, on-site inspection records, and rectification records.
[0007] S200 performs clause-level parsing of security specification texts to construct a specification knowledge graph.
[0008] S300 extracts scenario-based information from enterprise basic information, operation scenario information, equipment and facility information, risk identification information, on-site inspection records and rectification records to construct a scenario knowledge graph.
[0009] S400 establishes a cross-layer mapping relationship between the standard knowledge graph and the scenario knowledge graph to obtain a two-layer knowledge graph.
[0010] S500 generates a security question-and-answer intent frame based on the user's question.
[0011] S600 performs collaborative retrieval on a two-layer knowledge graph and vector index based on security question-answering intent frames to obtain a set of candidate evidence.
[0012] S700 performs applicability checks, conflict checks, and evidence compression on the candidate evidence set to obtain enhanced retrieval context.
[0013] S800 inputs the enhanced retrieval context into the security specification question-answering model to generate security specification question-answering results.
[0014] In the aforementioned implementation process, this application does not simply vectorize the normative text and directly input it into a large model. Instead, it first parses the normative clauses into computable obligation units, then parses the on-site problems of the enterprise into computable risk scenario units, and explicitly represents "whether a certain type of on-site fact triggers a certain normative obligation" through cross-layer mapping relationships. As a result, the system can complete the clause applicability judgment before generating the answer, and then perform evidence screening and answer generation, reducing the illusion risk when directly generating from a large model.
[0015] Furthermore, the two-layer knowledge graph is represented as follows: in, Represents a two-layer knowledge graph; This represents a set of nodes in a normative knowledge graph, including nodes related to laws, standards, clauses, obligations, applicable objects, and control measures. This represents a set of nodes in a scenario knowledge graph, including nodes such as enterprise, equipment, location, operation, hidden danger, evidence, and rectification. Represents the set of relations within a canonical knowledge graph; This represents the set of relationships within a scene knowledge graph; This represents the set of cross-layer mapping relationships between the canonical knowledge graph and the scenario knowledge graph. A set of attributes representing nodes and relationships; This represents a set of time attributes used to record the effective date, repeal date, inspection date, rectification date, and review date of the standard.
[0016] Furthermore, the normative obligation unit is represented as follows: in, Indicates the first Each standardized obligation unit; The term "applicable objects" refers to equipment, facilities, locations, personnel, positions, industries, or work activities. The applicable conditions are indicated, including equipment parameters, operating status, risk level, geographical scope, industry scope, and time status. The standard operating procedure requirements include the following: they must be set up, publicized, trained, inspected, marked and powered off, and must not be occupied, obstructed, or operated in violation of regulations. The indicators to be determined include distance, height, quantity, cycle, learning hours, power-on status, identification status, and protection status. This indicates the potential risks, legal consequences, or rectification directions that may result from violating the aforementioned normative obligations.
[0017] Furthermore, the scenario risk unit is represented as: in, Indicates the first Each scenario risk unit; Indicates the entity or location; Indicates equipment, facilities, or work areas; Indicates on-site behavior or operational status; Indicates an unsafe condition or unsafe behavior; The evidence data includes text inspection records, image recognition results, video screenshots, rectification photos, sensor data, or manual review records.
[0018] Furthermore, the matching score for cross-layer mapping relationships is expressed as: in, Represents the scenario risk unit With normative obligation unit Cross-layer matching scores; This indicates the semantic similarity between scenario risk units and normative obligation units; This represents the set of keywords extracted from the scenario risk unit; This represents the set of keywords extracted from the normative obligation unit; Indicates the degree of overlap between two keyword sets; This indicates the degree to which applicable conditions are met, and is used to determine whether conditions such as industry, equipment, operating status, and risk level are consistent. The normative validity score indicates whether a clause is currently valid, whether it is mandatory, and its normative level. , , , These are weighting coefficients, all of which are non-negative, used to adjust the impact of semantic similarity, keyword overlap, condition satisfaction, and norm validity on cross-layer matching.
[0019] Furthermore, the collaborative retrieval score is expressed as: in, Indicates user issues With candidate evidence Collaborative retrieval scores between them; This represents the vector semantic similarity between the user's question and the candidate evidence; This indicates the path relevance of the user's question to candidate evidence in the two-layer knowledge graph; This represents the cross-layer matching score between the scenario risk unit and the normative obligation unit corresponding to the candidate evidence; This represents the mismatch penalty value for candidate evidence, which increases when the evidence is geographically inapplicable, outdated, has inconsistent objects, or lacks certain conditions. These are non-negative weighting coefficients.
[0020] Furthermore, the credibility of the answer is expressed as: in, Indicate the answer Credibility; This indicates the number of pieces of evidence used to generate the answer; Indicates the first The weight of each piece of evidence; Indicates the first The verification value of each piece of evidence is set as follows: when the evidence meets the applicable conditions and there is no conflict, a higher value is taken; when the evidence lacks key conditions, a lower value is taken. This represents the conflict penalty coefficient, which increases when there are conflicts in normative hierarchy, version, or scenario facts among different pieces of evidence. This formula controls the strength of the answer output; when... When the value is below a preset threshold, the system outputs supplementary verification suggestions instead of a definitive conclusion.
[0021] The technical effects and advantages of this invention are as follows: This invention constructs a normative knowledge graph and a scenario knowledge graph and establishes cross-layer mapping to parse safety normative clauses into computable objects, conditions, behaviors, indicators, and consequences. Simultaneously, it organizes enterprise on-site inspection records, risk identification, and rectification information into traceable scenario risk units. This allows the question-and-answer system to complete clause applicability judgment and on-site fact matching before generating answers, helping to reduce clause mismatch, scenario disconnect, and large model illusion problems. By combining graph path retrieval and semantic vector retrieval in collaborative retrieval, and introducing cross-layer matching scores and conflict verification mechanisms, the relevance and effectiveness between cited clauses and user questions can be improved, avoiding unreliable conclusions due to normative level conflicts or missing conditions. Through evidence compression and answer credibility assessment, the system can output supplementary verification suggestions rather than mandatory conclusions when certainty is insufficient, thereby improving the traceability and on-site applicability of safety normative question-and-answer results. Attached Figure Description
[0022] Figure 1 A flowchart illustrating a security-compliant intelligent question-answering method based on a two-layer knowledge graph and retrieval-enhanced generation model, provided for embodiments of this application; Figure 2 A schematic diagram of the two-layer structure of the canonical knowledge graph and the scene knowledge graph provided in the embodiments of this application; Figure 3 This is a schematic diagram of the cross-layer mapping relationship construction process provided in the embodiments of this application; Figure 4 This is a schematic diagram of the collaborative retrieval and retrieval enhancement generation process provided in the embodiments of this application. Detailed Implementation
[0023] The technical solution of this application will be described below with reference to embodiments thereof. It should be understood that the described embodiments are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0024] In this embodiment, the safety specification text may include laws, administrative regulations, departmental rules, local regulations, normative documents, national standards, industry standards, local standards, and corporate policies. Basic corporate information may include the company name, industry category, address, factory area, job positions, and responsibility list. Work scenario information may include scenarios such as maintenance, cleaning, dredging, hot work, working at heights, confined spaces, conveying, briquetting, textile circular knitting machine operation, forklift operation, and power distribution operation. Equipment and facility information may include belt conveyors, briquetting machines, sorting machines, power distribution cabinets, air compressors, alarm devices, fire-fighting facilities, forklifts, combustible gas alarm devices, and safety interlock devices. On-site inspection records may include inspection time, problem description, basis for violation, on-site photos, and rectification requirements.
[0025] The first step is data acquisition and standardization (S100). Specifically, the system retrieves safety data from the standards library, company ledgers, risk reports, hazard investigation records, enforcement inspection records, rectification closure records, and on-site photos. For standard texts, the system hierarchically segments them according to document, chapter, clause, item, and appendix. For on-site inspection records, the system performs field processing according to company, equipment, location, problem, basis, and rectification direction. For image or video evidence, the system extracts visual objects such as equipment names, warning signs, emergency stop buttons, protective covers, power distribution tags, passageway occupancy, and fire-fighting facility obstructions, and binds the recognition results to the text records.
[0026] During data standardization, the system establishes a unified terminology list. For example, "power off with tag," "lock with tag," "do not close switch reminder," and "power off for maintenance" are uniformly mapped to energy isolation terms; "emergency stop button," "emergency stop rope," and "emergency stop control" are uniformly mapped to emergency stop terms; and "four-color risk map," "risk disclosure board," and "signature or higher risk warning sign" are uniformly mapped to risk disclosure terms. This terminology standardization avoids the same safety requirement being split into different knowledge nodes due to different expressions.
[0027] Based on this, a normative knowledge graph is constructed (S200). Clause-level parsing is performed on the security specification text to construct the normative knowledge graph. The normative knowledge graph is used to answer questions such as "What is the basis?", "What objects do the clauses apply to?", "What does the clause require?", and "How should violations be rectified?".
[0028] Specifically, the system first identifies the metadata of the standard text, including the standard name, standard number, issuing entity, implementation date, applicable industry, applicable region, mandatory status, and repeal status. Then, the system extracts obligations from the clause text, identifying standard verbs such as "shall," "must," "shall not," "prohibited," "appropriate," and "may," and distinguishes between mandatory obligations, prohibited obligations, recommended requirements, and explanatory requirements based on these verbs. For clauses containing numerical conditions, the system extracts verifiable indicators, such as "distance must not exceed a preset value," "height must not be lower than a preset value," "training time must not be less than a preset number of hours," and "distance from any point to the emergency stop device meets preset requirements."
[0029] The internal relationships within a normative knowledge graph include: the inclusion relationship between regulations and clauses; the binding relationship between clauses and applicable objects; the requirement relationship between clauses and control measures; the preventive relationship between clauses and risk types; the consequence relationship between clauses and legal liabilities; and the substitution relationship between old and new versions of clauses. Through these relationships, the normative knowledge graph can start with a problem object and find applicable clauses and corresponding control measures along the graph path.
[0030] Simultaneously, a scenario knowledge graph (S300) is constructed. Scenario-based extraction is performed from basic enterprise information, operational scenario information, equipment and facility information, risk identification information, on-site inspection records, and rectification records to construct the scenario knowledge graph. The scenario knowledge graph is used to answer questions such as "What happened on-site?", "Where is the risk?", "What is the evidence?", and "Has it been rectified?".
[0031] Specifically, the system extracts enterprise nodes, equipment nodes, location nodes, operation nodes, hazard nodes, basis nodes, and rectification nodes from on-site inspection records. For example, for the description "maintenance work was not marked and power was cut off," the system extracts the operation node "maintenance work," the hazard node "no power cut-off," the risk type node "accidental start or mechanical injury," and the control measure node "power off and mark." For the description "the distribution cabinet did not have the corresponding control equipment name affixed," the system extracts the equipment node "distribution cabinet," the defect node "missing the corresponding control equipment name label," and the control measure node "set up Chinese labels and indicate the equipment correspondence." For the description "fire hydrants were obstructed," the system extracts the fire protection facility node "fire hydrant," the hazard node "obstructed," and the control measure node "remove the obstruction and make it accessible."
[0032] The internal relationships of the scenario knowledge graph include: the enterprise has equipment, the equipment is located in a region, the equipment includes parts, the parts have potential hazards, the hazards are associated with operations, the hazards have evidence, the hazards correspond to corrective measures, and the corrective measures have responsible positions and completion status. Through these relationships, the system can organize scattered inspection issues into a traceable risk chain.
[0033] After the canonical knowledge graph and the scenario knowledge graph are constructed, a two-layer cross-layer mapping (S400) is performed. A cross-layer mapping relationship is established between the canonical knowledge graph and the scenario knowledge graph. This cross-layer mapping is a key step that distinguishes this application from ordinary knowledge base question answering. This step does not mechanically merge the two knowledge graphs, but rather determines whether a scenario risk unit triggers a canonical obligation unit.
[0034] In one implementation, the system addresses each scenario risk unit. Calculate its relationship with multiple normative obligation units Matching score When the matching score is higher than the first threshold, an "applicable" relationship is established; when the unsafe state in the scenario risk unit is contrary to the behavioral requirements in the normative obligation unit, a "violation" relationship is established; when the rectification record meets the behavioral and indicator requirements in the normative obligation unit, a "satisfaction" relationship is established; when on-site photos, inspection records, or sensor data can prove the existence of a certain hidden danger, an "evidence support" relationship is established.
[0035] For example, if the scenario risk unit is described as "the belt conveyor does not have an emergency stop button," the object slot of the regulatory obligation unit is "belt conveyor," the condition slot is "the conveyor length and operational accessibility meet preset conditions," and the behavior slot is "install an emergency stop button or emergency stop rope," then the system establishes a violation relationship based on the equipment object, operational risk, and control measures. If the scenario risk unit is described as "the circuit breaker of the distribution cabinet does not have a Chinese label and cannot identify the corresponding field equipment," and the behavior slot of the regulatory obligation unit is "panels and circuits should be labeled with the controlled object," then a violation relationship is established, and the risk consequence of "convenience of tagging and locking" is also associated.
[0036] Furthermore, the system establishes rules for risk trigger propagation. When a potential hazard node is associated with multiple risk consequences, the system propagates along the cross-layer mapping relationship to the regulatory layer to find more complete regulatory basis. For example, "unmarked power outage" not only triggers maintenance operation safety requirements but may also trigger multiple regulatory obligations such as equipment misstart, energy isolation, operating procedures, training and education, and on-site warning signs. The system sorts these bases according to the length of the graph path, the strength of evidence, and the priority of obligations, rather than simply returning the clause with the closest semantics.
[0037] Based on the aforementioned two-layer knowledge graph, the system generates a security question-and-answer intent frame (S500) according to the user's question. This security question-and-answer intent frame is used to transform natural language questions into searchable and verifiable structured queries.
[0038] The intent frame for security question answering is represented as follows: in, This represents the security question-and-answer intent frame corresponding to the user's question; This indicates the type of question and answer task, including compliance judgment, basis query, hidden danger classification, rectification suggestions, risk classification, list generation, or case retrieval; This refers to the object of the problem, including enterprises, equipment, facilities, locations, personnel, or positions. Indicates work activities or on-site behavior; Indicates the type of risk or the consequences of an accident; Indicates region, industry, regulatory level, or time condition; This indicates the output boundaries required by the user, including whether clauses need to be listed, whether corrective measures are needed, whether a checklist needs to be created, and whether uncertainties need to be output.
[0039] For example, when a user inputs "How to ensure compliance when cleaning a belt conveyor," the system generates a safety question-and-answer intent frame. The task type is rectification suggestion and compliance judgment; the object of the question is a belt conveyor; the work activity is cleaning; the risk types are entanglement, accidental start, and mechanical injury; and the output boundary provides evidence and measures. Based on this, the system retrieves evidence related to cleaning operations, power outages, emergency stops, idler roller protection, drum protection, on-site control panels, and warning signs from a two-layer knowledge graph, rather than simply searching for the words "cleaning" or "conveyor."
[0040] Subsequently, collaborative retrieval and evidence rearrangement (S600) are performed. Based on the security question-answering intent frame, collaborative retrieval is conducted in a two-layer knowledge graph and vector index to obtain a set of candidate evidence. Collaborative retrieval includes graph path retrieval, vector semantic retrieval, and cross-layer rearrangement.
[0041] Graph path retrieval starts from object nodes and operation nodes in the safety question-and-answer intent frame, expanding along relationships such as "belongs to equipment type," "triggers risk," "applicable clauses," "requires control measures," and "exists similar hazards" to obtain regulatory and scenario evidence. Vector semantic retrieval recalls semantically similar content from clause texts, inspection question texts, rectification texts, and question-and-answer history. Cross-layer re-ranking is based on collaborative retrieval scores. The candidate evidence is ranked, with priority given to those that are consistent with the subject, consistent with the conditions, valid in accordance with regulations, and have sufficient on-site evidence.
[0042] In one implementation, the system sets quotas for different types of evidence. For compliance assessment questions, candidate evidence must include at least one regulatory document and one scenario-based document; for rectification suggestion questions, candidate evidence must include at least one control measure document and one similar hazard document; for risk classification questions, candidate evidence must include at least one risk catalog document and one on-site operation document. By setting quotas for different types of evidence, the system avoids generating answers based solely on similar cases or abstract clauses that are detached from real-world scenarios.
[0043] After obtaining the candidate evidence set, applicability verification, conflict verification, and evidence compression are performed (S700). Applicability verification, conflict verification, and evidence compression are performed on the candidate evidence set to obtain enhanced retrieval context.
[0044] The applicability check is used to determine whether candidate clauses are applicable to the current problem. For example, if a clause applies to a belt conveyor, the system needs to confirm whether the equipment in the user's problem is a conveyor; if the clause's applicability conditions involve equipment length, the system needs to confirm whether length information exists on-site, and if missing, it is marked as a condition to be supplemented; if the clause's scope of application is a local regulation, the system needs to confirm whether the company's location is within the corresponding geographical area. If the applicable conditions are missing, the system will not use the clause as a basis for a definitive conclusion, but rather as a "basis to be verified".
[0045] Conflict checking is used to handle differences between multiple standards. The system prioritizes standards based on their level, mandatory nature, implementation time, geographical applicability, and the strictest requirements. When the same issue exists in national standards, local standards, and company regulations, the system prioritizes ensuring compliance with mandatory national standards, and prompts for stricter implementation when local regulations or company regulations impose stricter requirements. When old and new versions of clauses are recalled simultaneously, the system prioritizes clauses based on their time attributes. Determine if a version is valid.
[0046] Evidence compression is used to organize search results into a context suitable for large model inputs. The compressed context includes: the issue object, on-site facts, applicable clauses, clause requirements, applicable conditions, conflict resolution results, similar potential hazards, recommended measures, and uncertainties. Evidence compression does not change the substantive content of the clauses; it only removes duplicate content and irrelevant paragraphs, retaining the clause number, standard name, on-site evidence number, and map path.
[0047] Finally, security specification question-and-answer generation (S800) is performed. The enhanced search context is input into the security specification question-and-answer model to generate the results. The security specification question-and-answer model can be a locally deployed large language model or a large language model fine-tuned with security specification question-and-answer instruction data. To ensure controllable question-and-answer results, the system sets generation boundaries in the prompt template, requiring the model to answer only based on the enhanced search context, and prohibiting the addition of unretrieved clause numbers, the direct use of similar cases as legal basis, and the output of definitive conclusions when applicable conditions are missing.
[0048] In one implementation, the safety specification Q&A results include the following fields: judgment conclusion, applicable conditions, supporting clauses, on-site facts, risk description, corrective measures, items requiring supplementary verification, and credibility. For the question of "whether it is illegal or compliant," the system first outputs the conclusion, then the supporting evidence. For the question of "how to rectify," the system first outputs the rectification goals, then the specific measures. For the question of "what type of risk does a certain hazard belong to," the system outputs the risk type, triggering cause, and associated regulatory obligations.
[0049] For example, regarding the question "How should we respond when maintenance work is not marked with a power cut-off sign?", the system-generated answer includes: This situation falls under the category of insufficient energy isolation during equipment maintenance work; equipment operation should be stopped, power cut off, and "Do Not Close" or "Do Not Operate" signs should be posted before entering hazardous areas or contacting moving parts during maintenance, cleaning, or unblocking operations, and control measures should be taken to prevent accidental start-up; if the on-site distribution cabinet does not indicate the corresponding control equipment, the control object identification should be added to ensure that the power cut-off sign measures can be implemented; if the equipment also lacks emergency stop, interlocks, or protective covers, these should be included in the rectification closed loop. This answer is derived from regulatory obligations, on-site evidence, and cross-level mapping relationships, thus possessing traceability.
[0050] In addition, this application also includes map updating and Q&A feedback steps.
[0051] In some implementations, this application also includes a graph update step. After receiving new specification documents, inspection records, rectification photos, or review records, the system re-parses the clauses and extracts scenarios. If a new specification replaces an old one, the system updates the version relationships in the specification knowledge graph and marks the old specification node as historically valid. If a potential hazard has been rectified, the system adds a rectification node and a review node to the scenario knowledge graph, but does not delete the original hazard node to preserve the audit trail.
[0052] The priority of map updates is represented as follows: in, Represents data objects Update priority; This represents the time-varying factor, which takes a higher value when a new specification is released, repealed, or inspection records are updated. This indicates the scope of impact factor, which takes a higher value when the data object affects multiple industries, devices, or terms. This represents the question-and-answer call frequency factor, which takes a higher value when a certain type of clause or potential issue is frequently queried. This represents the controversy factor, which is a higher value when the question-and-answer feedback shows conflicting or uncertain answers. These are non-negative weighting coefficients. This formula is used to determine which specification nodes, scene nodes, or cross-layer mapping relationships should be updated first.
[0053] This application also provides a security specification intelligent question-answering system based on a two-layer knowledge graph and retrieval enhancement generation model, including a data acquisition module, a specification graph construction module, a scene graph construction module, a cross-layer mapping module, a retrieval enhancement module, a question-answer generation module, and a feedback update module.
[0054] The data acquisition module acquires safety specification texts, basic enterprise information, operational scenario information, equipment and facility information, risk identification information, on-site inspection records, and rectification records. The specification graph construction module builds a specification knowledge graph. The scenario graph construction module builds a scenario knowledge graph. The cross-layer mapping module establishes applicability, violation, satisfaction, and evidence support relationships between the specification knowledge graph and the scenario knowledge graph. The retrieval enhancement module performs graph path retrieval, semantic vector retrieval, evidence rearrangement, applicability verification, and conflict verification. The question-and-answer generation module uses a large-scale safety specification question-and-answer model to generate structured answers. The feedback update module updates the graph and index based on new specifications, new inspection records, and user feedback.
[0055] This application also provides an electronic device, including a processor, a memory, a communication interface, and a bus. The memory stores a computer program, and when the processor executes the computer program, it implements the security-compliant intelligent question-answering method based on a two-layer knowledge graph and retrieval-enhanced generative model as described in any embodiment of this application.
[0056] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in any embodiment of this application. The computer-readable storage medium may include a read-only memory, a random access memory, a magnetic disk, an optical disk, a removable storage device, or other media capable of storing program code.
[0057] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. For those skilled in the art, any equivalent substitutions, modifications, or improvements made to the technical solutions of this application without departing from the technical concept of this application should be included within the scope of protection of this application.
Claims
1. A security-compliant intelligent question-answering method based on a two-layer knowledge graph and retrieval-enhanced generative model, characterized in that, Includes the following steps: S100: Obtain safety specification texts, basic enterprise information, work scenario information, equipment and facility information, risk identification information, on-site inspection records, and rectification records; S200, perform clause-level parsing on the security specification text to construct a specification knowledge graph; S300: Extract the enterprise basic information, operation scenario information, equipment and facility information, risk identification information, on-site inspection records and rectification records in a scenario-based manner to construct a scenario knowledge graph; S400, establish a cross-layer mapping relationship between the standard knowledge graph and the scene knowledge graph to obtain a two-layer knowledge graph; S500 generates a security question-and-answer intent frame based on the user's question; S600, based on the security question-answering intent frame, a collaborative retrieval is performed in the two-layer knowledge graph and vector index library to obtain a set of candidate evidence; S700, perform applicability verification, conflict verification and evidence compression on the candidate evidence set to obtain retrieval enhancement context; S800, the enhanced retrieval context is input into the security specification question-answering model to generate security specification question-answering results.
2. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model as described in claim 1, characterized in that, The normative knowledge graph includes legal nodes, standard nodes, clause nodes, terminology nodes, applicable object nodes, operational activity nodes, equipment and facility nodes, risk type nodes, control measure nodes, and legal liability nodes; the scenario knowledge graph includes enterprise nodes, industry nodes, regional nodes, job node, equipment node, location node, operation node, hidden danger node, evidence node, rectification node, and review node.
3. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model as described in claim 1, characterized in that, The cross-layer mapping relationships include: the applicability relationship between scenario equipment and the applicable objects of the standard; the matching relationship between scenario operations and standard operation activities; the violation relationship between scenario hazards and standard prohibitions; the satisfaction relationship between scenario rectification and standard control measures; and the supporting relationship between on-site evidence and clause conclusions.
4. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model as described in claim 1, characterized in that, The clause-level parsing of the safety specification text includes: identifying the specification name, specification number, issuing authority, implementation date, expiration status, chapter structure, clause number, mandatory statements, prohibited statements, applicable conditions, and exception conditions, and binding the clause number with the mandatory statements, prohibited statements, applicable conditions, and exception conditions as a specification obligation unit.
5. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model according to claim 4, characterized in that, The normative obligation unit is represented by an obligation fingerprint, which includes an object slot, a condition slot, a behavior slot, an indicator slot, and a consequence slot. The object slot is used to represent the equipment, facilities, places, personnel, or work activities to which the clause applies. The condition slot is used to represent the industry, length, medium, risk level, work status, or time conditions required for the clause to take effect. The behavior slot is used to represent the action requirements that should be set up, should not be implemented, should be trained, should be publicized, should be marked off, or should be inspected. The indicator slot is used to represent verifiable parameters such as distance, height, quantity, cycle, training hours, or status. The consequence slot is used to represent risk consequences, illegal consequences, or rectification consequences.
6. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model according to claim 1, characterized in that, The process of generating a safety Q&A intent frame based on user questions includes: identifying the subject, work activity, equipment and facilities, risk type, Q&A task, geographical conditions, and evidence requirements in the user questions; the Q&A task includes one or more of the following: compliance judgment, clause basis query, hazard classification, rectification suggestion generation, risk level judgment, checklist generation, and similar case retrieval.
7. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model according to claim 1, characterized in that, The collaborative retrieval includes graph path retrieval and semantic vector retrieval; the graph path retrieval is used to obtain evidence of entity relationship, obligation relationship or risk relationship with the security question and answer intent frame; the semantic vector retrieval is used to obtain clause fragments, on-site description fragments and rectification fragments that are semantically similar to the user's question; the results of the collaborative retrieval are rearranged by cross-layer matching scores.
8. A security-compliant intelligent question-answering method based on a two-layer knowledge graph and retrieval enhancement generation model according to claim 7, characterized in that, The cross-layer matching score is determined based on semantic similarity, keyword overlap, applicability, standard validity, and evidence freshness. It is used to ensure that the security standard question-and-answer model prioritizes content that simultaneously satisfies the standard-layer basis and scenario-layer evidence.
9. The intelligent question-answering method for security specifications based on a two-layer knowledge graph and retrieval enhancement generation model according to claim 1, characterized in that, The conflict verification includes: when multiple regulatory clauses give different requirements for the same object, generating candidate conclusions according to the priority of laws and regulations, mandatory standards, local special provisions, new laws, and stricter control measures; when the scenario evidence is insufficient to meet the conditions for the application of the clauses, outputting information items to be supplemented, and not directly generating a definitive compliance conclusion.
10. A security-compliant intelligent question-answering system based on a two-layer knowledge graph and retrieval-enhanced generative model, characterized in that, include: The data acquisition module is used to acquire safety specification texts, basic enterprise information, work scenario information, equipment and facility information, risk identification information, on-site inspection records, and rectification records. The specification graph construction module is used to perform clause-level parsing of the security specification text and construct a specification knowledge graph. The scenario graph construction module is used to extract scenario-based data from enterprise-side data and on-site data to construct a scenario knowledge graph. The cross-layer mapping module is used to establish cross-layer mapping relationships between the canonical knowledge graph and the scenario knowledge graph to obtain a two-layer knowledge graph; The retrieval enhancement module is used to generate a security question-and-answer intent frame based on the user's question, and to perform collaborative retrieval and evidence verification based on the security question-and-answer intent frame; The question-and-answer generation module is used to generate security-compliant question-and-answer results based on the enhanced search context.