A two-vote analysis method based on a graph database and a large language model

The two-ticket analysis method constructed by graph database and large language model solves the problem of handling complex correlations in two-ticket analysis in new power systems, realizes intelligent control and dynamic risk warning of the whole process, and improves the safety and efficiency of power operation.

CN122240694APending Publication Date: 2026-06-19CHINA THREE GORGES CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-25
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing two-ticket analysis technology cannot efficiently handle complex relationships, deeply parse unstructured text, or dynamically adapt to complex operating scenarios, resulting in analysis efficiency, accuracy, and security control capabilities that cannot meet the needs of new power systems.

Method used

A graph database is used to construct a two-ticket association graph model. Combined with a large language model and a multi-agent collaborative framework, multi-hop queries, risk quantification, and intelligent analysis are realized. Intelligent analysis reports that conform to power industry standards are generated, and the system is optimized through closed-loop iterative optimization.

Benefits of technology

It achieves full-dimensional identification of risks associated with both types of tickets, improves operational safety and efficiency, provides interpretable intelligent decision support, and has strong scenario adaptability and continuous iteration capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122240694A_ABST
    Figure CN122240694A_ABST
Patent Text Reader

Abstract

This invention provides a two-ticket analysis method based on graph databases and large language models. For the first time, it constructs a dynamic graph database model for power system operation tickets and work tickets, enabling multi-dimensional relational representation of station components over time, breaking through the limitations of traditional relational databases. Simultaneously, it embeds dynamic risk weights to achieve real-time risk quantification of multi-hop paths, solving the challenge of risk assessment for cross-ticket operations. This patent establishes a power industry-specific two-layer retrieval enhancement generation architecture, combining multi-dimensional vector indexes of voltage level and equipment type to achieve accurate retrieval of professional content, significantly improving the compliance and accuracy of the analyzed content. This patent pioneers a master-slave multi-agent collaborative architecture, automating the decomposition and execution of complex analysis tasks and risk simulation verification through closed-loop collaboration of multiple types of agents, achieving automated control of the entire two-ticket analysis process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a two-vote analysis method based on graph databases and large language models. Background Technology

[0002] With the continuous advancement of the construction of new power systems, the power grid topology is becoming increasingly complex. The large-scale integration of new energy power plants and new power equipment has led to a significant increase in the frequency of cross-regional, multi-equipment, and multi-entity collaborative maintenance and operation. The data on two-tickets (work permits and work tickets) exhibits significant characteristics of massive volume, multi-source heterogeneity, and complex scenarios. The traditional two-ticket analysis model, which relies on manual review and rule matching, as well as the existing semi-automated two-ticket management technology, are no longer suitable for the safety control requirements of new power systems. The industry has put forward higher requirements for the intelligent full-process analysis of two-tickets, in-depth mining of correlation relationships, and accurate early warning of dynamic risks.

[0003] Currently, the industry has conducted relevant technological research and development on the intelligent management and analysis of two-tickets (work permits and tickets), resulting in various AI-based two-ticket management systems and analysis methods. These have, to a certain extent, achieved automated upgrades in two-ticket management, reduced the workload of substation operators in manually issuing and reviewing tickets, and improved the efficiency of basic two-ticket management. Representative related technical solutions are as follows: One is a Chinese invention patent application with publication number CN116091030A entitled "A Two-Ticket Management System Based on Artificial Intelligence". This solution discloses a two-ticket management system that includes a central server and topology analysis module, anti-misoperation database, anti-misoperation module, path search module, deep learning module, semantic recognition module, and two-ticket monitoring module connected to the central server. It adopts power grid intelligent analysis and intelligent reasoning technology to realize the automatic generation and automatic review of operation tickets, which can reduce the workload of operation personnel in issuing and reviewing tickets, while improving the standardization, security and efficiency of switching operations.

[0004] The second is a Chinese invention patent application with publication number CN117112848A entitled "A Method, Device, Equipment and Medium for Two-Ticket Analysis Based on Distribution Network". This solution discloses a technical solution that receives target query conditions, determines the target type ticket corresponding to the target query unit, matches the corresponding target query content according to the target query requirements, and finally generates the corresponding analysis results based on the matched target analysis conditions. This solution realizes the statistical analysis of two tickets in the distribution network and improves the efficiency and accuracy of two-ticket analysis.

[0005] The inventors have discovered that while the existing technical solutions have achieved technical improvements in automated management and basic analysis of the two-ticket system, they still suffer from many insurmountable technical defects and cannot meet the core requirements of comprehensive and intelligent management and control of the two-ticket system under the new power system. Specifically: Regarding the technical solution disclosed in CN116091030A, firstly, the system relies on traditional relational databases to store error prevention rules and operation records. However, relational databases struggle to efficiently handle complex relationships such as multi-level dependencies in equipment states and cross-substation operational linkages within the power grid topology. This leads to significant limitations in the breadth and depth optimization of the path search algorithm, making it unsuitable for complex scenarios involving multi-device collaborative operations and cross-substation coordinated work. Secondly, its semantic recognition module only parses fixed templates of dispatch instructions, lacking the ability to flexibly understand non-standardized, colloquial natural language instructions. Firstly, when faced with non-standard instruction scenarios, the generated operation tickets are prone to content deviations, limiting their applicability. Secondly, its deep learning module does not clearly define the coverage of training data. If the training samples are insufficient or the scenario coverage is limited, it will directly lead to insufficient generalization ability of the model under actual complex working conditions, making it difficult to cope with special operation scenarios such as sudden adjustments to the power grid. Finally, the error prevention verification of its two-ticket monitoring module is mainly based on static rules, which cannot dynamically integrate real-time power grid operation data from sources such as SCADA systems. This can easily lead to risk misjudgment or omission, resulting in insufficient reliability of safety control in scenarios where the power grid operation status changes dynamically.

[0006] Regarding the technical solution disclosed in CN117112848A, firstly, this method relies on preset static analysis conditions and lacks dynamic adaptability to complex operation scenarios. When special circumstances such as sudden power grid failures cause abnormalities in the execution time and process of the operation ticket, the fixed judgment threshold cannot accurately reflect the actual operation risk, resulting in insufficient accuracy in risk assessment. Secondly, the extraction of its target query content is limited to structured fields such as time and ticket number, without in-depth mining and analysis of unstructured text such as semantic details of dispatch instructions and work ticket remarks, resulting in a single analysis dimension and an inability to comprehensively cover potential risk points in the entire process of both tickets. Thirdly, this solution does not consider the coordination verification of operation tickets and work tickets under the same maintenance task, and cannot identify logical conflicts across tickets, easily overlooking major safety hazards such as contradictions between operation ticket steps and work ticket safety measures. Finally, the generation of its analysis results relies on hard-coded rules, and can only output labeled judgment results, failing to provide interpretable conclusions and targeted optimization suggestions, making it difficult to form effective decision support for power operation safety management.

[0007] In summary, existing technologies related to two-ticket analysis cannot achieve deep fusion of multi-source heterogeneous data on two tickets and efficient mining of complex relationships. They struggle to simultaneously balance efficiency, accuracy, scenario adaptability, and decision support capabilities, failing to meet the core requirements of intelligent full-process management and precise dynamic risk early warning in new power systems. Therefore, overcoming the numerous shortcomings of existing technologies and developing a two-ticket analysis method capable of efficiently handling complex relationships between two tickets, deeply parsing unstructured text, dynamically adapting to complex operational scenarios, and possessing interpretable intelligent decision support capabilities has become an urgent technical problem to be solved in this field. Summary of the Invention

[0008] This invention provides a two-ticket analysis method based on graph databases and large language models, which solves the core technical problems of existing power two-ticket analysis technologies, such as low efficiency in handling complex relationships, incomplete identification of cross-ticket risks, and poor adaptability to dynamic scenarios, resulting in analysis efficiency, accuracy, and safety control capabilities that cannot meet the safety production control requirements of new power systems.

[0009] To solve the above-mentioned technical problems, the technical solution adopted by this invention is: a two-vote analysis method based on graph databases and large language models, comprising the following steps: S1. Data preprocessing: Obtain the original data of the two tickets of the power system, clean and standardize the original data of the two tickets, and integrate multi-source related data such as power equipment ledger, historical fault records, and personnel scheduling data to obtain a structured two-ticket dataset; S2. Graph Database Modeling and Association Analysis: Based on the structured two-ticket dataset, construct a two-ticket association graph model in the graph database, which includes entities such as two tickets, equipment, personnel, and time intervals. Utilize the multi-hop query capability of the graph database to perform two-ticket association conflict detection and risk quantification analysis, and output the graph analysis results. S3. Retrieval Enhancement Generation Processing: Construct a vectorized knowledge base specifically for the power industry, generate retrieval instructions based on the graph analysis results, and retrieve matching power industry standards and case data from the vectorized knowledge base to obtain retrieval enhancement support data; S4. Multi-Agent Collaborative Task Processing: Through a preset multi-agent collaborative framework, the task of generating two analysis reports is broken down into multiple executable sub-tasks with priorities. Corresponding functional tools are scheduled to complete the execution and result verification of each sub-task, and integrated to obtain structured analysis intermediate results. S5. Intelligent Analysis Report Generation: The graph analysis results, retrieval enhancement support data, and structured analysis intermediate results are integrated into standardized prompt words, which are then input into a large language model adapted and fine-tuned for the power industry to generate two intelligent analysis reports that conform to power industry standards. S6. Closed-loop iterative optimization: Collect manual review feedback data from the two-vote intelligent analysis reports, and based on the feedback data, complete the vectorized knowledge base update, incremental fine-tuning of the large language model, and optimization of the two-vote association graph model to achieve closed-loop self-optimization of the system.

[0010] Preferably, in step S2, the two-ticket association graph model is configured with corresponding attribute parameters for nodes and edges, with nodes corresponding to entities and directed edges corresponding to the association relationships between entities. The attribute parameters of the nodes include ticket status, equipment risk level, and personnel qualification level. The attribute parameters of the edges include dynamic risk weight, operation order, and association timeliness. The dynamic risk weight includes at least equipment load rate and personnel workload coefficient.

[0011] Preferably, in step S2, the two-ticket association conflict detection includes at least the following: conflict detection of multiple overlapping operations on the same device, conflict detection of multiple high-risk tickets from the same person in the same time period, and detection of excessive concurrent operation load on the device; wherein, the conflict detection of overlapping operations is achieved by calculating the probability of time conflict P, and the calculation formula is: P = overlap duration / total operation duration. When P is greater than a preset conflict threshold, a high-risk warning is triggered.

[0012] Preferably, in step S3, the power industry-specific vectorized knowledge base adopts a two-layer architecture, including a structured procedure knowledge base and an unstructured fault case base. The vectorized knowledge base is configured with a multi-dimensional vector index based on voltage level and equipment type. During retrieval, the cosine similarity between the query vector and the knowledge base document vector is calculated to return the Top-K target document data with the highest matching degree.

[0013] Preferably, in step S4, the multi-Agent collaborative framework adopts a master-slave agent group architecture, which includes at least a scheduling agent, a verification agent, and a generation agent. The scheduling agent is responsible for task decomposition, priority allocation, and tool scheduling; the verification agent is responsible for logical consistency verification and risk simulation verification of sub-task results; and the generation agent is responsible for the structured integration of multi-source results. The three types of agents form a closed-loop collaborative execution link.

[0014] Preferably, the verification agent uses the Monte Carlo simulation method to predict the failure probability of the ticket operation sequence, and also supports counterfactual reasoning verification to achieve accurate quantification of risk level.

[0015] Preferably, in step S5, the large language model, after being adapted and fine-tuned for the power industry, completes the domain adaptation by fine-tuning the instructions of power industry dispatching terminology, safety procedures, and operating specifications. At the same time, a dynamic template engine is configured to automatically match the corresponding report generation template according to the risk level of the ticket. The large language model adopts the thinking chain technology to transform the topology analysis results into natural language descriptions that conform to power industry standards.

[0016] Preferably, in step S6, the collection of manual review feedback data is carried out by automatically identifying manual modification traces in the report through a negative sample mining algorithm, extracting erroneous content and labeling the error type to form a negative sample dataset; based on the negative sample dataset, incremental fine-tuning of the large language model, correction of the vectorized knowledge base, and optimization of the graph query algorithm are performed periodically.

[0017] Preferably, in step S1, the standardization process includes at least unifying device naming rules, filling in missing timestamps, and removing duplicate ticket records; the multi-source associated data and the original data of the two tickets are associated and mapped through the unique device number, employee number, and unique ticket identifier to generate a structured two-ticket dataset with a unique identifier.

[0018] Preferably, in step S2, when two new votes are added, the nodes, edges and corresponding attribute parameters of the two votes association graph model are updated in real time, and the full conflict detection process is automatically triggered; at the same time, the graph index is optimized regularly, and graph data in high-frequency query scenarios are cached to improve query response efficiency.

[0019] This invention provides a two-vote analysis method based on graph databases and large language models, which has the following beneficial effects: 1. Comprehensively enhance the ability to prevent and control safety risks in power operations. This invention is the first to construct a dynamic graph database model for two-ticket data, breaking through the technical limitations of traditional relational databases. Through multi-hop query capabilities, it can accurately uncover hidden safety hazards that are difficult to identify by traditional methods, such as cross-ticket logical conflicts, overlapping equipment operations, personnel overload operations, and cross-site operation linkage risks. Combined with the dynamic risk weights embedded in the edge attributes, it realizes real-time risk quantification calculation, achieving full-dimensional and comprehensive identification of two-ticket risks. It significantly reduces the probability of safety accidents in power maintenance, switching operations, and other operations from the source, providing a solid technical guarantee for safe power production.

[0020] 2. Significantly improves the efficiency of the entire two-ticket management process. This invention automates the entire process from two-ticket data collection, cleaning and standardization, correlation analysis, risk assessment to professional report generation. It reduces the two-ticket analysis work that traditionally required multiple working days to be completed manually to minutes, completely solving the industry pain point of low efficiency in traditional manual ticket review and analysis. It significantly reduces the workload of substation operation and dispatch management personnel in ticket issuance, review, and analysis, and significantly improves the overall efficiency of power two-ticket management and operation execution.

[0021] 3. This invention achieves compliant, accurate, and interpretable intelligent decision support. Through a dual-layer RAG architecture specific to the power sector, it enables precise retrieval of industry regulations and historical cases, effectively avoiding the illusion problem of general large models and ensuring that the analysis conclusions fully comply with current power safety regulations and industry technical specifications. At the same time, through domain-adapted fine-tuning of the large language model, the generated analysis report not only includes accurate risk warnings but also comes with implementable optimization and disposal solutions, authoritative regulatory references, and similar historical case references. It has strong interpretability and operability, providing professional and authoritative intelligent support for power operation safety management and dispatching decisions.

[0022] 4. Possessing strong scenario adaptability and continuous iteration capability, this invention's pioneering master-slave multi-agent collaborative architecture can flexibly decompose and adapt to two-ticket analysis scenarios of different voltage levels, different types of substations, and different levels of complexity, and can cope with the analysis needs of dynamic operation scenarios such as sudden power grid condition adjustments. At the same time, it constructs a closed-loop self-optimization system of "analysis-execution-feedback", which continuously iterates and optimizes graph query algorithms, large language models, and industry knowledge bases through automatic negative sample mining and manual review feedback. With plug-in extension design, it can quickly access new data sources, adapt to the business changes of new power systems such as new energy grid connection, new equipment commissioning, and industry standard updates, ensuring the long-term applicability, advancement, and stability of the system.

[0023] 5. Possessing broad industry promotion value and industrialization prospects, the technical solution of this invention can be fully adapted to all power business links such as power generation, transmission, substation and distribution, and is compatible with various standard ticket types such as electrical work tickets and switching operation tickets. It can be directly applied to the two-ticket management scenarios of power grid enterprises and power generation groups at all levels, providing a complete, implementable and highly reliable technical solution for the intelligent transformation of two-ticket management in the power industry, and possessing strong industry universality and industrialization promotion value. Attached Figure Description

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0025] This invention discloses a two-ticket analysis method based on graph database and large language model, which is applied to the intelligent analysis and safety management of the entire process of operation tickets and work tickets in power systems. It is compatible with standard ticket types such as electrical first type work ticket, electrical second type work ticket, line work ticket, and switching operation ticket in the power industry, covering all business links of power generation, transmission, substation and distribution.

[0026] like Figure 1 As shown, the specific implementation steps of this method are as follows: Step S1: Data Extraction and Preprocessing This step involves collecting, cleaning, standardizing, and integrating the two sets of data and related data from multiple sources to provide a high-quality structured dataset for subsequent analysis. The specific implementation process consists of three stages: S1.1 Raw Data Identification and Acquisition By connecting to the group's two-ticket management system and daily defect reporting system via API, raw data from operation tickets and work tickets is collected incrementally on a daily basis. The core key fields collected include: ticket number, ticket type, affiliated site, operation steps, equipment number, equipment name, work supervisor, operator, supervisor, authorized time, start time, end time, work content, safety measures, and remarks. The collected raw data is temporarily stored in a PostgreSQL intermediate database, and basic data validation rules are set to mark invalid data with missing ticket numbers, core timestamps, or equipment numbers, which then proceed to the subsequent completion process.

[0027] S1.2 Data Cleaning and Standardization 1. Perform three standardized processing operations on the collected raw data: Equipment naming standardization: Establish a unified equipment naming mapping dictionary for the group, and uniformly map non-standard names such as "10kV switchgear A", "10kV-A switchgear", and "10kV No.1 switchgear A bay" in different tickets to the unique equipment number corresponding to the group's equipment ledger (format example: CTG-SG-10kV-001-A), ensuring that the identification of the same equipment is unique in all tickets. 2. Timestamp standardization: All time fields are standardized using the format "YYYY-MM-DDHH:MM:SS". For missing operation end time and permission time, linear interpolation is performed by associating with personnel schedules and SCADA system equipment status change records to complete the missing time. 3. Deduplication and invalid data processing: A unique primary key is constructed based on "ticket number + station number" to remove duplicate ticket records generated by system synchronization. When there are multiple records for the same ticket number, only the latest version of the valid record is retained, and invalid ticket data that has been cancelled or revoked is removed.

[0028] S1.3 Multi-source data integration and structured output Integration of multiple business systems to complete related data integration: Integration with the equipment ledger management system to supplement attributes such as equipment model, rated load threshold, maintenance cycle, and historical maintenance records; integration with the historical fault management system to mark equipment that has experienced faults in the past 3 years, and to indicate the fault type and frequency; integration with the human resources management system to supplement personnel qualification level, safety training validity period, and historical work load data.

[0029] All integrated data is used to generate a unique snowflake ID as the primary key, converted into JSON structured format, and output to the graph database import interface to complete the full preprocessing process.

[0030] Step S2: Graph Database Modeling and Association Analysis This step, based on the preprocessed structured data, constructs a two-ticket relationship graph model. It leverages the multi-hop query capabilities of the graph database to mine complex relationships and quantify risks. The specific implementation process consists of three stages: S2.1 Construction of the Two-Vote Relationship Graph Model Using the Neo4j enterprise landscape database, a dynamic graph model with five dimensions of association—"site-equipment-tickets-personnel-time"—is constructed, defined as follows: 1. Node Definition: Five core node types are defined: Operation Ticket, Work Ticket, Equipment, Personnel, and Time Range. Each node is assigned a unique ID, corresponding one-to-one with the Snowflake ID in the preprocessed data. Each node type is configured with exclusive attributes. Operation Ticket / Work Ticket node attributes include ticket number, ticket type, site number, ticket status (not executed / in execution / completed / invalidated), risk level, work content, and permitted time. Equipment node attributes include unique equipment number, equipment name, voltage level, rated load threshold, affiliated site, and fault history markers. Personnel node attributes include employee ID, name, qualification level, department, and safety training validity period. Time Range node attributes include start time, end time, associated date, and work type.

[0031] Edge and attribute definition: Directed edges define the relationships between entities. Core edge types include "Personnel-Execution-Operation Ticket", "Personnel-Responsible-Work Ticket", "Equipment-Involved-Operation Ticket", "Equipment-Involved-Work Ticket", "Operation Ticket-Corresponding-Work Ticket", "Ticket-Attribution-Time Interval", "Personnel-Attribution-Time Interval", and "Equipment-Attribution-Site". Attributes are configured for each edge, including operation sequence, dynamic risk weight, association timeliness, and job type. The dynamic risk weight includes real-time equipment load rate, personnel workload coefficient, and equipment failure risk coefficient. The weight value ranges from 0 to 1, with higher values ​​corresponding to higher risks.

[0032] S2.2 Multi-hop query conflict detection and risk quantification Based on the completed graph model, multi-hop related queries are implemented using the Cypher query language, covering the detection of 12 typical risk patterns. The core detection logic is as follows: 1. Equipment Operation Time Conflict Detection: Query multiple tickets associated with the same equipment node within overlapping time intervals, calculate the time conflict probability P, and the calculation formula is: Overlap duration = max(0, min(ticket A end time, ticket B end time) - max(ticket A start time, ticket B start time)) P = Overlap duration / Total operation time of ticket A The preset conflict threshold is 0.3. When P > 0.3, a high-risk warning is triggered and the device is marked as a risk point of conflict in operation. At the same time, combined with the rated load threshold of the device, when the total load of multiple operations associated with the same device exceeds 90% of the rated load in the same time interval, an overload warning is triggered.

[0033] 2. Personnel workload conflict detection: Query two or more high-risk tickets associated with the same personnel node within the same time interval, or ticket tasks with a single day's work duration exceeding 8 hours. Combined with personnel qualification level, identify risks of unqualified work and overloaded work and trigger warnings.

[0034] 3. Cross-ticket logic conflict detection: Through a 3-hop association query, identify the operation ticket and work ticket corresponding to the same maintenance task, and verify whether the power outage range and grounding switch settings of the operation ticket match the safety measures of the work ticket. If there is any inconsistency, mark it as a cross-ticket logic conflict risk point.

[0035] After completing full conflict detection, calculate the comprehensive risk score R for each ticket. The calculation formula is: R = Σ (conflict probability × corresponding node risk weight). Based on the R value, the risk level of each ticket is divided into three levels: high, medium, and low. Finally, the graph analysis results containing all risk points, risk levels, conflict details, and comprehensive scores are output.

[0036] S2.3 Dynamic Graph Update and Performance Optimization A real-time data synchronization interface is set up so that when new ticket data is added to the two-ticket management system, incremental updates of graph database nodes, edges, and attributes are automatically triggered, and a full-scale conflict detection process is triggered simultaneously. The graph index is optimized weekly. For high-frequency query scenarios such as "query of all tickets for a certain station this week" and "query of related tickets for a certain device in the past month", a dedicated graph index is built. At the same time, the graph data accessed frequently is cached in memory to keep the query response time within 200ms. The graph database snapshot is backed up daily to support historical data backtracking and fault rollback.

[0037] Step S3: Retrieval Enhancement Generation (RAG) module execution This step involves building a dedicated knowledge base for the power industry, providing authoritative and compliant knowledge support for the two-ticket analysis, and resolving the illusion of large models. The specific implementation process consists of three stages: S3.1 Construction of Vectorized Knowledge Base in the Power Sector A two-tiered vectorized knowledge base specifically for the power industry is constructed. The first tier is a structured procedure knowledge base, which includes officially effective procedure documents such as the "Power Safety Work Procedures", industry technical standards, group internal equipment operation manuals, and site safety management specifications. The second tier is an unstructured case library, which includes historical fault cases, two-ticket violation cases, and typical maintenance operation cases from the group over the past 10 years.

[0038] The knowledge base documents are preprocessed and split into semantic paragraphs, with the length of each paragraph controlled within 512 tokens. The BERT-wwm pre-trained embedding model for the power domain is used to convert the split text into 768-dimensional embedding vectors, which are then stored in the Milvus vector database. Metadata is added to each vector, including document source, effective date, voltage level, and equipment type, to build a multi-dimensional vector index of "voltage level-equipment type". This supports precise filtering during retrieval and automatically excludes expired or invalid procedures.

[0039] S3.2 Dynamic Retrieval and Association Matching Based on the risk points, equipment types, and work content in the graph analysis results output in step S2, a search query statement is automatically generated (examples: "Safety Specifications for 220kV Main Transformer Power Outage Operations" and "Case Study on Handling Time Conflicts in Busbar Maintenance Operations"). The search query statement is converted into an embedding vector of the same dimension, and the cosine similarity between the query vector and the knowledge base vector is calculated using the following formula: Similarity = (Q × D) / (|Q| × |D|) Where Q is the query vector and D is the knowledge base document vector. The system returns the top-5 relevant paragraphs based on cosine similarity as supporting data for enhanced retrieval. It also labels each retrieval result with the corresponding procedure clause number and case number, providing authoritative and compliant citation basis for subsequent report generation.

[0040] S3.3 Knowledge Base Iterative Optimization The vector database is updated monthly by incorporating the latest power industry standards, new fault cases, and operating procedures added by the group. For manually reviewed and marked retrieval errors, the embedding model parameters and document segmentation strategies are adjusted to optimize retrieval accuracy. A knowledge citation statistics mechanism is established to cache frequently cited procedures and case content, thereby improving retrieval response efficiency.

[0041] Step S4: Multi-Agent Collaborative Task Processing This step utilizes a master-slave agent group architecture to automate the decomposition, scheduling, execution, and result verification of two-ticket analysis tasks, adapting to the flexible analysis needs of complex scenarios. The specific implementation process consists of three stages: S4.1 Multi-Agent Collaborative Architecture Design A master-slave agent group architecture is built based on the LangChain framework, with three types of core agents set up to form a closed-loop collaborative execution chain: 1. Scheduling Agent: The core control node, responsible for the overall planning, subtask decomposition, priority allocation, and tool scheduling of the two-ticket analysis report generation task; 2. Verification Agent: Responsible for verifying the logical consistency of the subtask output results, conducting risk simulation verification, and reviewing the accuracy of the data; 3. Generating Agent: Responsible for the structured integration and format standardization of results from multiple sub-tasks.

[0042] S4.2 Task Decomposition and Scheduling Execution After receiving the two analysis report generation instructions, the scheduling agent first breaks down the overall task into five priority subtasks, ordered from highest to lowest priority: 1) conflict detection and risk quantification subtask, 2) compliance verification subtask, 3) risk handling suggestion generation subtask, 4) report content structure integration subtask, and 5) report format standardization subtask. Timeout thresholds are set for each subtask, with a timeout threshold of 5 seconds for the conflict detection subtask and 10 seconds for the other subtasks.

[0043] The scheduling agent assigns corresponding execution tools to each subtask, including: calling the Neo4j graph database API to perform graph queries and conflict detection, calling Python scripts to perform risk score calculations, calling the RAG module interface to perform knowledge base retrieval, and calling PDF generation tools to perform report format conversion. When a high-risk conflict is detected, the scheduling agent automatically adjusts the task priority, prioritizing the use of the RAG module to retrieve corresponding emergency response procedures and historical cases, and generating response suggestions in advance.

[0044] S4.3 Result Validation and Fault Tolerance The verification agent performs logical consistency checks on the output of each subtask. Core verification items include: consistency between equipment load data and operation ticket time intervals, validity of the effective date of the retrieval procedure, logical accuracy of risk score calculation, and matching of tickets with personnel qualifications. Simultaneously, the verification agent uses the Monte Carlo simulation method to perform 1000 simulation samplings on the operation sequence of high-risk tickets, calculates the probability of failure in the operation sequence, completes secondary calibration of risk levels, supports counterfactual reasoning verification, and can simulate risk changes after adjusting the operation sequence and personnel configuration.

[0045] When a subtask times out or fails, a fault-tolerance mechanism is automatically triggered: for example, if a database API call times out, the system automatically switches to cached graph snapshot data for query execution; if the similarity of the RAG search results is less than 0.6, the system automatically expands the search keywords or switches to the general power regulations knowledge base for a secondary search to ensure the completion rate of the subtask. After all subtasks are completed, an Agent is generated to merge the output results of multiple tools, generating structured analysis intermediate results in JSON format. Simultaneously, a complete task execution log is recorded for subsequent fault tracing and performance optimization.

[0046] Step S5: Generation of Two-Vote Intelligent Analysis Report This step, based on multi-source analysis results, generates a two-vote analysis report conforming to power industry standards through a domain-adapted large language model. The specific implementation process consists of three stages: S5.1 Standardized Prompt Template Construction A standardized prompt word template specifically for building large language models. The template consists of three parts: 1. Fixed Constraint Section: Clearly define the model role as "Senior Expert in Power Industry Two-Ticket Analysis" and set core generation rules: strictly generate content based on the input graph analysis results and retrieved enhanced supporting data, prohibit the fabrication of unretrieved regulations and cases, and must mark all cited regulation numbers and case sources, and use standard power industry terminology throughout; 2. Dynamic Variable Section: Set up populated placeholders, including {Basic Ticket Information List}, {Details of Conflict Risk Points}, {Comprehensive Risk Score and Level}, {Retrieve Enhanced Support Data}, and {Compliance Verification Results}. Fill the corresponding output results of steps S2-S4 into the placeholders. 3. Generation Requirements Section: Clearly define the report's fixed structure, format specifications, length requirements, and the feasibility requirements for risk optimization recommendations.

[0047] S5.2 Large Language Model Generation and Processing The DeepSeek-V2 large language model, adapted and fine-tuned for the power industry, is adopted. The model completes domain adaptation through instruction fine-tuning. The fine-tuning dataset includes power industry dispatching terminology, safety procedures, two-ticket standard templates, historical high-quality two-ticket analysis reports, and manually reviewed and corrected samples, totaling 100,000 instruction data, enabling the model to fully grasp the professional terminology and standard requirements of two-ticket analysis in the power industry.

[0048] The completed prompts are input into the fine-tuned large language model. The model uses the Chain-of-Thought technique to first logically decompose the graph analysis results, and then combine them with the retrieved procedures and cases to complete risk analysis, compliance determination, and optimization suggestions. Finally, it outputs two intelligent analysis reports that meet the requirements.

[0049] The model is configured with a dynamic template engine that automatically switches report templates based on the overall risk level of each ticket: high-risk tickets use a detailed template, adding fault probability simulation results, emergency response plans, and historical case comparisons; medium- and low-risk tickets use a simplified template, highlighting key risk points and optimization suggestions. The generated report has a fixed structure including: ① Execution Overview: listing the scope of the tickets analyzed, related stations, equipment, personnel, and time intervals; ② Risk Analysis: listing each detected conflict risk point, risk level, triggering basis, and fault probability; ③ Compliance Verification: comparing against power safety regulations, listing compliance deviations and corresponding clauses; ④ Optimization Suggestions: providing feasible optimization solutions for each risk point; ⑤ Citation Basis: listing all referenced regulation clauses and historical case numbers. The report supports the insertion of Gantt charts, risk heatmaps, and other visual charts to intuitively display time conflicts and risk distribution.

[0050] S5.3 Report Output and Manual Review Interface The generated reports can be automatically converted to Markdown, PDF, and Word formats according to enterprise management requirements, and metadata such as report generation time, data version number, and responsible person can be added. A web-based manual review interface is also provided, allowing reviewers to annotate, modify, and mark errors on the reports. Reviewers' comments are automatically linked to the corresponding data nodes, synchronously generating a complete review log that records the modifications, reviewers, and review time. The final confirmed report is then synchronously archived in the group's two-invoice management system.

[0051] Step S6: Closed-loop iterative optimization This step involves collecting and applying manual feedback to achieve continuous self-optimization of the system, ensuring long-term adaptability to business needs. The specific implementation process consists of three stages: S6.1 Negative Sample Acquisition and Processing Establish an automatic negative sample mining and collection mechanism. Through a preset algorithm, automatically identify manual modification traces in reports, extract the content before and after modification, and label the error type, including factual errors, logical contradictions, non-standard terminology, incorrect reference to procedures, and misjudgment of risks, and store it in the negative sample database. Perform statistical analysis on the negative sample database every month, and optimize the graph query algorithm, data standardization rules, and model prompt word templates for high-frequency error types.

[0052] S6.2 Model and Knowledge Base Iterative Updates Each month, the graph relationship prediction model is retrained using two newly added votes and high-quality report samples confirmed by manual review, improving the accuracy of conflict detection. Quarterly, the large language model is incrementally fine-tuned based on the negative sample dataset and newly added high-quality samples to reduce the recurrence probability of similar errors. The application effects of novel embedding models and the large language model in this scenario are tested regularly, and model version upgrades are completed. For retrieval errors marked by manual review, the document segmentation and vector index of the knowledge base are corrected simultaneously, and the knowledge base version is updated.

[0053] S6.3 System Expansion and Adaptation The system adopts a plug-in architecture design, supporting the rapid access of new data sources, including weather API interfaces, drone inspection records, and real-time operation data of new energy power stations. It can incorporate weather factors and equipment on-site status into the risk assessment system. It supports the dynamic loading and updating of industry rules and enterprise specifications. The operation specifications of new equipment and the maintenance procedures for grid connection of new energy can be quickly implemented through plug-ins without modifying the core system code. It provides open API interfaces and custom rule configuration interfaces, allowing subordinate power stations to customize risk warning rules, analysis processes, and report templates according to their own business needs, adapting to the personalized management and control needs of different voltage levels and different types of power stations.

[0054] This invention provides a two-ticket analysis method based on graph databases and large language models. For the first time, it constructs power system operation tickets and work tickets into a dynamic graph database model, realizing multi-dimensional relational expressions of station components over time, breaking through the limitations of traditional relational databases. Simultaneously, it embeds dynamic risk weights to achieve real-time risk quantification of multi-hop paths, solving the problem of risk assessment for cross-ticket operations. This patent establishes a power industry-specific two-layer retrieval enhancement generation architecture, combining voltage level and equipment type multi-dimensional vector indexes to achieve accurate retrieval of professional content, significantly improving the compliance and accuracy of the analyzed content. This patent pioneers a master-slave multi-agent collaborative architecture, completing the automated decomposition and execution of complex analysis tasks and risk simulation verification through closed-loop collaboration of multiple types of agents, achieving automated control of the entire two-ticket analysis process. This patent develops a power industry-specific fine-tuning framework for general large language models, adapting the model to power industry scenarios through instruction fine-tuning, while designing a dynamic template engine to meet the report generation needs of different risk levels, effectively avoiding the illusion problem of large model generation. This patent also constructs a closed-loop self-optimizing architecture for the entire process of analysis, execution, and feedback. It achieves continuous iteration and upgrading of the model and knowledge base through negative sample mining, and supports plug-in expansion, which can fully adapt to various business needs of intelligent management and control of two tickets under the new power system.

[0055] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A two-vote analysis method based on graph databases and large language models, characterized in that, Includes the following steps: S1. Data preprocessing: Obtain the original data of the two tickets of the power system, clean and standardize the original data of the two tickets, and integrate multi-source related data such as power equipment ledger, historical fault records, and personnel scheduling data to obtain a structured two-ticket dataset; S2. Graph Database Modeling and Association Analysis: Based on the structured two-ticket dataset, construct a two-ticket association graph model in the graph database, which includes entities such as two tickets, equipment, personnel, and time intervals. Utilize the multi-hop query capability of the graph database to perform two-ticket association conflict detection and risk quantification analysis, and output the graph analysis results. S3. Retrieval Enhancement Generation Processing: Construct a vectorized knowledge base specifically for the power industry, generate retrieval instructions based on the graph analysis results, and retrieve matching power industry standards and case data from the vectorized knowledge base to obtain retrieval enhancement support data; S4. Multi-Agent Collaborative Task Processing: Through a preset multi-agent collaborative framework, the task of generating two analysis reports is broken down into multiple executable sub-tasks with priorities. Corresponding functional tools are scheduled to complete the execution and result verification of each sub-task, and integrated to obtain structured analysis intermediate results. S5. Intelligent Analysis Report Generation: The graph analysis results, retrieval enhancement support data, and structured analysis intermediate results are integrated into standardized prompt words, which are then input into a large language model adapted and fine-tuned for the power industry to generate two intelligent analysis reports that conform to power industry standards. S6. Closed-loop iterative optimization: Collect manual review feedback data from the two-vote intelligent analysis reports, and based on the feedback data, complete the vectorized knowledge base update, incremental fine-tuning of the large language model, and optimization of the two-vote association graph model to achieve closed-loop self-optimization of the system.

2. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S2, the two-ticket association graph model is configured with corresponding attribute parameters for nodes and edges, with nodes corresponding to entities and directed edges corresponding to the association relationships between entities. The attribute parameters of the nodes include ticket status, equipment risk level, and personnel qualification level. The attribute parameters of the edges include dynamic risk weight, operation order, and association timeliness. The dynamic risk weight includes at least equipment load rate and personnel workload coefficient.

3. The two-vote analysis method based on graph database and large language model according to claim 2, characterized in that, In step S2, the two-ticket association conflict detection includes at least the following: conflict detection of multiple overlapping operations on the same device at the same time, conflict detection of multiple high-risk tickets from the same person at the same time, and detection of excessive concurrent operation load on the device. The time overlapping operation conflict detection is achieved by calculating the time conflict probability P, and the calculation formula is: P = overlap duration / total operation duration. When P is greater than the preset conflict threshold, a high-risk warning is triggered.

4. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S3, the power industry-specific vectorized knowledge base adopts a two-layer architecture, including a structured procedure knowledge base and an unstructured fault case base. The vectorized knowledge base is configured with a multi-dimensional vector index based on voltage level and equipment type. During retrieval, the cosine similarity between the query vector and the knowledge base document vector is calculated to return the Top-K target document data with the highest matching degree.

5. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S4, the multi-Agent collaborative framework adopts a master-slave agent group architecture, which includes at least a scheduling agent, a verification agent, and a generation agent. The scheduling agent is responsible for task decomposition, priority allocation, and tool scheduling; the verification agent is responsible for logical consistency verification and risk simulation verification of sub-task results; and the generation agent is responsible for the structured integration of multi-source results. The three types of agents form a closed-loop collaborative execution link.

6. The two-vote analysis method based on graph database and large language model according to claim 5, characterized in that, The verification agent uses Monte Carlo simulation to predict the failure probability of the ticket operation sequence, and also supports counterfactual reasoning verification to achieve accurate quantification of risk level.

7. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S5, the large language model, after being fine-tuned for the power industry, completes the domain adaptation by fine-tuning the instructions of power industry dispatching terminology, safety procedures, and operating specifications. At the same time, a dynamic template engine is configured to automatically match the corresponding report template according to the risk level of the ticket. The large language model adopts the thinking chain technology to transform the topology analysis results into natural language descriptions that conform to the power industry standards.

8. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S6, the collection of feedback data is manually reviewed. The negative sample mining algorithm automatically identifies the manual modification traces in the report, extracts the erroneous content and labels the error type to form a negative sample dataset. Based on the negative sample dataset, the incremental fine-tuning of the large language model, the correction of the vectorized knowledge base and the optimization of the graph query algorithm are completed periodically.

9. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S1, the standardization process includes at least unifying equipment naming rules, filling in missing timestamps, and removing duplicate ticket records; the multi-source associated data and the original data of the two tickets are associated and mapped through the unique equipment number, employee number, and unique ticket identifier to generate a structured two-ticket dataset with a unique identifier.

10. The two-vote analysis method based on graph database and large language model according to claim 1, characterized in that, In step S2, when two new tickets are added, the nodes, edges and corresponding attribute parameters of the two-ticket association graph model are updated in real time, and the full-scale conflict detection process is automatically triggered; at the same time, the graph index is optimized regularly, and graph data in high-frequency query scenarios are cached to improve query response efficiency.

Citation Information

Patent Citations

  • Two-ticket management system based on artificial intelligence

    CN116091030A

  • Two-ticket analysis method and device based on power distribution network, equipment and medium

    CN117112848A