A power supply bureau intelligent document question and answer optimization method and system based on semantic analysis
By identifying and quantifying feature words in unstructured fault reports of power systems through semantic analysis, the problem of traditional systems being unable to understand complex fault logic and physical dependencies is solved, enabling more accurate fault diagnosis and knowledge management, and ensuring the safety and stability of power systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
- Filing Date
- 2026-05-06
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional intelligent document question-and-answer systems are unable to understand the logical and physical dependencies of complex faults when processing unstructured fault reports from power systems, and are unable to parse descriptive features and non-standard terms, leading to a decline in the accuracy and reliability of fault diagnosis.
A semantic analysis-based approach is used to identify target objects, operating parameters, and non-standard terms in query requests to obtain physical operating characteristic information, quantify descriptive feature words, and perform causal logic matching in a knowledge association network to output diagnostic results.
It improves the accuracy and reliability of intelligent document question-and-answer systems in power fault diagnosis, enabling them to parse descriptive fault reports, map non-standard terms, determine fault evolution paths, and enhance the intelligent transformation of knowledge management and the safe and stable operation of power systems.
Smart Images

Figure CN122132543A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent document processing technology in the power industry, and more specifically, to an intelligent document question-and-answer optimization method and system for power supply bureaus based on semantic analysis. Background Technology
[0002] Against the backdrop of the intelligent transformation of power systems, power supply bureaus are facing unprecedented technical challenges in their operation and maintenance work. Traditional intelligent document question-and-answer systems perform well in processing standardized fault code queries, but they have significant shortcomings when faced with descriptive fault reports generated by new intelligent sensors. The keyword matching mechanisms used by these systems cannot effectively parse semantic expressions containing dynamic changes, such as "increased voltage fluctuations" and "abnormal temperature gradients," and they have even less difficulty in identifying time-dimensional fault characteristics such as "instantaneous" and "intermittent."
[0003] As the intelligence level of power distribution network equipment increases, the fault phenomena recorded by maintenance personnel show a clear trend towards natural language. This is typically manifested in: the extensive use of non-standard terms such as "gradually deteriorating" and "intermittent" in descriptions of abnormal equipment conditions; fault characteristic descriptions often including complex time conditions such as "parameters deviating from baseline values within a certain period"; and the presence of descriptions implying causal relationships, such as "abnormal transformer oil temperature accompanied by increased line harmonics," when multiple devices experience coordinated faults. Existing systems cannot establish a quantitative correlation between these descriptions and the physical characteristics of power equipment, leading to serious deviations in semantic understanding.
[0004] A more prominent problem arises in the analysis of complex fault scenarios. When queries involve multiple factors such as "misoperation of protection devices under high temperature conditions," traditional systems can only return isolated condition matching results, failing to identify the physical relationship between ambient temperature, equipment load, and protection logic. This deficiency prevents the system from reconstructing the actual fault evolution path, resulting in the omission of critical historical cases. Furthermore, the system lacks the ability to dynamically map dialect terms commonly used by maintenance personnel (such as "ignition" corresponding to "arc discharge"), further reducing query accuracy.
[0005] The existing technical architecture suffers from three fundamental limitations: First, the semantic parsing layer lacks the ability to transform descriptive features into quantitative indicators of equipment operating parameters; second, the terminology mapping mechanism does not consider the corrective effect of real-time operating status of power equipment on semantic understanding; and finally, the knowledge retrieval module lacks multi-dimensional correlation analysis capabilities based on physical causal chains. These problems lead to a significant decline in the system's response quality when faced with unstructured queries generated by modern power systems, severely limiting the application value of intelligent question-answering systems in the field of fault diagnosis.
[0006] There is currently no effective technical solution to the above problems. Summary of the Invention
[0007] The purpose of this invention is to provide a semantic analysis-based intelligent document question-and-answer optimization method and system for power supply bureaus. It aims to solve the problem that traditional systems cannot understand the complex logical and physical dependencies between different devices and events when facing complex faults, nor can they construct a complete fault evolution path from scattered text descriptions. This invention enhances the intelligent transformation of knowledge management in power supply enterprises, improves the accuracy of question-and-answer in scenarios such as distribution network operation specification query and equipment parameter analysis, and ultimately ensures the safe and stable operation of the power system.
[0008] In a first aspect, the present invention provides a method for optimizing intelligent document question-and-answering for power supply bureaus based on semantic analysis, comprising the following steps: S1. Parse the input query request and identify the target object, running parameters, descriptive feature words, and non-standard terms in the query request; S2. Obtain the physical operation feature information corresponding to the target object, and quantize the descriptive feature words according to the physical operation feature information to obtain quantized semantic features; S3. Based on the context information of the query request, map the non-standard terms to standard technical entities; S4. Based on the quantized semantic features and the standard technical entities, determine the fault evolution path between the target object and the operating parameters, and output the diagnostic results.
[0009] The intelligent document question-and-answer optimization method for power supply bureaus based on semantic analysis provided by this invention can effectively parse descriptive fault reports, map non-standard terms to standard technical entities, quantify semantic features, and determine the fault evolution path through causal logic matching, thereby improving the accuracy and reliability of the intelligent document question-and-answer system in power fault diagnosis.
[0010] Secondly, this invention provides a power supply bureau intelligent document question-answering optimization system based on semantic analysis, comprising: The identification module is used to parse the input query request and identify the target object, running parameters, descriptive feature words and non-standard terms in the query request; The processing module is used to obtain the physical operation feature information corresponding to the target object, and to quantize the descriptive feature words according to the physical operation feature information to obtain quantized semantic features; The mapping module is used to map the non-standard terms to standard technical entities based on the context information of the query request; The output module is used to determine the fault evolution path between the target object and the operating parameters based on the quantized semantic features and the standard technical entities, and to output the diagnostic results.
[0011] As can be seen from the above, the intelligent document question-and-answer optimization method for power supply bureaus based on semantic analysis provided by the present invention can accurately determine the fault evolution path by parsing query requests, quantifying descriptive feature words, mapping non-standard terms and performing causal logic matching. It has the ability to effectively parse descriptive fault reports, map non-standard terms to standard technical entities, quantify semantic features, and determine the fault evolution path through causal logic matching, thereby improving the accuracy and reliability of the intelligent document question-and-answer system in power fault diagnosis.
[0012] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0013] Figure 1 This is a flowchart of a semantic analysis-based intelligent document question-answering optimization method for power supply bureaus, provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of a semantic analysis-based intelligent document question-answering optimization system for power supply bureaus, provided in an embodiment of the present invention.
[0015] Label Explanation: 100. Identification module; 200. Processing module; 300. Mapping module; 400. Output module. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0017] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0018] In traditional intelligent document question-and-answer systems for power supply bureaus, the system struggles to effectively understand deep semantics and perform causal logic matching when processing queries containing non-standard terminology, descriptive keywords, and multi-device event associations. Specifically, the system relies on keyword matching mechanisms, which fail to capture the time-series changes, trend characteristics, and physical causal relationships between cross-device events described in the query, resulting in relevance and incompleteness in the search results. Descriptive keywords such as "instantaneous" and "violent fluctuations" cannot be quantified, non-standard terminology is ambiguous due to a lack of unified standards, and the system cannot logically connect equipment operating parameters in discrete records with fault phenomena, thus affecting the accuracy of fault diagnosis and system reliability.
[0019] For example, during the intelligent operation of a distribution network, a key feeder in a substation in a certain area frequently experienced instantaneous tripping under specific operating conditions. The reports submitted by maintenance personnel detailed the environmental conditions, load parameters, and protection device behavior at the time of the tripping, including descriptive information such as high ambient temperature, a sudden increase in feeder load current before tripping, and the activation of a specific type of relay protection device without any obvious external fault signal. When queries involved the inherent relationships between multiple conditions, the system could only recognize isolated terms and could not understand the deep semantic connections between "high ambient temperature," "high load rate," and "protection device malfunction." This resulted in search results containing a large number of irrelevant documents, such as cases about the effects of normal temperature or single load tripping, while historical records truly reflecting multi-factor coupled faults were omitted. Consequently, the system could not provide a complete fault evolution path, making it difficult for the maintenance team to locate the root cause, forcing them to take temporary operational measures, affecting power supply continuity and equipment maintenance efficiency.
[0020] If the aforementioned issues are not addressed, the intelligent document question-and-answer system will be unable to activate the deep information accumulated in the fault case database, rendering historical data unusable due to insufficient semantic understanding. Maintenance personnel's trust in the system will continue to decline, leading them to rely more on personal experience or manual document review when facing complex faults, rather than relying on system-assisted decision-making. Furthermore, this limitation results in decreased knowledge management efficiency, the system's inability to integrate multi-dimensional information to form effective diagnostic recommendations, and the potential to increase operational risks by providing incomplete information, ultimately hindering the ability to ensure the safe and stable operation of the power system.
[0021] For reference, see the appendix. Figure 1 This invention provides a semantic analysis-based intelligent document question-answering optimization method for power supply bureaus, comprising the following steps: S1. Parse the input query request and identify the target object, runtime parameters, descriptive keywords, and non-standard terms in the query request; S2. Obtain the physical operation feature information corresponding to the target object, and quantify the descriptive feature words according to the physical operation feature information to obtain quantified semantic features; the physical operation feature information includes the type information of the target object, the physical threshold information of the monitoring parameters, and historical operation statistics. S3. Based on the context information of the query request, map non-standard terms to standard technical entities; S4. Based on quantified semantic features and standard technical entities, perform causal logic matching in a pre-defined knowledge association network to determine the fault evolution path between the target object and the operating parameters, and output the diagnostic results.
[0022] This application proposes a semantic analysis-based optimization method for intelligent document question answering in power supply bureaus. This method aims to enhance the deep semantic understanding and correlation matching capabilities of intelligent document question answering systems in power supply bureaus for "soft fault" reports in power operation and maintenance scenarios, particularly for natural language queries containing time-series features, trend descriptions, and causal chains across multiple devices and events. The core lies in enabling the system to "sensitize" and "quantify" these descriptions based on specific contexts and to understand non-standard terminology through a context-sensitive dynamic semantic dictionary and quantification rules.
[0023] In step S1, the input query request is parsed to identify the target object, operating parameters, descriptive keywords, and non-standard terms in the query request. When the power supply bureau's maintenance personnel ask a question to the intelligent question-and-answer system, the system first activates an intelligent query parsing module (this module can be specifically a text analysis unit based on Natural Language Processing (NLP) technology, such as using a combination of rule-based and statistical methods). This module analyzes each word in the query, especially those describing time and trends (such as "instantaneous," "violent fluctuations," and "slow rise") and those non-standard, colloquial equipment names or phenomenon descriptions (such as "square box" and "malfunction"). Through a preset lexical analyzer and named entity recognizer, this module identifies these words as the basic logical elements constituting the fault phenomenon, such as equipment name, parameter type, time modifiers, and trend descriptive words. For example, for a query of "the number of instantaneous voltage drops on a certain line has increased abnormally", this module will identify "line" as the equipment type, "voltage" as the parameter, "instantaneous" as the time characteristic, "drop" as the event, and "the number of drops has increased abnormally" as the trend description.
[0024] The lexical analyzer (also known as a tokenizer or lexical unit analyzer) is the first step in natural language processing. Its main function is to decompose the input continuous sequence of characters (i.e., the user's query text) into a series of smallest units with independent meaning; these units are called lexical units or tokens. For Chinese text, the lexical analyzer typically performs word segmentation, dividing the sentence into individual words. For example, when a user inputs "the number of times the instantaneous voltage drops on a certain line has increased abnormally," the lexical analyzer will decompose it into independent tokens such as "certain," "line," "instantaneous," "voltage," "drops," "number of times," "abnormal," and "increase." The lexical analyzer identifies these tokens using a pre-defined dictionary and rules (such as regular expressions) and removes irrelevant punctuation marks or spaces, providing a clean and structured token sequence for subsequent semantic analysis.
[0025] Named Entity Recognizers (NERs) operate based on the lexical sequence output by a lexical analyzer. The main function of a NER is to identify entities with specific meanings in text and categorize them into predefined categories. In the power industry, these predefined categories can include equipment names (e.g., "transformer," "switchgear"), parameter types (e.g., "voltage," "current," "temperature"), time modifiers (e.g., "instantaneous," "within a specific time period"), trend descriptions (e.g., "drastic fluctuations," "abnormal increases," "slow rise"), and fault phenomena (e.g., "tripping," "insulation breakdown"). NERs are typically trained and used with rule-based methods, statistical models (e.g., Conditional Random Fields, CRFs), or deep learning models (e.g., Recurrent Neural Networks, Transformers). For example, after receiving word elements such as “line”, “instantaneous”, “voltage”, “drop”, and “abnormal increase in frequency”, the named entity recognizer will identify that “line” is an entity of type “device name”, “voltage” is an entity of type “parameter”, “instantaneous” is an entity of type “time modifier”, “drop” is an entity of type “event”, and “abnormal increase in frequency” is an entity of type “trend description”.
[0026] Lexical analyzers and named entity recognizers are basic functional modules in natural language processing technology. These are existing technologies and will not be elaborated upon here.
[0027] In step S2, the physical operation feature information corresponding to the target object is obtained, and the descriptive feature words are quantified based on the physical operation feature information to obtain quantified semantic features. The physical operation feature information includes the type information of the target object, the physical threshold information of the monitoring parameters, and historical operation statistics. The system maintains a dynamic semantic dictionary, which dynamically adjusts the understanding of descriptive words based on the specific type of power equipment mentioned in the current query (e.g., transformer, line, or relay protection device), relevant monitoring parameters (e.g., voltage, current, temperature), and the normal operating range and physical limit values of these parameters defined in the power equipment ontology knowledge. For example, when the word "instantaneous" appears in the query, if the intelligent query parsing module identifies that it refers to the action time of the relay protection device, the system will query the quantization rule corresponding to "relay protection device - instantaneous" from the rule base and quantize it into a time scale of less than 100 milliseconds; but if it refers to the change in transformer oil temperature, the system will quantize it into a time scale of less than 5 minutes. Similarly, for terms like "drastic fluctuations" or "abnormal increases," the system combines the physical thresholds of specific parameters with historical operating data. It calculates the rate of change or standard deviation of the parameter from the average within a specific time window and compares this calculation with preset thresholds to determine the degree of deviation from the normal range, thus providing a more precise quantitative explanation. For example, for "drastic fluctuations in the intensity of partial discharge signal in the transformer," the system calculates the standard deviation of the discharge signal over the past hour. If this standard deviation exceeds three times the standard deviation of the transformer model under normal operating conditions, it is classified as "drastic fluctuation."
[0028] In step S3, non-standard terms are mapped to standard technical entities based on the context information of the query request. To address the issue of maintenance personnel using non-standard terms, the system maintains a dedicated "domain expert consensus mapping table." This table continuously collects and analyzes the associations between non-standard expressions (such as "little black box" or "crazy") used by maintenance personnel in daily communication and fault reports, and their corresponding standard device names, fault modes, or operational behaviors in the knowledge system. When the intelligent query parsing module identifies a non-standard term, the system prioritizes searching for the most likely standardized counterpart in this mapping table. The search process combines contextual information such as device type and parameters identified in the current query for filtering and sorting. For example, if the query mentions "that square box," and the intelligent query parsing module identifies that the context of the current conversation is related to a new type of smart switch (e.g., model "smart switch-x"), the system will associate it with that model of smart switch based on the records in the mapping table (e.g., "square box" maps to "smart switch-x," the context is "smart switch," and the priority is high). This mapping table also dynamically adjusts the priority or confidence level of different mapping relationships based on feedback from operations and maintenance personnel (e.g., confirmation or correction of the system's recommended mapping results) and usage frequency, ensuring that the system can always find the most accurate interpretation.
[0029] In step S4, based on quantified semantic features and standard technical entities, causal logic matching is performed in a pre-defined knowledge association network to determine the fault evolution path between the target object and operating parameters, and the diagnostic results are output. With the help of a dynamic semantic dictionary and an expert consensus mapping table, the system can transform natural language queries from maintenance personnel, even "soft fault" reports containing vague descriptions and non-standard terminology, into more precise and physically meaningful structured information. The system no longer simply matches keywords but can extract the dynamic changes in equipment status, time-series features, and potential causal clues contained in the query. For example, for the query "the transformer partial discharge signal intensity fluctuates drastically," the system can understand that "drastic fluctuation" is not just "fluctuation," but refers to a specific intensity of trend that may indicate insulation degradation. This understanding is achieved by comparing the quantified trend information with the fault precursor patterns defined in the power equipment ontology knowledge. Based on this precisely understood semantic information, the system matches historical cases and physical laws in the fault case library and power equipment ontology knowledge that are highly consistent logically and causally. This can be achieved through a graph database-based inference engine. The system uses the semantically processed elements (equipment, parameters, events, trends) in the query as starting nodes or constraints for the graph query, and performs path searches within the power equipment ontology knowledge graph. For example, if the query elements include "line," "voltage drop," and "relay protection action," the system will search the graph for a causal path from "line fault" to "voltage drop" and then to "relay protection action." It attempts to connect the various elements in the query to form one or more possible fault causal chains and provides corresponding diagnostic paths, rather than simply returning isolated document fragments.
[0030] The core technical concept of this solution lies in introducing a context-sensitive dynamic semantic dictionary and a domain expert consensus mapping table. This enables the intelligent question-answering system to transcend traditional keyword matching and deeply understand the descriptive language and non-standard terminology with time and trend characteristics in "soft fault" reports in power operation and maintenance scenarios. It allows the system to, like an experienced power engineer, dynamically quantify and contextualize these ambiguous expressions based on specific equipment types and operating parameters. It transforms informal, colloquial knowledge into standardized information that the system can understand and reason about, thereby achieving deep semantic understanding and causal correlation matching of complex fault modes, ultimately providing a comprehensive and accurate fault diagnosis path.
[0031] The following example will provide a more detailed explanation of the above technical solution: Suppose that in a power supply bureau, maintenance personnel A discovers that a critical feeder at a regional substation is frequently experiencing unexplained momentary tripping. Over several weeks, personnel A submits multiple descriptive reports containing information on ambient temperature, load current curves, and relay protection operation data. These reports detail specific parameters at the time of each trip, such as "ambient temperature exceeding 40 degrees Celsius," "feeder load current instantaneously reaching 85% of its rated value before tripping," and "a specific type of relay protection device operating without any obvious external fault signal." Personnel A attempts to use an intelligent question-and-answer system to search for "historical cases where an ambient temperature exceeding 40 degrees Celsius and a load rate higher than 80% caused a specific type of protection device to malfunction."
[0032] First, in step S1, the system parses the query request from maintenance personnel A. The intelligent query parsing module identifies "ambient temperature," "load rate," and "protection device" as target objects, "above 40 degrees," "above 80 percent," and "malfunction" as descriptive features, and "specific model protection device" as a non-standard term.
[0033] Next, in step S2, the system acquires the physical operating characteristic information corresponding to these target objects. For example, for "ambient temperature," the system acquires its normal operating range and historical statistical data; for "load rate," the system acquires its physical threshold information and historical operating statistical data. Then, the system quantifies the descriptive feature words based on this physical operating characteristic information. For example, "above 40 degrees" is quantified into a specific temperature value (e.g., if the real-time monitored temperature is 40.5 degrees Celsius, then the quantified specific temperature value is 40.5), "above 80 percent" is quantified into a specific load current value (e.g., if the rated load is 100A and the real-time load is 85A, then the quantified specific load current value is 85A), and "maloperation" is quantified according to the action time threshold of the relay protection device (e.g., if the actual action time far exceeds the normal action time threshold, it may be quantified as a numerical state of "abnormal action time delay X milliseconds").
[0034] Subsequently, in step S3, the system maps the non-standard term "specific model protection device" to a standard technical entity based on the context information of the query request. The system queries the "domain expert consensus mapping table" and, combined with the context information such as "protection device" already identified in the query, maps "specific model protection device" to the corresponding standard relay protection device entity in the knowledge base, such as "relay protection device model X".
[0035] Finally, in step S4, the system performs causal logic matching in a pre-defined knowledge association network based on quantified semantic features and standard technical entities (such as "relay protection device model X"). The system uses this information as query input to the knowledge graph and performs causal logic matching through a pre-established semantic knowledge graph in the power field. This knowledge graph defines power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them. The system searches the graph for the causal path from "abnormal ambient temperature" to "overload" and then to "relay protection device model X malfunction". In this way, the system can determine the fault evolution path between "ambient temperature exceeding 40 degrees and load rate exceeding 80%" and "relay protection device model X malfunction" and output diagnostic results, such as a similar fault record from many years ago, caused by insufficient thermal stability of a certain component inside a specific batch of protection devices, which describes in detail the mechanism by which the performance of the component deteriorates under high temperature and high load, leading to misjudgment of the protection logic.
[0036] As can be seen from the above examples, the method of this application intelligently parses query requests, quantifies descriptive feature words, and uses contextual information to map non-standard terms to standard technical entities. Finally, it performs causal logic matching in a knowledge association network, thereby accurately identifying the deep semantics and causal relationships of complex faults. Unlike traditional question-answering systems that rely solely on keyword matching, the method of this application can connect three seemingly independent factors—"ambient temperature," "load rate," and "protection device malfunction"—through deep semantic association, finding a similar fault record from many years ago caused by insufficient thermal stability of a component inside a specific batch of protection devices. This capability enables the system to provide a comprehensive fault diagnosis path, solving the problem that traditional systems, when faced with complex faults, cannot understand the complex logical and physical dependencies between different devices and events, nor can they construct a complete fault evolution path from scattered text descriptions. The method in this application can effectively handle "soft fault" reports with multiple variables and strong descriptive features, as well as multi-causal chain fault queries, thereby enhancing the intelligent transformation of knowledge management in power supply enterprises, improving the accuracy of question and answer in scenarios such as distribution network operation specification queries and equipment parameter analysis, and ultimately ensuring the safe and stable operation of the power system.
[0037] In some embodiments, step S2, which involves quantizing descriptive feature words based on physical operation feature information to obtain quantized semantic features, includes: S21. Based on physical operation characteristic information, calculate the rate of change of the parameters corresponding to the descriptive feature words within a preset time window or the degree of deviation from the preset standard value; S22. By comparing the rate of change or degree of deviation with their respective preset thresholds, the descriptive feature words are quantified to obtain quantified semantic features.
[0038] Step S21 aims to transform the vague descriptive features in the query request into quantifiable dynamic indicators. In practice, the rate of change can be calculated by obtaining the initial and final values of a parameter within a preset time window (e.g., the past hour, the past 24 hours, or the past 7 days), and then calculating the ratio of its relative change to the initial value, or by calculating the average change per unit time. For example, for the description "slow voltage increase," the system can calculate the average rate of voltage increase over the past hour. The degree of deviation can be calculated by comparing the current value of the parameter with a preset standard value (e.g., the rated operating value of the equipment, historical average value, or safety threshold), calculating its absolute difference or relative percentage deviation. For example, for the description "abnormal temperature increase," the system can calculate the percentage deviation of the current temperature from the upper limit of the equipment's normal operating temperature. The physical operating characteristic information includes the type information of the target object, the physical threshold information of the monitored parameters, and historical operating statistics. This information provides the necessary context and benchmark for calculating the rate of change and the degree of deviation. For example, when calculating the voltage change rate, it is necessary to know the type of the target object (e.g., transformer, line) in order to obtain the historical operating statistics and normal fluctuation range of its corresponding voltage parameters.
[0039] Step S22 aims to standardize and normalize the dynamic indicators calculated in S21, transforming them into quantitative semantic features with clear physical meaning. Specifically, the calculated rate of change or degree of deviation can be compared with a series of tiered thresholds, thereby quantifying descriptive features into different levels or states. For example, a voltage deviation of less than 5% is considered "slight deviation," 5% to 10% is "moderate deviation," and greater than 10% is "severe deviation." Alternatively, the calculated rate of change or degree of deviation can be compared with a single preset threshold to determine whether it exceeds the normal range, generating Boolean (yes / no) or binary (normal / abnormal) quantitative semantic features accordingly. For example, if the temperature change rate exceeds a certain threshold, it is quantified as "rapid temperature change." These preset thresholds can be set based on the operating procedures of power equipment, industry standards, expert experience, or historical fault data analysis results. For example, for the term "instantaneous," if it refers to the action time of a relay protection device, the preset threshold can be set to 100 milliseconds; if it refers to the change in transformer oil temperature, the preset threshold can be set to 5 minutes. For terms such as "drastic fluctuations" or "abnormal increases", the threshold can be set based on the multiple of the standard deviation of the parameter within a specific time window, such as exceeding 3 times the normal standard deviation.
[0040] This method, through the aforementioned quantification process, transforms the descriptive feature words contained in the query request into quantified semantic features with clear physical meaning. These quantified semantic features, along with the target object identified in the preceding steps, the operating parameters, and the subsequently mapped standard technical entities, serve as inputs for causal logic matching within a pre-defined knowledge association network. This process first calculates the rate of change or deviation from a pre-defined standard value of the parameters corresponding to the descriptive feature words within a pre-defined time window based on the physical operating characteristic information. It utilizes the type information of the target object, the physical threshold information of the monitoring parameters, and historical operating statistics as a foundation, ensuring that the understanding of descriptive words incorporates the actual operating status and historical performance of the equipment. Subsequently, the descriptive feature words are quantified by comparing the calculated rate of change or deviation with pre-defined thresholds, thereby obtaining the quantified semantic features. These pre-defined thresholds are set based on the operating procedures of power equipment, industry standards, or expert experience, ensuring the objectivity and consistency of the quantification results. This precise quantification process enables the system to more accurately understand query intent, capture time-series anomalies and trend deviations in equipment operation, and thus perform deeper causal reasoning within the knowledge association network. This allows the system to determine the fault evolution path between the target object and the operating parameters, and output more comprehensive and accurate diagnostic results. This significantly improves the system's deep semantic understanding and correlation matching capabilities for "soft fault" reports, solving the problem of inaccurate diagnostic results or omission of key information in traditional systems due to their inability to understand deep semantics.
[0041] As a specific implementation, when a query request includes "severe fluctuations in transformer partial discharge signal strength," the system first identifies "transformer" as the target object, "partial discharge signal strength" as the operating parameter, and "severe fluctuations" as the descriptive term. In step S21, the system calculates the rate of change of the "partial discharge signal strength" parameter within a preset time window or the degree of deviation from a preset standard value based on the physical operating characteristic information. Specifically, the system can obtain the average value and standard deviation of the "partial discharge signal strength" of this type of transformer under normal operating conditions from historical operating statistics. Simultaneously, the system obtains the real-time data of the "partial discharge signal strength" involved in the current query request over the past hour. Then, the system calculates the standard deviation of this parameter over the past hour. In step S22, the system quantifies "severe fluctuations" by comparing the calculated standard deviation with a preset threshold. For example, the system can preset a rule that if the standard deviation of the "partial discharge signal strength" over the past hour exceeds three times the standard deviation of this type of transformer under normal operating conditions, then "severe fluctuations" is quantified as "severe fluctuation anomaly." If the standard deviation exceeds 1 but does not exceed 3, it is quantified as "moderate abnormal fluctuation". If it does not exceed 1, it is quantified as "normal fluctuation". These quantification rules can be stored in a configurable rule base, which is stored in key-value pairs. The key is a combination of descriptive terms and device type and parameters, and the value is the corresponding quantification threshold or calculation method. In this way, the system transforms the vague natural language description of "violent fluctuation" into the specific quantitative semantic feature of "severe abnormal fluctuation", providing accurate input for subsequent fault diagnosis.
[0042] Through the above technical solution, this method can transform descriptive feature words contained in query requests, such as "instantaneous," "violent fluctuations," or "abnormal increase," into quantitative semantic features with clear physical meaning. This enables the system to dynamically capture the rate of change of the parameters corresponding to the descriptive feature words within a preset time window or the degree of deviation from a preset standard value, overcoming the problem of traditional methods lacking a specific mechanism to dynamically capture these subtle anomalies. By comparing the rate of change or the degree of deviation with a preset threshold, this method ensures the objectivity, reliability, and consistency of the quantification results. Therefore, the generated quantitative semantic features can more accurately reflect subtle anomalies in the device's operating status, providing more accurate and reliable input for subsequent causal logic matching in a preset knowledge association network. This significantly improves the accuracy of fault evolution path determination and the comprehensiveness of diagnostic results, effectively solving the problem of inaccurate diagnosis or omission of key information caused by insufficient semantic understanding when traditional systems process "soft fault" reports.
[0043] In some embodiments, the specific steps in step S22 include: S221. Obtain real-time and historical data of multiple operating parameters involved in the query request; S222. Identify the joint behavior of multiple operating parameters over a specific time series; S223. Based on the joint behavior and the preset parameter association rules, determine whether the joint behavior conforms to the preset combination pattern; S224. When the joint behavior conforms to the combination pattern, and the rate of change or deviation of any of the multiple operating parameters does not reach their respective preset thresholds, the descriptive feature words are quantified according to the combination pattern to obtain quantized semantic features. Otherwise, the descriptive feature words are quantified by comparing the rate of change or deviation with their respective preset thresholds to obtain quantized semantic features.
[0044] The acquisition of real-time data for multiple operating parameters involved in the query request aims to provide foundational data for subsequent joint behavior analysis. When dealing with complex faults, a single parameter is often insufficient to fully reflect the equipment status, requiring multi-dimensional data support. This step can be achieved by interfaceing with real-time data platforms such as Supervisory Control and Data Acquisition (SCADA), Energy Management System (EMS), or Data Acquisition and Monitoring System (DCS), using API calls or data subscriptions to obtain real-time monitoring data related to the target object and operating parameters in the query request. Alternatively, data can be directly collected from sensor nodes and transmitted to a data processing center via a smart sensor network deployed on power equipment, where it is then filtered and extracted according to the query request.
[0045] Identifying the joint behavior of multiple operating parameters over a specific time series aims to capture the dynamic correlations and collaborative change patterns among these parameters, rather than analyzing each parameter in isolation. This is crucial for understanding the complex interactions between parameters in "soft faults." This identification process can employ time series analysis methods, such as Granger causality tests, cross-correlation analysis, or Dynamic Time Warping (DTW) algorithms, to detect the synchronicity, lag, or trend consistency of different parameter sequences over time. Furthermore, machine learning models, such as Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs), or attention mechanisms, can be used to model multivariate time series data, automatically learning and identifying joint behavior patterns among parameters.
[0046] Based on the joint behavior and predefined parameter association rules, it is determined whether the joint behavior conforms to a predefined combination pattern. This aims to utilize domain knowledge and experience to interpret and classify the identified joint behaviors, matching them with known fault or anomaly patterns. The predefined parameter association rules can be stored in a rule engine. These rules are defined by domain experts based on historical fault data and equipment operating mechanisms; for example, "when parameter A increases and parameter B decreases, it may indicate fault C." The combination pattern can be a specific combination of parameter changes described by these rules. Alternatively, data mining techniques, such as association rule learning (Apriori algorithm) or sequence pattern mining, can be used to automatically discover frequently occurring parameter joint behavior patterns from large amounts of historical operating data and use them as predefined combination patterns.
[0047] When the combined behavior conforms to the combination pattern, and the rate of change or deviation of any of the multiple operating parameters does not reach its respective preset threshold, the descriptive feature word is quantified according to the combination pattern to obtain the quantified semantic feature. Otherwise, the descriptive feature word is quantified by comparing the rate of change or deviation with its respective preset threshold to obtain the quantified semantic feature. This provides a flexible and intelligent quantification strategy, prioritizing the overall trend and combination pattern among parameters, and avoiding ignoring potential composite anomalies due to a single parameter not reaching its threshold. When a specific combination pattern is detected (e.g., "temperature rises slowly while vibration frequency increases slightly"), even if a single parameter (e.g., temperature change rate or vibration deviation) has not yet reached its respective independent alarm threshold, the system can still quantify the descriptive feature word (e.g., "unstable equipment operation") as "early warning" or "potential risk level 2" according to the preset quantification rules of the combination pattern. If the combination pattern is not met or a parameter has reached its independent threshold, the system reverts to traditional single-parameter threshold comparison for quantification. This process can be implemented using a decision tree or a rule-based expert system. The nodes of the decision tree can determine whether the joint behavior conforms to a combination pattern and whether a single parameter reaches a threshold. Based on the judgment result, different branches are selected for quantization. For example, if the combination pattern matches and no single parameter exceeds the limit, the combination pattern quantization logic is executed; otherwise, the single parameter quantization logic is executed.
[0048] This solution introduces the identification of the joint behavior of multiple operating parameters and the judgment of combination patterns, enabling the quantitative processing of descriptive feature words to go beyond the threshold comparison of a single parameter. Instead, it allows for a more macroscopic and relational understanding of the equipment's operating status. This is closely integrated with the steps of parsing query requests, acquiring physical operating characteristic information, and calculating the rate of change or degree of deviation, forming a more complete semantic analysis chain. In this way, even when a single parameter has not reached the traditional alarm threshold, the system can identify potential fault risks in advance based on the joint behavior and combination patterns between parameters. This allows the quantitative semantic features to more accurately reflect the true state of "soft faults," significantly improving the accuracy and foresight of subsequent causal logic matching and fault evolution path determination.
[0049] The following example illustrates this. Suppose an operations and maintenance (O&M) personnel query "Transformer operation is unstable, partial discharge signal fluctuates slightly, oil temperature rises slowly." First, the system acquires real-time data on multiple operating parameters related to the transformer, such as partial discharge signal strength, oil temperature, load current, and vibration frequency. This data can be transmitted in real-time from the transformer intelligent monitoring terminal to the data center via the MQTT protocol. Next, the system identifies the combined behavior of these parameters over the past hour. For example, by performing cross-correlation analysis on the time series of partial discharge signal strength and oil temperature, a positive correlation is found, meaning that as the oil temperature rises, the partial discharge signal strength also shows a slight upward trend. Simultaneously, analysis of the load current and vibration frequency reveals that they remain stable within the same time period, without significant changes. Then, based on the identified combined behavior and preset parameter association rules, the system determines whether it conforms to a preset combination pattern. For example, the system might have a rule: "When the transformer oil temperature rises slowly and the partial discharge signal fluctuates slightly, while the load current and vibration frequency are normal, it conforms to the combination pattern of 'early degradation of transformer insulation.'" In this case, the system determines that the current combined behavior conforms to this combination pattern. Finally, the system performs quantification processing. Assuming that the rate of change of partial discharge signal intensity and the deviation of oil temperature do not reach their respective preset alarm thresholds (e.g., the rate of change of partial discharge signal intensity does not exceed 5%, and the deviation of oil temperature from the normal value does not exceed 2°C), the system will quantify the descriptive term "unstable operation" based on the "early degradation of transformer insulation" combination pattern, for example, quantifying it as "early degradation of insulation risk level 1". If the combination pattern is not met, or the rate of change of partial discharge signal intensity exceeds 5%, the system will revert to quantification based solely on comparing the rate of change or deviation of a single parameter with its respective threshold.
[0050] Through the above technical solution, this application effectively solves the problem of incomplete quantification results caused by isolated analysis when traditional methods handle multi-parameter joint anomalies. By introducing the identification of the joint behavior of multiple operating parameters and the judgment of combination patterns, the system can capture the characteristics of "soft faults" where a single parameter has not yet reached the alarm threshold, but the coordinated changes of multiple parameters indicate potential risks. This makes the quantification of descriptive feature words more accurate and forward-looking, enabling earlier identification of subtle anomalies and potential fault trends in equipment operation. This solution significantly improves the semantic understanding capability of intelligent document question-and-answer systems for complex fault scenarios, providing more reliable and comprehensive quantitative semantic features for subsequent fault evolution path determination, thereby improving the accuracy and timeliness of fault diagnosis. This helps power supply bureau maintenance personnel to discover and handle potential problems earlier, ensuring the safe and stable operation of the power system.
[0051] In some embodiments, the specific steps in step S3 include: S31. Identify the core functional symptoms in the query request; S32. Generate a candidate list of standard technical entities for non-standard terms; S33. Obtain real-time operational anomaly data related to the core functional symptoms for each standard technical entity in the candidate list; use the real-time operational anomaly data as context information for the query request; S34. Based on real-time operational anomaly data, assess the abnormal correlation between each standard technical entity and the core functional symptoms, and select the standard technical entity with the highest abnormal correlation as the mapping result of the non-standard terminology, thereby mapping the non-standard terminology to the standard technical entity.
[0052] The identification of core functional symptoms in query requests aims to focus on key fault characteristics, avoid interference from non-core symptoms, and ensure the targeting of the mapping process. This step can utilize Natural Language Processing (NLP) techniques, such as deep learning-based text classification models or rule matching engines, to perform semantic analysis on the query requests and identify key phrases or words describing abnormal equipment states, fault phenomena, or performance degradation. These key phrases or words represent symptoms that users are concerned about and that directly point to functional problems with the equipment. Alternatively, a pre-defined symptom dictionary and grammatical rules can be used to extract verb and noun combinations directly related to power equipment faults from the query requests, such as "tripping," "overload," "abnormal temperature," and "data transmission interruption," and then combined with the context for preliminary screening to determine the symptoms that best reflect the core functional problems of the equipment.
[0053] Generating a candidate list of standard technical entities for non-standard terms aims to provide a set of potential mapping options, laying the foundation for subsequent evaluation. This step can be achieved by constructing a knowledge base or ontology containing technical entities such as the names, components, and failure modes of all standard equipment in the power sector. When a non-standard term is identified, the system uses methods such as fuzzy matching, word vector similarity calculation, or rule-based pattern matching to retrieve standard technical entities that are semantically similar to or may refer to the non-standard term from this knowledge base, forming a candidate set. Alternatively, a sequence labeling model can be trained using historical query logs and manually labeled data. This model can identify non-standard terms in query requests and generate a list of standard technical entities containing multiple possible mappings from a predefined standard technical entity library.
[0054] The process involves acquiring real-time operational anomaly data related to core functional symptoms for each standard technical entity in the candidate list, and using this data as context information for query requests. The aim is to dynamically associate symptoms with entities using real-time data, ensuring the mapping reflects the current operational status. This step can be achieved by interfaceing with multiple data sources, such as Supervisory Control and Data Acquisition (SCADA) systems, equipment condition monitoring systems, and historical databases, to query sensor data, alarm logs, and event records related to candidate standard technical entities in real time. After preprocessing, these data can be used to extract abnormal indicators related to core functional symptoms, such as packet loss rate, communication latency, and error codes. Alternatively, a real-time data stream service can be subscribed to to obtain the latest status of operational parameters related to candidate entities. For example, for a candidate smart meter, its recent meter reading data, communication status, and self-test reports can be obtained, and abnormal patterns related to specific symptoms can be identified.
[0055] Based on real-time operational anomaly data, the abnormal correlation between each standard technical entity and core functional symptoms is evaluated. The standard technical entity with the highest abnormal correlation is selected as the mapping result for non-standard terms, thereby mapping non-standard terms to standard technical entities. This aims to achieve precise selection through quantified correlation, improve mapping accuracy, and provide reliable input for fault diagnosis. This evaluation can be based on statistical methods, such as calculating the co-occurrence frequency, mutual information, or Pearson correlation coefficient between real-time operational anomaly data and core functional symptoms in historical fault cases. The entity with the highest abnormal correlation means that its real-time operational status shows the strongest association with the currently queried core functional symptoms. Alternatively, a machine learning-based classifier or regression model can be constructed. This model takes real-time operational anomaly data as input features and outputs a correlation score between each candidate standard technical entity and the core functional symptoms. The model can be trained on historical fault data to learn the association between different anomaly patterns and specific equipment faults.
[0056] This application's solution first identifies the core functional symptoms in the query request, focusing user attention on the most critical fault manifestations, thus providing clear semantic anchors for subsequent mapping of non-standard terms. Based on this, the system generates a candidate list of potential standard technical entities for non-standard terms, providing a comprehensive selection space for the mapping process. Subsequently, by acquiring real-time operational anomaly data related to these candidate entities and the core functional symptoms, this solution introduces dynamic, real-time contextual information, enabling mapping decisions to fully consider the current actual operating state of the equipment. Finally, based on this real-time anomaly data, the system quantitatively evaluates the anomaly correlation between each candidate entity and the core functional symptoms, and selects the entity with the highest correlation as the final mapping result. This mechanism ensures that the mapping of non-standard terms is no longer a static, simple matching based on preset rules, but a dynamic, intelligent reasoning based on real-time operating status. By accurately mapping non-standard terms to standard technical entities, this solution provides high-quality input for subsequent causal logic matching, significantly improving the accuracy of fault evolution path determination and the reliability of diagnostic results, thereby effectively solving the problem of inaccurate mapping caused by the lack of direct or dynamically updated contextual information related to core functional symptoms in traditional methods.
[0057] The following is a concrete example to illustrate this. Suppose the operations and maintenance personnel enter the query: "The data transmission of that box is abnormal."
[0058] First, the system identifies the core functional symptoms in the query request. By performing semantic analysis on "abnormal data transmission," it is determined to be the core functional symptom of this query.
[0059] Next, the system generates a candidate list of standard technical entities for the non-standard term "box". For example, the system identifies standard technical entities such as "smart meter", "remote terminal unit (RTU)" or "advanced protection relay" through a preset domain expert consensus mapping table or knowledge base, and includes them in the candidate list.
[0060] The system then retrieves real-time operational anomaly data related to the core functional symptoms for each standard technical entity in the candidate list. For example, for a "smart meter," the system queries its most recent meter reading data, communication status logs (such as disconnection records and signal strength), and self-test reports; for a "remote terminal unit (RTU)," the system retrieves its packet loss rate, communication latency, and error codes; and for an "advanced protection relay," the system checks its communication port status and error logs. This real-time data is used as contextual information for the query request.
[0061] Finally, based on this real-time operational anomaly data, the system assesses the correlation between each standard technical entity and the core functional symptom "data transmission anomaly." If the smart meter's communication status shows frequent disconnections, a high data packet loss rate, and a communication module fault alarm in its self-test report, while the communication status of the remote terminal unit (RTU) and advanced protection relays are normal, the system will determine that the smart meter has the highest correlation with "data transmission anomaly." Therefore, the system maps the "box" to the "smart meter" and uses this as the precise input for subsequent fault diagnosis.
[0062] Through the above technical solution, this application effectively solves the problem of inaccurate mapping of non-standard terms in traditional methods. By introducing the identification of core functional symptoms and dynamic evaluation of real-time operational anomaly data, this solution makes the mapping process of non-standard terms more targeted, real-time, and accurate. This not only avoids mapping deviations caused by insufficient or outdated contextual information but also ensures that the mapping results can truly reflect the current operating status and fault performance of the equipment. Therefore, this solution provides more accurate and reliable input for subsequent causal logic matching, significantly improving the accuracy of fault evolution path determination. This enables the power supply bureau's intelligent document question-and-answer system to provide more comprehensive and accurate diagnostic results when processing complex "soft fault" queries containing non-standard terms, greatly enhancing the system's practicality and decision support capabilities.
[0063] In some embodiments, the specific steps in step S31 include: S311. Extract all descriptive symptoms from the query request; S312. Obtain the pre-defined association rules and priority information of each descriptive symptom in the power equipment ontology knowledge; S313. Based on association rules, analyze the logical relationships between each descriptive symptom and identify whether there are conflicts or redundancies among each descriptive symptom to obtain the analysis and identification results; S314. Based on the analysis and identification results, and in conjunction with priority information, determine the core functional symptoms.
[0064] To ensure accurate identification of core functional symptoms, this application first extracts all descriptive symptoms from the query request in step S311. This step aims to comprehensively acquire all descriptive information related to faults or anomalies in the user query, providing a complete data foundation for subsequent analysis. Specifically, Natural Language Processing (NLP) techniques, such as rule-based pattern matching, statistical learning models (e.g., Conditional Random Fields, CRF), or deep learning models (e.g., BERT, LSTM), can be used to perform Named Entity Recognition (NER) to identify and extract all words or phrases describing device status, behavior, or phenomena from the query text. Furthermore, a pre-defined symptom dictionary and syntactic analysis can also be used to identify words in the query statement that function as predicates or objects and describe abnormal states.
[0065] Subsequently, in step S312, the pre-defined association rules and priority information for each descriptive symptom in the power equipment ontology knowledge are obtained. This step provides a structured knowledge framework and judgment basis for the logical analysis of the symptoms. Association rules define causal, parallel, and exclusionary relationships between different symptoms, while priority information is used for decision-making when symptoms conflict or are redundant. This information can be stored in a knowledge graph, where nodes represent entities such as symptoms, equipment, and parameters, edges represent the relationships between them, and attributes (such as priority and confidence level) are attached. Alternatively, this information can be stored in a relational database, using a predefined table structure to store symptom identifiers, associated symptom identifiers, association types (e.g., "mutually exclusive," "containment," "precursor"), priority values, etc.
[0066] Next, in step S313, the logical relationships between the descriptive symptoms are analyzed according to the association rules, and conflicts or redundancies among the descriptive symptoms are identified to obtain the analysis and identification results. This step aims to handle inconsistencies or duplicate information in the query that may exist, ensuring the accuracy of subsequent core symptom identification. By applying preset association rules, the system can determine the rationality of symptom combinations. In specific implementation, a rule reasoning engine can be used to take the extracted descriptive symptoms as input and make logical judgments by comparing them with the association rules in the ontology knowledge. For example, if the rule defines "overload" and "underload" as mutually exclusive symptoms, the system will mark them as conflicting when they occur simultaneously. Another approach is to use graph algorithms to perform path analysis or subgraph matching in the knowledge graph to identify symptom combinations that do not conform to the preset pattern or to discover semantically duplicated symptoms.
[0067] Finally, in step S314, the core functional symptoms are determined based on the analysis and identification results and in conjunction with priority information. This step, after handling conflicting and redundant symptoms, selects the key symptoms that best represent the user's query intent and the nature of the fault from the remaining valid symptoms. Priority information plays a decision-making support role in this process. Specifically, conflicting or redundant symptoms can be filtered out based on the analysis and identification results obtained in step S313. Then, for the remaining symptoms, they are sorted according to the priority information obtained in step S312, and the symptom with the highest priority is selected as the core functional symptom. For example, if "equipment overheating" and "insulation aging" coexist, and "insulation aging" has a higher priority (because it is usually a deeper cause), then "insulation aging" is selected as the core symptom. Furthermore, factors such as the syntactic position of the symptom in the query, word frequency, and semantic distance from the query target object can be combined using a weighted scoring model to comprehensively determine the core symptoms based on priority information.
[0068] The solution presented in this application ensures the accurate identification of core functional symptoms through the synergistic effect of the aforementioned steps. First, by comprehensively extracting all descriptive symptoms from the query request, information omissions are avoided, providing a complete and rich data foundation for subsequent analysis. Second, by acquiring pre-defined association rules and priority information from the power equipment ontology knowledge, a structured knowledge framework and judgment criteria are provided for symptom analysis, making the analysis process evidence-based. Based on this, the logical relationships between descriptive symptoms are analyzed according to association rules, and potential conflicts or redundancies are identified, effectively handling inconsistencies and eliminating the risk of misidentification. Finally, by combining the analysis and identification results with priority information, the system can comprehensively judge and accurately determine the core functional symptoms. This organic combination of steps enables the system to perform deep understanding and reasoning on complex natural language queries, much like an experienced engineer, thus providing a solid and reliable foundation for subsequently mapping non-standard terms to standard technical entities. By providing a verified and accurate core functional symptom, the accuracy and reliability of non-standard term mapping are greatly improved, thereby optimizing the diagnostic results of the entire intelligent document question-answering system.
[0069] The following is a concrete example to illustrate this. When maintenance personnel input a query into the intelligent question-and-answer system, such as "the data transmission of that box is abnormal," the system first activates a query parsing unit. In step S311, this unit uses natural language processing technology to identify the core functional symptom "data transmission abnormality" and the non-standard, colloquial device term "box." In step S312, the system obtains association rules and priority information related to "data transmission abnormality." For example, the rules may define a strong association between "data transmission abnormality" and "communication module failure," and a weak association with "power module failure," with "communication module failure" having a higher priority than "power module failure." Simultaneously, the system may also obtain mutually exclusive rules between "data transmission abnormality" and "device offline." In step S313, the system analyzes the logical relationships of "data transmission abnormality" based on these association rules. If the query also includes "device offline," the system identifies the conflict between the two. If the query only contains "data transmission abnormality," the symptom is initially deemed valid. In step S314, the system determines "data transmission anomaly" as a core functional symptom based on the analysis and identification results and priority information. For example, if the query mentions both "data transmission anomaly" and "communication module failure," and "communication module failure" is given a higher priority in the ontology knowledge (because it may be a deeper cause of the data transmission anomaly), the system may identify "communication module failure" as the core functional symptom. In this way, the system can accurately identify the core intent of the query, even if the query contains non-standard terminology or potential symptom conflicts.
[0070] Through the above technical solution, this application effectively resolves potential conflicts or redundancies among symptoms, significantly improving the accuracy of identifying core functional symptoms. This provides more precise contextual information for subsequently mapping non-standard terms to standard technical entities, thereby ensuring the reliability of the mapping results. Ultimately, the solution of this application enables the intelligent document question-and-answer system to provide more accurate and targeted diagnostic results when processing queries containing complex descriptions and non-standard terms, greatly enhancing the system's practical value and decision-making support capabilities in power operation and maintenance scenarios.
[0071] In some embodiments, the specific steps in step S313 include: S3131. Based on association rules, conduct a preliminary analysis of the logical relationships between descriptive symptoms, and preliminarily identify whether there are conflicts or redundancies in the descriptive symptoms, and obtain preliminary analysis and identification results; S3132. Determine whether the preliminary analysis and identification results indicate the existence of symptom combinations not explicitly covered by association rules, or whether the preliminary analysis and identification results have multiple logical interpretations; S3133. If the preliminary analysis and identification results indicate that there are symptom combinations that are not explicitly covered by the association rules, or if the preliminary analysis and identification results have multiple logical interpretations, then obtain the real-time operational data corresponding to the descriptive symptoms; S3134. Based on real-time operating data and combined with the preset real-time anomaly judgment logic in the knowledge of the power equipment itself, determine whether there are other associations, conflicts or redundancies in the symptom combination that are not covered by the association rules, and obtain the real-time judgment result. S3135. If the real-time judgment result indicates the existence of other associations, conflicts, or redundancies not covered by the association rules, then the real-time judgment result shall be used as the analysis and identification result; S3136. Otherwise, the preliminary analysis and identification results shall be taken as the analysis and identification results.
[0072] The association rules are pre-defined knowledge sets used to describe known logical relationships between symptoms of power equipment faults. These rules can be constructed based on the experience of power industry experts, historical fault data analysis, or equipment physical models. For example, the association rules can be stored as an IF-THEN rule set for pattern matching in a knowledge base, or a semantic network can be constructed in the form of an ontology, connecting symptoms, equipment, and faults through predefined semantic relationships (such as "cause," "accompany," and "mutually exclusive"). Preliminary analysis of the logical relationships between the descriptive symptoms refers to using the aforementioned association rules to perform preliminary reasoning and matching on the descriptive symptoms extracted from the query request to understand whether there are known causal, accompanying, or mutually exclusive relationships between them. This can be achieved by using a rule engine to perform pattern matching on the symptom set to identify symptom combinations that conform to pre-defined rules, or by using a graph traversal algorithm to find the paths and relationship types between symptom nodes in the semantic network. Preliminary identification of whether there are conflicts or redundancies in the descriptive symptoms refers to determining, based on the preliminary analysis, whether there are contradictory symptoms (conflicts) or symptoms expressing the same meaning (redundancy) in the symptom set. Conflict identification can be accomplished by checking whether mutually exclusive rules are triggered simultaneously, while redundancy identification can be accomplished by checking whether multiple symptoms point to the same underlying physical phenomenon or failure mode, or by calculating the semantic similarity or logical distance between symptoms. The preliminary analysis and identification results are the output of this step, including a preliminary judgment on the logical relationship of symptoms and the identification of potential conflicts or redundancies.
[0073] Determining whether the preliminary analysis results indicate the existence of symptom combinations not explicitly covered by the association rules, or whether the preliminary analysis results have multiple logical interpretations, aims to identify uncertainties in the analysis process. Determining whether there are symptom combinations not explicitly covered by the association rules means checking whether there are symptom combinations in the preliminary analysis results that cannot be explained by existing association rules; that is, the relationship between these symptoms is unknown or ambiguous. This can be determined by checking whether the rule engine has unmatched rules or cannot reach a clear conclusion when processing symptom combinations, or by evaluating the connection density or path completeness of the symptom combinations in the knowledge graph. If the connections are sparse or the paths are incomplete, they may not be explicitly covered. Determining whether the preliminary analysis results have multiple logical interpretations means checking whether there are multiple possible logical interpretations in the preliminary analysis results, making it impossible to determine a unique and clear symptom relationship. This can be determined by checking whether the rule engine has triggered multiple competing or uncertain rules, leading to ambiguity in the judgment of symptom relationships, or by evaluating the number and confidence of paths from symptom combinations to different failure modes in the semantic network. If multiple paths with similar confidence exist, multiple interpretations may exist.
[0074] If the preliminary analysis and identification results indicate the existence of symptom combinations not explicitly covered by the association rules, or if the preliminary analysis and identification results have multiple logical interpretations, then the real-time operational data corresponding to the descriptive symptom is obtained. Real-time operational data refers to the current or recent operational status data of the power equipment related to the descriptive symptom, such as sensor data, telemetry data, alarm information, etc. The acquisition method can be through interface with SCADA (Supervisory and Data Acquisition) systems, EMS (Energy Management System), or equipment IoT platforms to query the operational data of relevant equipment in real time, or by subscribing to the real-time data streams of relevant equipment through a data bus or message queue and retrieving them from the cache when needed.
[0075] Based on the real-time operational data and combined with the pre-defined real-time anomaly judgment logic within the power equipment ontology knowledge, the system determines whether the symptom combination exhibits novel correlations, conflicts, or redundancies, thus obtaining a real-time judgment result. The pre-defined real-time anomaly judgment logic within the power equipment ontology knowledge consists of rules or models used to process real-time data, aiming to dynamically assess the authenticity, correlation, and presence of anomalies in the symptoms. This can be achieved by constructing anomaly detection algorithms based on statistical or machine learning models, such as classifiers or clustering models trained on historical data, to determine whether real-time data deviates from normal patterns; or by using expert system rules that include dynamic thresholds, trend analysis rules, or physical constraints, such as "If the temperature rises by Y degrees within X minutes and the current exceeds Z, an overheating fault may exist." Determining whether the symptom combination exhibits novel correlations, conflicts, or redundancies refers to using real-time data and the real-time anomaly judgment logic to perform a deeper dynamic analysis of the symptom combination to discover novel correlations that static rules fail to capture, confirm or eliminate conflicts, or identify deeper levels of redundancy. This can be achieved by inputting real-time data into an anomaly detection model to assess the degree and pattern of anomalies in symptom combinations, thereby inferring whether novel associations or conflicts exist, or by verifying physical causal chains between symptoms through real-time data. The real-time judgment result is the output of this step, representing a more accurate and timely assessment of the relationships between symptom combinations based on real-time data and dynamic logic.
[0076] If the real-time judgment result indicates the existence of a new type of association (i.e., other associations not covered by the association rules), conflict, or redundancy, then the real-time judgment result is adopted as the analysis and identification result. This means that when real-time data analysis discovers new and more accurate symptom relationships, the system prioritizes adopting the results of dynamic analysis. Otherwise, the preliminary analysis and identification result is adopted as the analysis and identification result. If real-time data analysis fails to discover new associations, conflicts, or redundancies, or if the real-time data itself does not provide sufficient information to overturn the preliminary analysis, then the preliminary analysis result based on static rules is retained.
[0077] This solution addresses the accuracy issues in symptom analysis when encountering situations not covered by association rules or with multiple logical interpretations by introducing real-time operational data and real-time anomaly judgment logic. Specifically, the system first performs a preliminary analysis of the logical relationships between descriptive symptoms based on preset association rules, obtaining preliminary analysis and identification results. This provides a basic analytical framework for subsequent judgments, avoiding the limitations caused by relying entirely on static rules. Subsequently, the system determines whether the preliminary analysis and identification results contain symptom combinations not explicitly covered by association rules or multiple logical interpretations, thereby identifying uncertainties in the analysis process and ensuring timely triggering of supplementary mechanisms when rules are insufficient. If such situations exist, the system proactively acquires real-time operational data corresponding to the descriptive symptoms, introducing a dynamic data source to compensate for the shortcomings of association rules in covering new or ambiguous symptoms. Next, based on the real-time operational data and combined with the preset real-time anomaly judgment logic in the power equipment ontology knowledge, the system dynamically judges whether there are new association relationships, conflicts, or redundancies in symptom combinations, thereby using real-time data to dynamically evaluate the logic between symptoms and solving the problem that rules cannot handle unknown or complex scenarios. Ultimately, if the real-time judgment result indicates the existence of new relationships, conflicts, or redundancies, then the real-time judgment result is taken as the final analysis and identification result, ensuring that the latest dynamic analysis is used when new problems are discovered, thereby improving the accuracy of identification; otherwise, the preliminary analysis and identification result is taken as the final analysis and identification result, which preserves the analysis efficiency when there are no anomalies and avoids unnecessary resource consumption.
[0078] This approach is closely integrated with step S31, which identifies core functional symptoms, and step S3, which maps non-standard terms. By dynamically verifying symptom relationships, it ensures the accuracy and robustness of core functional symptom identification. This enables subsequent non-standard term mapping and fault evolution path determination to be based on more accurate semantic understanding, thereby significantly improving the diagnostic accuracy and reliability of the entire intelligent document question-answering system when dealing with complex, ambiguous, or novel "soft faults."
[0079] The following is a specific example to illustrate this. Suppose that maintenance personnel query "a feeder tripped, ambient temperature is high, and the protection device malfunctioned." In step S311 above, the system extracts three descriptive symptoms from the query request: "feeder tripped," "high ambient temperature," and "protection device malfunction." In step S312, the system obtains the preset association rules and priority information for these symptoms in the power equipment's knowledge. Moving to step S3131 of this solution, the system performs a preliminary analysis of these symptoms based on the association rules. For example, the system might find that "feeder tripped" is the result, and "protection device malfunction" is one of the causes, but there is no direct, explicit preset association rule between "high ambient temperature" and "protection device malfunction," or there are multiple possibilities (high temperature may affect the protection device or other equipment, causing a trip). In this case, the preliminary analysis and identification results indicate the existence of symptom combinations or multiple logical interpretations not explicitly covered by the association rules.
[0080] In step S3132, the system determines that the preliminary analysis and identification results do indeed indicate the existence of symptom combinations and multiple logical interpretations not explicitly covered by the association rules. Therefore, in step S3133, the system actively acquires real-time operational data corresponding to these descriptive symptoms. For example, the system queries real-time data from the feeder, protection device, and environmental sensors, including internal temperature sensor data from the protection device, CPU load, memory usage, voltage and current waveforms before and after the feeder trips, and historical trends from the environmental temperature sensor.
[0081] In step S3134, the system analyzes the real-time operational data based on preset real-time anomaly judgment logic within the power equipment's inherent knowledge. For example, the preset real-time anomaly judgment logic might include: if the internal temperature of the protection device rises synchronously with the ambient temperature and exceeds its design threshold, a thermal stability problem may exist; if the CPU load of the protection device rises abnormally instantaneously before tripping, it may indicate a software or hardware fault; if the voltage and current waveforms show no external short-circuit characteristics, but the protection device operates, this further supports the judgment of maloperation. Through real-time data analysis, the system discovered that when the ambient temperature of this type of protection device exceeds 40 degrees Celsius, the temperature of a key internal component also rises synchronously and approaches its operating limit, while its internal diagnostic log shows intermittent self-test failure records. This indicates a new correlation between "high ambient temperature" and "maloperation of the protection device," namely, high temperature leads to a decrease in the thermal stability of the internal components of the protection device, thereby triggering maloperation. This analysis result is the real-time judgment result.
[0082] Since the real-time judgment result indicates the existence of a novel correlation, in step S3135, the system uses this real-time judgment result as the final analysis and identification result. Through the above technical solution, this method effectively solves the limitations encountered by traditional methods in symptom analysis, namely, the potential inaccuracy of analysis when symptom combinations are not explicitly covered by association rules or have multiple logical interpretations. By introducing real-time operational data and real-time anomaly judgment logic, this method can dynamically discover novel correlations, confirm or eliminate conflicts, and identify redundancy, thereby significantly improving the accuracy of identifying logical relationships between descriptive symptoms. This makes the subsequent step S31, which identifies core functional symptoms, more reliable, thereby optimizing the mapping of non-standard terms and the determination of fault evolution paths, ultimately improving the overall performance and user trust of the power supply bureau's intelligent document question-and-answer system in complex fault diagnosis scenarios.
[0083] In some embodiments, step S3 may further include: S35. Based on the mapping results, adjust the association weights between non-standard terms and standard technical entities.
[0084] "Based on the mapping result" refers to the final mapping result determined during the process of mapping non-standard terms to standard technical entities. This result can be a specific standard technical entity, or it can be the system's confidence level or priority score for the mapping. Its purpose is to provide a basis for subsequent weight adjustments, ensuring that the adjustments are based on actual mapping results rather than preset rules. For example, the system can record each successful mapping pair (non-standard term, standard technical entity) or record the confidence score for each mapping. "Adjusting the association weight between the non-standard term and the standard technical entity" refers to modifying the numerical value stored internally in the system that represents the strength of the association between the non-standard term and the standard technical entity. This weight can reflect the reliability of the mapping relationship, its frequency of use, or the system's level of trust in it. Its purpose is to enable the system to dynamically optimize the mapping strategy based on actual operational feedback, improving the accuracy and efficiency of future mappings. For example, when the system successfully maps a non-standard term to a standard technical entity, and the mapping is accepted by the user or verified as correct by subsequent operations, the system can increase the association weight of the mapping pair; or, the system can provide a user interface that allows operations and maintenance personnel to explicitly confirm or correct the mapping results, thereby adjusting the weight; or, as the mapping between a non-standard term and a standard technical entity is frequently used and verified to be effective, the system can automatically increase its association weight.
[0085] When the system successfully maps a non-standard term to a standard technical entity in steps S31 to S34, this mapping result is not merely a temporary match, but a valuable learning signal. This application dynamically adjusts the association weights between the non-standard term and the standard technical entity based on this mapping result. This means that if a specific non-standard term (e.g., "box") is successfully and with high confidence mapped to a standard technical entity (e.g., "smart meter") in a specific context (e.g., "data transmission anomaly"), the system will strengthen the association between "box" and "smart meter" in the "data transmission anomaly" context. This dynamic adjustment mechanism allows the system to learn from each successful mapping and continuously optimize its internal "domain expert consensus mapping table." Over time and with more queries processed, mappings that are frequently verified as accurate will gradually increase in association weight, thus gaining higher priority in future mapping processes. Conversely, if a mapping is found to be inaccurate or rarely used, its weight may decrease. This mechanism solves the problem of fixed association weights in traditional systems. Without dynamic weight adjustment, even if the system finds an accurate mapping through S31-S34, the successful experience of this mapping will not feed back into the mapping mechanism itself. This means that when faced with new, similar queries, the system still needs to perform complex re-evaluations, or the mapping may fail due to unreasonable preset weights. Through dynamic weight adjustment, the system can optimize the association strength based on actual mapping feedback, avoiding mismatches in preset weights, thereby improving the accuracy and robustness of the system in handling non-standard terms. It enables the system to learn and adapt to new terms or data changes, optimizing the efficiency and reliability of future mappings.
[0086] In one specific implementation, when the system evaluates and selects the standard technical entity with the highest anomaly relevance based on real-time operational anomaly data in step S34, thereby mapping the non-standard term "box" to the standard technical entity "smart meter," the system immediately initiates a weight adjustment mechanism. Specifically, the system maintains a "domain expert consensus mapping table," which records the association weights between non-standard expressions (such as "box"), standard entities (such as "smart meter"), and functional symptom contexts (such as "data transmission anomaly"). At this time, the system updates the association weights of the mapping pair based on the successful mapping result. For example, an Exponentially Weighted Moving Average (EWMA) method can be used for updating. Assuming the old association weight is W_old, the current successful mapping weight contribution is W_current (for example, it can be set to 1.0 to indicate a successful mapping), and the learning rate is alpha (for example, it can be set to 0.1). Then, the new association weight W_new will be calculated as: W_new = alpha * W_current + (1 - alpha) * W_old. In this way, each successful mapping positively impacts the association weights, gradually adjusting them towards higher confidence levels. Furthermore, to further accelerate learning and adaptation, the system can provide a user feedback interface. For example, after outputting diagnostic results, the system can ask operations personnel, "Is this mapping result accurate?" If the operations personnel select "yes," the system will further increase the weight of that mapping pair; if the operations personnel select "no" and provide a correct mapping, the system will decrease the weight of the original mapping and increase the weight of the user-specified new mapping. This interactive feedback mechanism enables the system to adapt to new non-standard terms and their corresponding standard technical entities more quickly and accurately, thereby continuously optimizing the accuracy of the mapping table.
[0087] Through the above technical solution, this application can dynamically adjust the association weights between non-standard terms and standard technical entities based on the actual mapping results during the process of mapping non-standard terms to standard technical entities. This effectively solves the problem that the association weights in traditional methods remain fixed, causing the system to be unable to dynamically adapt to new data or optimize mapping accuracy. Specifically, when the system successfully completes a mapping from a non-standard term to a standard technical entity, the successful mapping experience is fed back into the weight adjustment mechanism, strengthening the association weights of the mapping pair. This allows the system to learn from each successful mapping and continuously optimize its internal mapping table. Over time and with the processing of more queries, the association weights of mapping relationships that are frequently verified as accurate will gradually increase, thus gaining higher priority and confidence in future mapping processes. This dynamic adjustment mechanism enables the system to optimize association strength based on actual mapping feedback, avoiding mismatch problems that may be caused by preset weights, and significantly improving the accuracy and robustness of the system in processing non-standard terms. It allows the system to continuously learn and adapt to new non-standard terms or data changes, thereby optimizing the efficiency and reliability of future mappings and ensuring that the intelligent document question-answering system maintains a high level of performance and accuracy during long-term operation.
[0088] In some embodiments, the specific steps in step S4 further include: S41. Using quantified semantic features and standard technical entities as query inputs to the knowledge graph, causal logic matching is performed through a pre-established semantic knowledge graph in the power field to determine the fault evolution path between the target object and operating parameters; the semantic knowledge graph in the power field defines power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them. S42. Output the diagnostic results based on the fault evolution path.
[0089] The purpose of using the quantized semantic features and standard technical entities as query inputs to the knowledge graph is to transform the preprocessed and standardized query information into a format that the knowledge graph can understand and process. This serves as the starting point or constraint for graph queries, ensuring the accuracy and semantic consistency of the input information, thereby enabling effective reasoning and matching within complex knowledge networks. This can be achieved by converting the quantized semantic features and standard technical entities (e.g., unique identifiers or standard fault codes of devices) into node or relational attributes in graph database query languages (such as Cypher for Neo4j or SPARQL for RDF graphs), serving as starting conditions for graph traversal or pattern matching. Alternatively, an intermediate layer can be constructed to map these features and entities to predefined query templates, which are then translated into query statements from the underlying knowledge graph, thus enabling query input.
[0090] The pre-established semantic knowledge graph for the power industry is a knowledge base specifically built for the power sector, containing rich semantic information. Its core function is to provide structured professional knowledge in the power field, including the complex relationships between equipment, parameters, faults, and causes, to support accurate causal reasoning. This knowledge graph can be constructed based on ontology, using standard formats such as RDF / OWL. Through expert experience, historical document mining, and machine learning methods, it formally defines and stores entities such as power equipment, operating parameters, fault phenomena, and fault causes, as well as the causal, temporal, and attribute relationships between them. Alternatively, it can be stored and managed using graph databases (such as Neo4j and JanusGraph), where nodes represent entities in the power field (such as transformers, voltage, overload), edges represent relationships between entities (such as "cause," "belongs to," "monitor"), and attributes are attached to nodes and edges to store detailed information.
[0091] The aforementioned causal logic matching refers to finding logical paths within a knowledge graph that can explain or predict the cause or evolution of a specific phenomenon (such as a failure) based on a query input. Its aim is to reveal deep causal relationships between events, rather than simple surface associations. This can be achieved by using graph traversal algorithms (such as Depth-First Search (DFS) or Breadth-First Search (BFS)) or pathfinding algorithms (such as Dijkstra's algorithm for finding the shortest path, or the heuristic-based A* algorithm) to search for causal chains from the query entity to the failure phenomenon within the knowledge graph. Alternatively, a rule-based reasoning engine can be used, combining causal rules defined in the knowledge graph (e.g., if A occurs and B occurs, it may lead to C), to logically reason about the query input and deduce possible failure evolution paths.
[0092] Determining the fault evolution path between the target object and the operating parameters is the result of causal logic matching, that is, identifying a series of causal events and state changes from the initial anomaly (the abnormal state of the target object and operating parameters) to the final fault phenomenon. Its purpose is to provide a clear and traceable fault development chain. The fault evolution path can be represented as an ordered sequence of nodes and edges in a knowledge graph, where each node represents an event or state, and each edge represents a causal or temporal relationship. Alternatively, it can be achieved by generating a structured text description that details where the fault started, what intermediate steps it went through, what the final result was, and identifies the key equipment and parameters involved.
[0093] The semantic knowledge graph for the power sector defines power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them. It emphasizes the core content and structure of the knowledge graph, meaning it not only stores data but, more importantly, explicitly encodes the complex semantic relationships between various entities in the power sector, especially causal and temporal relationships, which are the foundation for advanced reasoning. In the graph, different types can be defined for entities (such as "transformer," "overload," and "insulation aging"), and different predicates can be defined for relationships (such as "leads to," "is a component of," and "prior to"). Causal relationships can be represented as edges of "cause-leads to-effect," and temporal relationships can be represented as edges of "event A-prior to-event B." Relationships can be further refined by adding attributes; for example, causal relationships can have a "confidence" attribute, and temporal relationships can have a "time interval" attribute, to enhance the expressive power of the graph and the accuracy of reasoning.
[0094] Based on the described fault evolution path, diagnostic results are output. This involves presenting the fault evolution path obtained through causal logic matching in a user-understandable manner as the final response to the query request. Its purpose is to provide maintenance personnel with clear fault diagnosis information and decision support. Diagnostic results may include the root cause of the fault, possible intermediate steps, involved equipment and parameters, and recommended troubleshooting or repair solutions. This information can be presented in the form of structured text, charts, or visual paths. Diagnostic results can also be linked to historical fault cases, equipment maintenance manuals, and other documents, providing links or summaries of relevant documents for users to further access detailed information.
[0095] This application enhances the accuracy of causal logic matching by introducing a semantic knowledge graph in the power industry, addressing the shortcomings of knowledge association networks in terms of definition and structure. This allows for more accurate determination of fault evolution paths and improved diagnostic reliability. Upon receiving query information that has been parsed, quantified semantically, and mapped from non-standard terms to standard technical entities, this solution uses these precise quantified semantic features and standard technical entities as input to the knowledge graph. This input accurately reflects the query intent, avoiding matching biases caused by semantic ambiguity. Subsequently, the system performs causal logic matching using a pre-established semantic knowledge graph in the power industry. This knowledge graph is the core of this solution; it specifically and explicitly defines power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them. This structured, semantically rich knowledge representation enables the matching process to perform deep reasoning based on professional knowledge in the power industry, effectively handling complex causal chains and temporal dependencies between devices. By searching and identifying these predefined causal and temporal relationships in the graph, the system can determine the complete fault evolution path between the target object and the operating parameters, thereby identifying fault modes across devices and multiple event chains, ensuring the completeness and accuracy of the path. Ultimately, based on this determined fault evolution path, the system outputs detailed diagnostic results. This precise path based on graph matching ensures that the diagnostic results comprehensively reflect the fault evolution process, significantly improving the reliability and practicality of the output. This solution overcomes the shortcomings of traditional knowledge association networks in definition and structure by combining refined query inputs (quantified semantic features and standard technical entities) with a structured, semantically rich semantic knowledge graph of the power sector. Quantified semantic features provide precise numerical or trend-based representations of "soft faults," standard technical entities resolve ambiguities in non-standard terminology, and the semantic knowledge graph of the power sector provides the knowledge foundation for deep causal reasoning. The synergistic effect of these three elements enables the system to delve deeper from surface phenomena to the essential causes and evolution mechanisms of faults, thus providing a more comprehensive and accurate diagnostic path than traditional methods when facing complex fault scenarios.
[0096] For example, suppose an operations and maintenance personnel input the query: "The partial discharge signal strength of a certain transformer fluctuates drastically, and the cooling fan speed decreases abnormally, causing the transformer temperature to rise continuously." First, the system will parse, quantify, and map the query to entities according to the steps S1, S2, and S3 described above. For example, "drastic fluctuations in partial discharge signal strength" will be quantified as a specific anomaly level and trend characteristic, while "abnormal decrease in cooling fan speed" will be identified as a specific abnormal operating parameter and mapped to the standard technical entity "cooling fan failure." "Continuous rise in transformer temperature" will be identified as a fault phenomenon. Next, in step S41, the system submits these quantified semantic features (such as "high partial discharge signal anomaly level" and "cooling fan speed below threshold") and standard technical entities (such as "transformer" and "cooling fan") as query inputs to a pre-established semantic knowledge graph in the power field. The knowledge graph explicitly defines: entities, such as "transformer," "cooling fan," "partial discharge," "temperature rise," and "insulation aging"; causal relationships, such as "partial discharge leads to insulation aging," "insulation aging leads to temperature rise," "cooling fan failure leads to decreased cooling capacity," and "decreased cooling capacity leads to temperature rise"; and temporal relationships, such as "partial discharge precedes insulation aging" and "cooling fan failure may occur simultaneously with temperature rise." The system utilizes the graph database's inference engine to perform causal logic matching within the knowledge graph. It searches for all possible paths from the starting points of "abnormal partial discharge signal" and "abnormal cooling fan speed" to the fault phenomenon of "continuously rising transformer temperature." For example, the system might identify two main fault evolution paths: Path 1: Abnormal partial discharge signal -> intensified partial discharge -> accelerated insulation aging -> increased internal transformer losses -> continuously rising transformer temperature. Path 2: Abnormally reduced cooling fan speed -> cooling fan failure -> decreased cooling capacity -> poor transformer heat dissipation -> continuously rising transformer temperature. The system comprehensively considers these two paths and determines the most likely fault evolution path based on the confidence level or weight defined in the knowledge graph. For example, if the causal relationship between "intensified partial discharge" and "accelerated insulation aging" is defined in the knowledge graph with a high confidence level, and the causal relationship between "cooling fan failure" and "poor heat dissipation" is also clear, the system will identify this as a complex fault caused by multiple factors. Finally, in step S42, the system outputs the diagnostic results based on the determined fault evolution path. The results may include: Diagnostic conclusion: The continuous rise in transformer temperature is the result of the combined effect of insulation aging caused by partial discharge and cooling fan failure; Root cause: Partial discharge and cooling fan failure; Evolution process: Partial discharge leads to insulation degradation, increasing internal heat generation; at the same time, the cooling fan failure leads to insufficient heat dissipation, and the combination of the two results in an abnormal temperature rise.
[0097] Through the above technical solution, this application effectively addresses the shortcomings of traditional knowledge association networks in terms of definition and structure. Specifically, it lacks clear definitions of power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them, leading to inaccurate matching and an inability to effectively handle complex fault scenarios involving multiple devices and event chains, thus affecting the accuracy and comprehensiveness of diagnostic results. Specifically, by using quantified semantic features and standard technical entities as query inputs and utilizing a pre-established semantic knowledge graph in the power field for causal logic matching, the system can accurately identify and trace the complete evolution path from abnormal phenomena to final faults based on professional knowledge in the power field. This knowledge graph clearly defines power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them, enabling the system to deeply understand the internal mechanisms of complex faults and effectively handle fault scenarios involving multiple factors and multiple stages. Therefore, this solution can provide more comprehensive and accurate fault diagnosis results, significantly improving the diagnostic reliability and practicality of intelligent document question-and-answer systems in power operation and maintenance scenarios, thereby assisting maintenance personnel in quickly locating the root cause of faults and improving fault handling efficiency.
[0098] Reference Appendix Figure 2 This invention provides a semantic analysis-based intelligent document question-answering optimization system for power supply bureaus (this semantic analysis-based intelligent document question-answering optimization system for power supply bureaus adopts the semantic analysis-based intelligent document question-answering optimization method for power supply bureaus described in the above embodiments, and the specific process is described in the corresponding steps above), including: The identification module 100 is used to parse the input query request and identify the target object, running parameters, descriptive feature words and non-standard terms in the query request; The processing module 200 is used to obtain the physical operation feature information corresponding to the target object, and to quantize the descriptive feature words according to the physical operation feature information to obtain quantized semantic features; The mapping module 300 is used to map non-standard terms to standard technical entities based on the context information of the query request. Output module 400 is used to determine the fault evolution path between the target object and the operating parameters based on quantized semantic features and standard technical entities, and output the diagnostic results.
[0099] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0100] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for optimizing intelligent document question-answering for power supply bureaus based on semantic analysis, characterized in that, Includes the following steps: S1. Parse the input query request and identify the target object, running parameters, descriptive feature words, and non-standard terms in the query request; S2. Obtain the physical operation feature information corresponding to the target object, and quantize the descriptive feature words according to the physical operation feature information to obtain quantized semantic features; S3. Based on the context information of the query request, map the non-standard terms to standard technical entities; S4. Based on the quantized semantic features and the standard technical entities, determine the fault evolution path between the target object and the operating parameters, and output the diagnostic results.
2. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 1, characterized in that, The physical operation characteristic information includes the type information of the target object, the physical threshold information of the monitoring parameters, and historical operation statistics.
3. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 1 or 2, characterized in that, In step S2, the step of quantizing the descriptive feature words based on the physical operation feature information to obtain quantized semantic features includes: S21. Based on the physical operation characteristic information, calculate the rate of change of the parameter corresponding to the descriptive feature word within a preset time window or the degree of deviation from the preset standard value; S22. The descriptive feature words are quantified by comparing the rate of change or the degree of deviation with their respective preset thresholds to obtain the quantized semantic features.
4. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 3, characterized in that, The specific steps in step S22 include: S221. Obtain real-time and historical data of multiple operating parameters involved in the query request; S222. Identify the joint behavior of the multiple operating parameters over a specific time series; S223. Based on the joint behavior and a preset parameter association rule, determine whether the joint behavior conforms to a preset combination pattern; S224. When the joint behavior conforms to the combination pattern, and the rate of change or the degree of deviation of any of the multiple operating parameters does not reach their respective preset thresholds, the descriptive feature words are quantized according to the combination pattern to obtain the quantized semantic features; otherwise, the descriptive feature words are quantized by comparing the rate of change or the degree of deviation with their respective preset thresholds to obtain the quantized semantic features.
5. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 1, characterized in that, The specific steps in step S3 include: S31. Identify the core functional symptoms in the query request; S32. Generate a candidate list of standard technical entities for the non-standard terms; S33. Obtain real-time operational anomaly data related to the core functional symptoms for each standard technical entity in the candidate list; the real-time operational anomaly data serves as the context information for the query request; S34. Based on the real-time operational anomaly data, assess the abnormal correlation between each standard technical entity and the core functional symptoms, and select the standard technical entity with the highest abnormal correlation as the mapping result of the non-standard term, thereby mapping the non-standard term to the standard technical entity.
6. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 5, characterized in that, The specific steps in step S31 include: S311. Extract all descriptive symptoms from the query request; S312. Obtain the preset association rules and priority information for each of the described descriptive symptoms; S313. Based on the association rules, analyze the logical relationship between each of the descriptive symptoms, and identify whether there is conflict or redundancy among each of the descriptive symptoms, to obtain the analysis and identification results; S314. Based on the analysis and identification results, and in conjunction with the priority information, determine the core functional symptoms.
7. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 6, characterized in that, The specific steps in step S313 include: S3131. Based on the association rules, perform a preliminary analysis of the logical relationships between the descriptive symptoms, and preliminarily identify whether there are conflicts or redundancies among the descriptive symptoms, to obtain preliminary analysis and identification results; S3132. Determine whether the preliminary analysis and identification results indicate the existence of symptom combinations not explicitly covered by the association rules, or whether the preliminary analysis and identification results have multiple logical interpretations; S3133. If the preliminary analysis and identification results indicate that there are symptom combinations not explicitly covered by the association rules, or if the preliminary analysis and identification results have multiple logical interpretations, then obtain the real-time running data corresponding to the descriptive symptoms; S3134. Based on the real-time running data and combined with the preset real-time anomaly judgment logic, determine whether the symptom combination has other associations, conflicts or redundancies not covered by the association rules, and obtain the real-time judgment result; S3135. If the real-time judgment result indicates the existence of other associations, conflicts, or redundancies not covered by the association rules, then the real-time judgment result shall be used as the analysis and identification result; S3136. Otherwise, the preliminary analysis and identification results shall be used as the analysis and identification results.
8. The intelligent document question-answering optimization method for power supply bureaus based on semantic analysis according to claim 1, characterized in that, The specific steps in step S4 also include: S41. Using the quantified semantic features and the standard technical entities as query inputs to the knowledge graph, the fault evolution path between the target object and the operating parameters is determined through a pre-established semantic knowledge graph in the power field; the semantic knowledge graph in the power field defines power equipment, operating parameters, fault phenomena, fault causes, and the causal and temporal relationships between them. S42. Output the diagnostic results based on the fault evolution path.
9. A power supply bureau intelligent document question-answering optimization system based on semantic analysis, characterized in that, include: The identification module is used to parse the input query request and identify the target object, running parameters, descriptive feature words and non-standard terms in the query request; The processing module is used to obtain the physical operation feature information corresponding to the target object, and to quantize the descriptive feature words according to the physical operation feature information to obtain quantized semantic features; The mapping module is used to map the non-standard terms to standard technical entities based on the context information of the query request; The output module is used to determine the fault evolution path between the target object and the operating parameters based on the quantized semantic features and the standard technical entities, and to output the diagnostic results.