Fire-fighting water supply system fault knowledge extraction method and system, processing equipment and storage medium

By building a fire water supply system fault knowledge extraction model based on UIE framework and T5 large language model, the problem of low fault diagnosis accuracy in traditional methods is solved, and the conversion from unstructured text to structured knowledge is realized, which improves the fault diagnosis efficiency and accuracy of the fire water supply system.

CN120258107APending Publication Date: 2025-07-04CHINA UNIV OF PETROLEUM (BEIJING)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510282888.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing fault diagnosis methods for fire water supply systems are difficult to process complex semantic information due to the lack of a systematic diagnostic knowledge base and traditional knowledge extraction model, resulting in low fault diagnosis accuracy and insufficient generalization capabilities, and it is impossible to quickly and accurately extract structured fault knowledge from unstructured text data.

Method used

The fire water supply system fault knowledge extraction model is adopted based on the UIE framework and the T5 major language model. Through fine-tuning and pre-training, a fire water supply system fault knowledge extraction model is built to realize the transformation from unstructured text data to structured knowledge, and fuse it into a knowledge graph.

Benefits of technology

It realizes the rapid and accurate extraction of fault knowledge of structured fire water supply system from unstructured text data, supports more accurate fault diagnosis and maintenance, and builds a comprehensive and accurate fault knowledge graph, which improves diagnostic efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258107A_ABST
    Figure CN120258107A_ABST
Patent Text Reader

Abstract

The invention relates to a fire-fighting water supply system fault knowledge extraction method and system, processing equipment and a storage medium, and the method is characterized in that the method comprises the steps: carrying out the fine adjustment of a trained fire-fighting water supply system fault knowledge extraction model for a fire-fighting water supply system knowledge extraction task, the fire-fighting water supply system fault knowledge extraction model is constructed based on a UIE framework and a T5 big language model; performing fault knowledge extraction on the fire-fighting water supply system fault knowledge prediction by adopting the fine-tuned fire-fighting water supply system fault knowledge extraction model; and performing fusion processing on the extracted fault knowledge to enable the fault knowledge to reach the standard of constructing the knowledge graph, and storing the fault knowledge to form the fault knowledge graph of the fire-fighting water supply system, and the method can be widely applied to the technical field of artificial intelligence and big data application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and big data applications, and particularly to a method, system, processing device and storage medium for extracting fault knowledge of a fire water supply system. Background Art

[0002] Under the background of the increasingly severe current fire prevention situation, China's fire protection policies are actively promoting intelligent and information-based technological innovations. As a key infrastructure in the fire protection field, the reliability of the fire water supply system is directly related to the effectiveness of fire prevention and control. Existing fire water supply systems often fail to meet the expected reliability requirements due to untimely maintenance or inaccurate detection, and regular diagnosis has become a necessary means. However, in actual operation, traditional manual diagnosis methods are difficult to ensure the timely discovery and accurate diagnosis of faults due to problems such as insufficient diagnostic experience and lack of systematic professional guidance.

[0003] The fault diagnosis of the fire water supply system urgently needs to rely on a professional and systematic diagnostic knowledge base, and this knowledge is usually hidden in the fault data of the fire water supply system. At present, the fault data of the fire water supply system are mostly unstructured text information, including maintenance records, operation manuals and fault reports, etc. It is necessary to convert them into structured fault knowledge before they can be used to build a knowledge base and then used for fault diagnosis of the fire water supply system.

[0004] However, these text data have context relevance and contain rich syntactic and semantic information. Traditional knowledge extraction models are difficult to process these complex semantic information, resulting in poor accuracy and insufficient generalization ability, and unable to quickly and accurately extract structured fire water supply system fault knowledge from unstructured text data. In recent years, large language models have achieved remarkable results in various natural language tasks with their powerful context learning and inference capabilities, and are an effective means to solve the above problems. Summary of the Invention

[0005] Aiming at the above problems, the purpose of the present invention is to provide a method, system, processing device and storage medium for extracting fault knowledge of a fire water supply system, which can quickly and accurately extract structured fire water supply system fault knowledge from unstructured text data.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, a method for extracting fault knowledge of a fire water supply system is provided, including:

[0007] For the task of extracting knowledge of the fire water supply system, fine-tune the trained fire water supply system fault knowledge extraction model, wherein the fire water supply system fault knowledge extraction model is constructed based on the UIE framework and the T5 large language model;

[0008] Adopt the fine-tuned fire water supply system fault knowledge extraction model to extract fault knowledge from the fire water supply system fault knowledge corpus;

[0009] Fuse and process the extracted fault knowledge to meet the standards for constructing a knowledge graph, and store it to form a fault knowledge graph of the fire water supply system.

[0010] Furthermore, the construction process of the fire water supply system fault knowledge extraction model is as follows:

[0011] Collect and preprocess the fire water supply system fault data, and construct the fault knowledge ontology of the fire water supply system;

[0012] Based on the UIE framework and the T5 large language model, construct a fire water supply system fault knowledge extraction model;

[0013] Train the constructed fire water supply system fault knowledge extraction model to obtain a trained fire water supply system fault knowledge extraction model.

[0014] Furthermore, the collection and preprocessing of the fire water supply system fault data, and the construction of the fault knowledge ontology of the fire water supply system include:

[0015] Collect and preprocess the fire water supply system fault data;

[0016] Based on the preprocessed fire water supply system fault data, model the key components of the fire water supply system, and establish the equipment entities in the fault knowledge ontology of the fire water supply system. Among them, the key components of the fire water supply system include the hierarchy and structural relationship of the fire water supply system;

[0017] Based on the preprocessed fire water supply system fault data, consider the logical relationship between the fire water supply system fault symptoms, fault causes and solutions, and establish the fault entities in the fault knowledge ontology of the fire water supply system.

[0018] Furthermore, the construction of the fire water supply system fault knowledge extraction model based on the UIE framework and the T5 large language model includes:

[0019] Taking the constructed fault knowledge ontology as a specification, construct the architecture of the fire water supply system fault knowledge extraction model;

[0020] Based on the UIE framework, the T5 large language model and the constructed architecture of the fire water supply system fault knowledge extraction model, construct the fire water supply system fault knowledge extraction model.

[0021] Furthermore, the fire water supply system fault knowledge extraction model includes an SSI structure guiding layer, a T5-based encoding and decoding layer, and an SEL structure extraction layer;

[0022] The SSI structure guiding layer is used for the extraction mechanism based on the schema mode, controls the type of information generated during extraction, creates the SSI sequence, and after feature splicing with the text sequence, serves as the input for the next layer;

[0023] The encoding and decoding layer based on T5 is used to encode the sequence after splicing the SSI sequence and the text sequence into a linear SEL, and then convert it into a text sequence through the decoding form;

[0024] The SEL structure extraction layer is used to locate and associate different IE structures, re-encode the decoded information, and convert it into the information to be extracted.

[0025] Furthermore, training the constructed fire water supply system fault knowledge extraction model to obtain a trained fire water supply system fault knowledge extraction model includes:

[0026] Obtain a number of general text data;

[0027] Based on the obtained number of general text data, pre-train the fire water supply system fault knowledge extraction model:

[0028] Divide the general text data into text-structure parallel corpus D Pair , unstructured text data set D Text , structured data set D Record ;

[0029] Use the text-structure parallel corpus D Pair to pre-train the fire water supply system fault knowledge extraction model;

[0030] Use the unstructured text data set D Text to improve the semantic representation of the fire water supply system knowledge extraction model;

[0031] Use the structured data set D Record to perform structured generation pre-training;

[0032] Merge the tasks of text-to-structure pre-training, improving the semantic representation of the fire water supply system knowledge extraction model, and structured generation pre-training to obtain the objective function of the fire water supply system fault knowledge extraction model;

[0033] For the pre-trained fire water supply system fault knowledge extraction model, perform domain adaptation training on the fault structured and semi-structured data including the standard guide data and abnormal alarm log data of the fire water supply system knowledge.

[0034] Further, for the task of extracting knowledge from the fire water supply system, the trained fire water supply system fault knowledge extraction model is fine-tuned. The model can understand and learn the historical data patterns and characteristics of fire water supply system faults, including:

[0035] Based on the fault knowledge ontology of the fire water supply system, given a labeled fire water supply system fault data, the cross-entropy loss function is used to perform preliminary fine-tuning on the fire water supply system fault knowledge extraction model;

[0036] The rejection mechanism is adopted to perform secondary fine-tuning on the preliminarily fine-tuned fire water supply system fault knowledge extraction model.

[0037] In the second aspect, a fire water supply system fault knowledge extraction system is provided, including:

[0038] A fine-tuning module for fine-tuning the trained fire water supply system fault knowledge extraction model for the task of extracting knowledge from the fire water supply system, wherein the fire water supply system fault knowledge extraction model is constructed based on the UIE framework and the T5 large language model;

[0039] A fault knowledge extraction module for extracting fault knowledge from the fire water supply system fault knowledge corpus by using the fine-tuned fire water supply system fault knowledge extraction model;

[0040] A fusion processing module for fusing and processing the extracted fault knowledge to meet the standards for constructing a knowledge graph and storing it to form a fault knowledge graph of the fire water supply system.

[0041] In the third aspect, a processing device is provided, including computer program instructions, wherein when the computer program instructions are executed by the processing device, they are used to implement the steps corresponding to the above-mentioned fire water supply system fault knowledge extraction method.

[0042] In the fourth aspect, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein when the computer program instructions are executed by a processor, they are used to implement the steps corresponding to the above-mentioned fire water supply system fault knowledge extraction method.

[0043] Due to the above technical solutions adopted by the present invention, it has the following advantages:

[0044] 1. Based on the large language model, the present invention constructs a fire water supply system fault knowledge extraction model, which can solve the problem of processing complex unstructured texts and realize the rapid and accurate conversion of unstructured fault texts into structured knowledge.

[0045] 2. Relying on the powerful language understanding ability, context learning and inference ability of the large language model, the present invention can quickly and accurately extract structured fire water supply system fault knowledge from unstructured text data.

[0046] 3. By adopting the Universal Information Extraction (UIE) framework and based on the Text–Text transfer Transformer (T5) pre-trained model, the present invention constructs a large model for extracting fire water supply system fault knowledge, which can uniformly model text data, efficiently complete entity recognition, event recognition and knowledge extraction, and more clearly display the text semantics and content relationships.

[0047] 4. The knowledge graph constructed by the present invention can comprehensively and accurately cover the knowledge content related to the fire water supply system faults. Relevant personnel can quickly obtain the required information from the stored knowledge graph data, providing a data basis for improving the efficiency and accuracy of fire water supply system fault diagnosis.

[0048] In summary, the present invention can be widely applied in the fields of artificial intelligence and big data application technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0050] Figure 1 is a schematic diagram of the method provided by an embodiment of the present invention;

[0051] Figure 2 is a schematic diagram of the fault knowledge ontology of the fire water supply system provided by an embodiment of the present invention;

[0052] Figure 3 is a schematic diagram of the structure of the large model for extracting fire water supply system knowledge provided by an embodiment of the present invention;

[0053] Figure 4 is a schematic diagram of the greedy algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The exemplary embodiments of the present invention will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0055] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0056] Although the terms first, second, third, etc. may be used herein to describe multiple elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or section from another. Unless the context clearly indicates otherwise, terms such as "first", "second", and other numerical terms when used herein do not imply an order or sequence. Thus, the first element, component, region, layer, or section discussed below may be referred to as the second element, component, region, layer, or section without departing from the teachings of the example embodiments.

[0057] At present, traditional knowledge extraction models are difficult to process such complex semantic information, resulting in poor accuracy and insufficient generalization ability, and unable to quickly and accurately extract structured fire water supply system fault knowledge from unstructured text data. The present invention proposes a method, system, processing device and storage medium for extracting fire water supply system fault knowledge. Among them, the method includes: for the fire water supply system knowledge extraction task, fine-tuning the trained fire water supply system fault knowledge extraction model; using the fine-tuned fire water supply system fault knowledge extraction model to extract fault knowledge from the fire water supply system fault knowledge corpus; fusing and processing the extracted fault knowledge to make it meet the standard for constructing a knowledge graph, and storing it to form a fault knowledge graph of the fire water supply system; the fire water supply system fault knowledge extraction model is constructed based on the UIE framework and the T5 large language model. The present invention is based on the Universal Information Extraction (UIE) framework and the Text–Text transfer Transformer (T5) pre-trained model, and establishes a large model framework for fire water supply system knowledge extraction. By converting unstructured fault text data into structured triple relationships and using the Neo4j graph database for efficient storage and query, a fault knowledge graph of the fire water supply system is formed to support more accurate fault diagnosis and maintenance operations. This innovative method will provide strong support for the intelligent fault diagnosis of the fire water supply system and fill the gap in current research.

[0058] Embodiment 1

[0059] As Figure 1 shown, this embodiment provides a method for extracting fire water supply system fault knowledge, including the following steps:

[0060] 1) Collect and preprocess the fire water supply system fault data, and construct the fault knowledge ontology of the fire water supply system.

[0061] Specifically, as Figure 2 shown, the fault knowledge ontology of the fire water supply system is a standardized description of the concepts and relationships between concepts in the field of fire water supply system faults. This description can determine the concept nodes and related relationships in the fire water supply system fault knowledge graph, and is the basis and template for constructing the knowledge graph, laying the foundation for the subsequent extraction of fire water supply system fault knowledge. The fault knowledge ontology of the fire water supply system is divided into two main parts: the equipment entity and the fault entity of the fire water supply system. Therefore, the specific process of this step is:

[0062] 1.1) Collect and preprocess the fire water supply system fault data.

[0063] Specifically, the fault data of the fire water supply system includes standard guide data such as fault manuals and alarm procedures, behavioral data such as wind analysis tables, operation and maintenance work orders, and experience feedback, and abnormal alarm log data.

[0064] Specifically, the preprocessing of the standard guide data in the fault data of the fire water supply system is as follows: Use the tool library to extract the text content from PDF data such as guide specifications, process non-text content such as headers, footers, and watermarks, remove redundant spaces and line breaks, and finally convert it into standard format text data. The preprocessing of the behavioral data in the fault data of the fire water supply system is as follows: Export the required field data through the operation and maintenance system of the fire water system. The preprocessing of the abnormal alarm log data in the fault data of the fire water supply system is as follows: Develop a log parsing tool to determine the data field delimiter (such as space, comma, tab), and use regular expressions to extract fields.

[0065] Specifically, the preprocessing also includes data cleaning of the collected fault data of the fire water supply system. By filling, deleting, or interpolating missing data, handling missing values in the data, deleting or correcting abnormal values or error records in the data, and deleting duplicate data records, etc., the multi-source heterogeneous data, that is, the fault data of the fire water supply system, is converted into standard text data.

[0066] 1.2) Based on the preprocessed fault data of the fire water supply system, model the key components of the fire water supply system, and establish the equipment entities in the fault knowledge ontology of the fire water supply system. Among them, the key components of the fire water supply system include the hierarchical and structural relationships of the fire water supply system.

[0067] For example, it is necessary to clarify how different units, sub-units, and specific faulty components such as water tanks, water pumps, pipes, and sprinkler heads in the fire water supply system are organized and associated. This helps to accurately represent the structural relationship of the equipment in the knowledge graph and provides basic information on equipment structure for fault diagnosis.

[0068] 1.3) Based on the preprocessed fault data of the fire water supply system, consider the logical relationship among the fault symptoms, fault causes, and solutions of the fire water supply system, and establish the fault entities in the fault knowledge ontology of the fire water supply system.

[0069] Specifically, classify the faults of the fire water supply system (such as water tank leakage, water pump failure, pipe blockage, etc.) and their root causes (such as material aging, component wear, improper operation, etc.), and associate these faults and causes with the corresponding corrective measures (such as repairing the water tank, replacing water pump components, cleaning the pipeline, etc.). In this way, the ontology can represent the fault information of the fire water supply system more structurally and systematically, thus realizing a more effective fault diagnosis and maintenance process.

[0070] 2) Based on the UIE framework (Universal Information Extraction) and the T5 (Text to Text Transfer Transformer) large language model, a fault knowledge extraction model for the fire water supply system is constructed, specifically as follows:

[0071] 2.1) Taking the constructed fault knowledge ontology as a specification, construct the architecture of the fault knowledge extraction model for the fire water supply system.

[0072] Specifically, after completing the construction of the fire water supply system fault ontology, the fault knowledge extraction of the fire water supply system can be carried out. Knowledge extraction refers to taking the fault knowledge ontology as a specification and using entity recognition and relationship extraction models in deep learning to extract the triple relationships of named entities from unstructured data. Triples are the basis for constructing a knowledge graph, including two main parts: Entity-Attribute-Value (EAV) and Entity-Relation-Entity (ERE). Among them, the Entity-Attribute-Value triple describes a certain attribute of an entity (such as a water system device) and its corresponding value, and the Entity-Relation-Entity (ERE) triple describes the relationship between two entities or events.

[0073] To extract the relationships between entities or the attributes and values of entities from the text and further represent them in the form of triples:

[0074] G = {E, R, T}(1)

[0075] Where G represents the knowledge graph; E represents the set of entities {e1, e1, ……, e E}, and the entity e is the most basic component element in the knowledge graph and is represented by a node in the graph; R represents the set of relationships {r1, r2, ……, r R}, and the relationship r is the edge of the knowledge graph, representing the connection between different entities, including attributes and relationships; F represents the set of facts {f1, f2, ……, f F}, and each fact f is defined as a triple relationship (h, r, t) ∈ f, where h represents the head entity, r represents the relationship and attribute, and t represents the tail entity. Identifying unstructured data as structured data and storing it in a set is convenient for subsequent analysis, query, and reasoning.

[0076] For example: Suppose there is the following operating data and fault feedback data: "The automatic sprinkler system put into production on June 1, 2024, with the outlet pressure of the fire pump being 0.8 MPa, the outlet pressure of the pressure reducing valve being 0.5 MPa, and the network pressure being 1 MPa." And the fault feedback data: "On August 2, 2024, after the alarm, the signal valve closed but the signal display was normal, and the monitoring and stabilizing system started and stopped." From this, the following triples can be extracted: EAV triples: (fire pump, outlet pressure, 0.8 MPa), (pressure reducing valve, outlet pressure, 0.5 MPa), and ERE triples: (monitoring and stabilizing system, cause of fault, signal valve closed).

[0077] 2.2) Based on the UIE framework, the T5 large language model, and the constructed fault knowledge extraction model architecture of the fire water supply system, construct a fault knowledge extraction model for the fire water supply system.

[0078] Specifically, to efficiently extract the fault knowledge triples of the fire water supply system from a large amount of text data, the embodiments of the present invention are based on the UIE framework and the T5 large language model to perform fault knowledge extraction of the fire water supply system and establish a fault knowledge extraction model for the fire water supply system. Among them, the construction of the UIE framework includes constructing an SSI (Structure Schema Instructor) and an SEL (Structured Extraction Language). The SSI uses a pattern-based prompt, which is concatenated with the text sequence as the input; the SEL has a unified representation form as the final output, aiming to represent the information generated by different information extraction (IE) tasks in one sequence. The embodiments of the present invention implement the process from the SSI and the text sequence to the SEL, using the T5 large language model, which can regard the input and output as text sequences and flexibly process various tasks.

[0079] Specifically, the core idea of the fault knowledge extraction model for the fire water supply system constructed by the embodiments of the present invention is: generally model different information extraction tasks, adaptively generate the target structure, and collaboratively learn the general information extraction ability from different knowledge sources. More specifically, the fault knowledge extraction model for the fire water supply system includes three layers. Among them, the first layer is the SSI structure guiding layer, which is used to control the type of information generated during the extraction process based on the schema mode (database mode) extraction mechanism, create a pattern-based prompt, that is, the SSI sequence, and perform feature concatenation with the text sequence as the input of the next layer; the second layer is the encoding and decoding layer based on T5, which is used to encode the sequence after the SSI sequence and the text sequence are concatenated into a linear SEL (the role of the decoder), and then convert it into a text sequence in the form of decoding (the role of the encoder); the third layer is the SEL structure extraction layer, which is used to locate and associate different IE structures, re-encode the decoded information, and convert it into the information to be extracted, where:

[0080] ①SSI (Structure Schema Instructor) structure guiding layer

[0081] Specifically, to enable the model to generate a unified structure for different IE tasks, for example, given a sentence: "On August 2, 2024, after the alarm, the signal valve closed but the signal showed normal, and the monitoring and voltage stabilizing system started and stopped. It is expected to generate 3 entity structures (fault location: signal valve, fault cause: closed, time: August 2, 2024) and an event structure (fault location: signal valve, fault symptom: signal showed normal, fault consequence: monitoring and voltage stabilizing system started and stopped). To achieve this goal, an SSI structure guiding layer is adopted. The SSI structure guiding layer is an extraction mechanism based on the schema (database) mode, which controls what to locate, associate, and generate in the source text in UIE, can adaptively control the type of information generated during the extraction process, creates a pattern-based prompt, and uses it as a prefix during the generation process. This prompt includes three types of fragments:

[0082] Entity name: The name of the information to be located in a specific information extraction task, such as "fault location" in the named entity recognition task.

[0083] Association name: The association between entity names, such as "fault consequence" in relation extraction.

[0084] Special characters: [spot], [asso], [text]. These names are added before the entity name, association name, and input text. [spot] is the entity name (type) extracted from the text sequence, and [asso] is the relationship extracted from the text sequence.

[0085] Specifically, all tokens (the smallest unit for processing text) in the SSI structure guiding layer will be concatenated and placed before the original sequence, that is:

[0086] s⊕x = [s1, s2, …, s |s| , x1, x2, …, x |x| = [[spot], …, [spot], …, [asso], …, [asso], …, [text], x1, x2, …, x |x| (2)

[0087] where x = [x1, x2, …, x |x| is the text sequence; x |x| is a segment of the text sequence or a word, phrase, sentence, etc. within it; s = [s1, s2, …, s |s| is the structure pattern guide, and s |s| is a combination of special characters and entity names, a combination of special characters and association names, or the single special character [text] (usually as the last unit of the SSI).

[0088] ② T5-based encoding and decoding layer

[0089] Specifically, with the given structure pattern guide (s) and text sequence (x) as inputs, the model extracts target information by generating a linearized SEL. The present invention uses an encoder-decoder architecture to perform the process of text-structure extraction language. Given a text sequence x and a structure pattern guide s, the model first calculates the hidden representation H of each token:

[0090] H = Encoder(s1,…,s |s| ,x1,…,x |x| ) (3)

[0091] where Encoder(·) is a Transformer encoder, including structures such as input encoding, multi-head attention mechanism, residual connection and normalization, and feed-forward neural network, encodes the input text sequence x and structure pattern guide s, and then outputs to the T5-based decoder layer. The model decodes the input into a linearized SEL in an autoregressive manner. At the i-th step of decoding, the model generates the i-th token y in the SEL sequence i , and decodes the state

[0092]

[0093] where Decoder(·) is a Transformer decoder, including structures such as input encoding, causal multi-head attention mechanism, residual connection and normalization, multi-head attention mechanism, and feed-forward neural network, continuously predicts the conditional probability p(y i |y <i ,x,s) during the decoding process to ensure the accuracy of the result.

[0094] Then, when the output is <eos>When (End of Sequence) is reached, it serves as a marker for the end of the sequence, indicating that Decoder(·) has completed the prediction.

[0095] ③ SEL Structure Extraction Layer

[0096] Specifically, the SEL structure extraction layer re-encodes the decoded information by locating and associating different IE structures, and converts it into the information to be extracted.

[0097] Each SEL expression includes three types of semantic units:

[0098] Entity Name: Represents the type of information to be located in the original text.

[0099] Association Name: Represents the specific type of association to be extracted from the original text.

[0100] Position and Content of Information Fragment: Represents the text fragment to be located or associated in the original text, and uses ":" to represent the mapping from the position and content of the information fragment to its tag or association name. And two structure indicators "(" and ")" are used to form a hierarchical structure between the extracted information, which is expressed as:

[0101] y = UIE(s ⊕ x) (5)

[0102] where y = [y1, y2, …, y |y| is a SEL sequence, and y |y| is a unit in the SEL sequence, expressed as Entity Name or Association Name: Corresponding Text Fragment.

[0103] 3) Train the constructed fire water supply system fault knowledge extraction model to obtain a trained fire water supply system fault knowledge extraction model.

[0104] Specifically, to enable the fire water supply system fault knowledge extraction model to adapt to different knowledge extraction tasks, it needs to be pre-trained and fine-tuned. The fire water supply system fault knowledge extraction model needs to encode text, map text to structures, and decode valid structures, specifically as follows:

[0105] 3.1) Obtain a large amount of general text data.

[0106] Specifically, the obtained general text data is divided into three parts: text-structure parallel corpus D Pair , unstructured text data set (a large number of ordinary texts) D Text and structured data set D Record . Among them, the text-structure parallel corpus D Pair Each instance of is a parallel pair, namely the token sequence a and the structured record b; the structured dataset D Record Each instance of is a structured record b.

[0107] 3.2) Based on the obtained general text data, pre-train the fire water supply system fault knowledge extraction model:

[0108] 3.2.1) Use the text-structure parallel corpus D Pair Pre-train the fire water supply system fault knowledge extraction model. To enable the pre-training of the fire water supply system fault knowledge extraction model to capture the basic text-to-structure mapping ability, in this embodiment, D Pair ={(a,b)} is used for the fire water supply system fault knowledge extraction model:

[0109] 3.2.1.1) For a given parallel sample (a,b), extract the type s to be located + and the positive associated type s in the structured record b a+ , and form the positive sample s + =s s+ ∪s a+ , where s s+ is the positive location type extracted from the structured record b.

[0110] 3.2.1.2) Considering the singularity of training the model only with positive samples, construct the sampled negative location type s s- and the negative associated type s in the structured record b a- , and splice out the final dataset s meta =s + ∪s s- ∪s a- .

[0111] 3.2.1.3) Based on the obtained dataset s meta , determine the text-to-structure pre-training objective function as:

[0112]

[0113] where θ e and θ d are the parameters of the encoder and decoder respectively; p is the conditional probability.

[0114] 3.2.2) Use the unstructured text dataset D Text to improve the semantic representation of the fire water supply system knowledge extraction model.

[0115] Specifically, during the text-to-structure pre-training process, to effectively alleviate the semantic forgetting of the token segment entity name and the associated name E, the present invention adopts the pre-training objective of MLM (Masked Language Modeling), which can train the model to understand and generate natural language by randomly masking a part of the tokens in the input sequence and letting the model predict.

[0116] Specifically, on the unstructured text dataset D Text the MLM pre-training is adopted to extract the knowledge model of the fire water supply system, thereby improving the semantic representation ability of the fire water supply system knowledge extraction model:

[0117]

[0118] Among them, is the pre-training loss function on the dataset D Text ; a ′ is the original text of the masked part of the token units; a″ is the masked token unit segment.

[0119] 3.2.3) Use the structured dataset D Record for structured generation pre-training.

[0120] Specifically, to enable the fire water supply system fault knowledge extraction model to capture the generation ability for the SEL structure extraction layer and the schema (database schema) defined structure, use the structured dataset D Record to pre-train the fire water supply system fault knowledge extraction model.

[0121] Specifically, take the decoder of the fire water supply system fault knowledge extraction model as a structured language model, and take each record in the structured dataset D Record as an SEL expression, then:

[0122]

[0123] Among them, is the pre-training loss function on the dataset D Record ; b <i is all the elements before the i-th element in the structured record b.

[0124] 3.2.4) Combine the tasks of text-to-structure pre-training, improving the semantic representation of the fire water supply system knowledge extraction model, and structured generation pre-training to obtain the objective function of the fire water supply system fault knowledge extraction model.

[0125] Specifically, by pre-training for structured generation, the decoder of the fire water supply system fault knowledge extraction model can capture the regularity of SEL and the interaction between different tags, combine all the above tasks, and obtain the objective function of the fire water supply system fault knowledge extraction model

[0126]

[0127] Specifically, in actual operation, each type of corpus is represented as a triple. For the text-structure parallel corpus D Pair in the data (d, b), construct the (s, a, b) triple by sampling the meta-schema for each text-record; for the unstructured text dataset D Text in the text data a, construct the (None, a′, a″) triple; for the structured dataset D Record in the structured record b, construct the (None, None, b) input triple.

[0128] 3.2.5) For the pre-trained fire water supply system fault knowledge extraction model, perform domain adaptation training on fault structured and semi-structured data such as standard guide data and abnormal alarm log data including fire water supply system knowledge.

[0129] 4) For the fire water supply system knowledge extraction task, fine-tune the trained fire water supply system fault knowledge extraction model, and the model can understand and learn the historical data patterns and features of fire water supply system faults.

[0130] Specifically, after the pre-training of the fire water supply system fault knowledge extraction model is completed, fine-tuning can make it adapt to a specific fire water supply system knowledge extraction task:

[0131] 4.1) Based on the fault knowledge ontology of the fire water supply system, given a labeled fire water supply system fault data D task =(s, a, b), adopt the cross-entropy loss function to perform preliminary fine-tuning on the fire water supply system fault knowledge extraction model:

[0132]

[0133] 4.2) To alleviate the exposure bias in the decoding process of the autoregressive model, adopt the Rejection Mechanism to perform secondary fine-tuning on the preliminarily fine-tuned fire water supply system fault knowledge extraction model:

[0134] 4.2.1) Given an instance (s, a, b), encode the structured record b using SEL.

[0135] 4.2.2) With probability p ∈ Randomly insert [NULL] (special character) into SPOTNAME (entity name) and ASSONAME (associated name) to construct negative samples (SPOTNAME, [NULL]) and (ASSONAME, [NULL]), so as to achieve the purpose of not making a decision to avoid misclassification and rejecting false generation when facing uncertain or difficult-to-classify data.

[0136] 5) Adopt the fine-tuned fire water supply system fault knowledge extraction model to extract fault knowledge from the fire water supply system fault knowledge corpus.

[0137] 6) Integrate and process the extracted fault knowledge to meet the standards for constructing a knowledge graph, and store it in the Neo4j graph database to form a fault knowledge graph of the fire water supply system, specifically:

[0138] 6.1) It is difficult to ensure the accuracy and standardization of the fire water supply system fault data filled in manually, and the fault knowledge initially obtained through deep learning methods for knowledge extraction is usually uneven, and there may be cases of multiple meanings for one word or multiple words with the same meaning, resulting in deviations or repetitions in the extracted triples. Therefore, in this embodiment, the extracted fault knowledge is integrated and processed to meet the standards for constructing a knowledge graph.

[0139] Specifically, when calculating the similarity, considering that the merging of fire water supply system fault data is essentially clustering according to its relevance, the present invention uses a greedy algorithm to solve:

[0140] 6.1.1) Assume that m and n are two entities to be merged, and the goal of the similarity between m and n is to find a clustering scheme with the minimum cost:

[0141]

[0142] Among them, P mn represents the probability that m and n are in the same class; r mn represents that m and n are assigned to the same class; represents the cost of cutting the edge between m and n; represents the cost of retaining the edge between x and y.

[0143] 6.1.2) Repeat the above step 6.1.1) to obtain the results of entity disambiguation for all entities.

[0144] Specifically, as Figure 4 shown, the solid line indicates that there is a relationship between two entities. After classifying them into the same class, it will incur a cost to the final result The dotted line indicates that there is no relationship between the two entities. Classifying them into one category will incur a cost for the final result. As can be seen from the figure, for those with a higher similarity, the probability of being truncated is lower, and thus they are retained (entity merging); for those with a lower similarity, the probability of being retained is lower.

[0145] 6.2) By performing entity disambiguation on each triple relationship extracted through the above method, the fault knowledge after fusion processing is obtained and stored in the Neo4j graph database, thus completing the construction of the fault knowledge graph of the fire water supply system.

[0146] Embodiment 2

[0147] This embodiment provides a fire water supply system fault knowledge extraction system, including:

[0148] A fine-tuning module for fine-tuning the trained fire water supply system fault knowledge extraction model for the fire water supply system knowledge extraction task, where the fire water supply system fault knowledge extraction model is constructed based on the UIE framework and the T5 large language model.

[0149] A fault knowledge extraction module for extracting fault knowledge from the fire water supply system fault knowledge corpus using the fine-tuned fire water supply system fault knowledge extraction model.

[0150] A fusion processing module for fusing and processing the extracted fault knowledge to meet the standards for constructing a knowledge graph and storing it to form a fault knowledge graph of the fire water supply system.

[0151] The system provided in this embodiment is used to execute the above method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.

[0152] Embodiment 3

[0153] This embodiment provides a processing device corresponding to the fire water supply system fault knowledge extraction method provided in Embodiment 1. The processing device can be a processing device applicable to a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of Embodiment 1.

[0154] The processing device includes a processor, a memory, a communication interface, and a bus. The processor, the memory, and the communication interface are connected through the bus to complete communication with each other. The memory stores a computer program that can run on the processing device. When the processing device runs the computer program, it executes the fire water supply system fault knowledge extraction method provided in this Embodiment 1.

[0155] In some implementations, the memory may be a high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0156] In other implementations, the processor may be various types of general-purpose processors such as a central processing unit (CPU) or a digital signal processor (DSP), which are not limited herein.

[0157] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0158] Those skilled in the art can understand that the structure of the above-mentioned computing device is only a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computing device to which the solution of the present invention is applied. The specific computing device may include more or fewer components, or combine certain components, or have different component arrangements.

[0159] Embodiment 4

[0160] This embodiment provides a computer program product corresponding to the fire water supply system fault knowledge extraction method provided in Embodiment 1. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing the fire water supply system fault knowledge extraction method described in Embodiment 1 are uploaded.

[0161] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.

[0162] The computer-readable storage medium provided in the above-mentioned embodiment has the same implementation principle and technical effects as the above method embodiment, and will not be elaborated herein.

[0163] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows and / or one or more blocks in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0164] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more flows and / or one or more blocks in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0166] The above embodiments are only used to illustrate the present invention. The structures, connection methods, manufacturing processes, etc. of each component can be changed. Any equivalent transformation and improvement made on the basis of the technical solution of the present invention should not be excluded from the protection scope of the present invention.< / eos>

Claims

1. A method for extracting fault knowledge of a fire water supply system, characterized in that, Including: For the knowledge extraction task of the fire water supply system, fine-tune the trained fire water supply system fault knowledge extraction model, where the fire water supply system fault knowledge extraction model is constructed based on the UIE framework and the T5 large language model; Use the fine-tuned fire water supply system fault knowledge extraction model to extract fault knowledge from the fire water supply system fault knowledge anticipation; Fuse and process the extracted fault knowledge to meet the standards for constructing a knowledge graph, and store it to form a fault knowledge graph of the fire water supply system.

2. The method for extracting fault knowledge of a fire water supply system according to claim 1, wherein, The construction process of the fire water supply system fault knowledge extraction model is as follows: Collect and preprocess the fire water supply system fault data, and construct the fault knowledge ontology of the fire water supply system; Based on the UIE framework and the T5 large language model, construct a fire water supply system fault knowledge extraction model; Train the constructed fire water supply system fault knowledge extraction model to obtain a trained fire water supply system fault knowledge extraction model.

3. The method for extracting fault knowledge of a fire water supply system according to claim 2, characterized in that, The collection and preprocessing of the fire water supply system fault data to construct the fault knowledge ontology of the fire water supply system includes: Collect and preprocess the fire water supply system fault data; Based on the preprocessed fire water supply system fault data, model the key components of the fire water supply system, and establish the equipment entities in the fault knowledge ontology of the fire water supply system, where the key components of the fire water supply system include the hierarchy and structural relationship of the fire water supply system; Based on the preprocessed fire water supply system fault data, consider the logical relationship between the fire water supply system fault symptoms, fault causes, and solutions, and establish the fault entities in the fault knowledge ontology of the fire water supply system.

4. The method for extracting fault knowledge of a fire water supply system according to claim 3, wherein, The construction of the fire water supply system fault knowledge extraction model based on the UIE framework and the T5 large language model includes: Taking the constructed fault knowledge ontology as a specification, construct the architecture of the fire water supply system fault knowledge extraction model; Based on the UIE framework, the T5 large language model, and the constructed architecture of the fire water supply system fault knowledge extraction model, construct the fire water supply system fault knowledge extraction model.

5. The method for extracting fault knowledge of a fire water supply system according to claim 4, wherein, The fire water supply system fault knowledge extraction model includes an SSI structure guidance layer, a T5-based encoding and decoding layer, and an SEL structure extraction layer; The SSI structure guidance layer is used to control the type of information generated during the extraction process based on the schema-based extraction mechanism, create an SSI sequence, and splice the features with the text sequence as the input of the next layer; The T5-based encoding and decoding layer is used to encode the sequence spliced by the SSI sequence and the text sequence into a linear SEL, and then transform it into a text sequence through decoding; The SEL structure extraction layer is used to locate and associate different IE structures, re-encode the decoded information, and transform it into the information to be extracted.

6. The method for extracting fault knowledge of a fire water supply system according to claim 2, characterized in that, The training of the constructed fire water supply system fault knowledge extraction model to obtain a trained fire water supply system fault knowledge extraction model includes: Obtain a number of general text data; Based on the obtained number of general text data, pre-train the fire water supply system fault knowledge extraction model: Divide the general text data into text-structured parallel corpus D Pair , unstructured text data set D Text , structured data set D Record ; Adopt text-structure parallel corpus D Pair Pre-train the fault knowledge extraction model for the fire water supply system; Adopt the unstructured text dataset D Text Improve the semantic representation of the knowledge extraction model for the fire water supply system; Adopt a structured dataset D record Perform structured generation pre-training; Merge the tasks of text-to-structure pre-training, improving the semantic representation of the knowledge extraction model for the fire water supply system, and structured generation pre-training to obtain the objective function of the fire water supply system fault knowledge extraction model; For the pre-trained fire water supply system fault knowledge extraction model, perform domain adaptation training on fault-structured and semi-structured data including standard guide data and abnormal alarm log data of the fire water supply system knowledge.

7. The method for extracting fault knowledge of a fire water supply system according to claim 2, wherein, For the fire water supply system knowledge extraction task, fine-tune the trained fire water supply system fault knowledge extraction model, and the model can understand and learn the historical data patterns and characteristics of the fire water supply system faults, including: Based on the fault knowledge ontology of the fire water supply system, given a labeled fire water supply system fault data, use the cross-entropy loss function to perform preliminary fine-tuning on the fire water supply system fault knowledge extraction model; Adopt a rejection mechanism to perform secondary fine-tuning on the fire water supply system fault knowledge extraction model after preliminary fine-tuning.

8. A fault knowledge extraction system for a fire water supply system, characterized in that, Including: A fine-tuning module for fine-tuning the trained fire water supply system fault knowledge extraction model for the fire water supply system knowledge extraction task, where the fire water supply system fault knowledge extraction model is constructed based on the UIE framework and the T5 large language model; A fault knowledge extraction module for extracting fault knowledge from the fire water supply system fault knowledge corpus using the fine-tuned fire water supply system fault knowledge extraction model; A fusion processing module for fusing and processing the extracted fault knowledge to meet the standards for constructing a knowledge graph and storing it to form a fault knowledge graph of the fire water supply system.

9. A processing device, characterized in that, Including computer program instructions, where the computer program instructions, when executed by a processing device, are used to implement the steps corresponding to the fire water supply system fault knowledge extraction method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores computer program instructions, where the computer program instructions, when executed by a processor, are used to implement the steps corresponding to the fire water supply system fault knowledge extraction method described in any one of claims 1-7.