Electric power engineering knowledge base construction method and system based on large model

By building a large model-based power engineering knowledge base, the problem of difficult management and dissemination of complex power engineering knowledge is solved, efficient knowledge extraction and query is achieved, and the query efficiency and effectiveness of the knowledge base are improved.

CN120216703APending Publication Date: 2025-06-27GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510281374.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively manage and disseminate complex power engineering knowledge, resulting in difficulty in understanding and data omissions in the knowledge extraction process, which in turn affects the efficiency of knowledge query and the effectiveness of answers.

Method used

Using a large model-based method, we obtain the original knowledge data and professional term database of power engineering, conduct model training and entity extraction and annotation, build a multimodal knowledge graph, and fuse it with the pre-trained large model to form a power engineering knowledge base.

Benefits of technology

It improves the quality and completeness of knowledge extraction, realizes multimodal fusion of knowledge, enhances the query efficiency of the knowledge base and the effectiveness of answers, and can respond more accurately to natural language queries and conduct knowledge inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216703A_ABST
    Figure CN120216703A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power engineering knowledge base construction method based on a large model. The method comprises the following steps: acquiring electric power engineering original knowledge data and an electric power engineering terminology library; the original knowledge data comprises text data and image and drawing data; performing model training on the pre-training language model based on the text data and an electric power engineering terminology library to obtain a language model; performing naming recognition on the text data based on a language model to obtain a plurality of entities in the electric power engineering, and performing extraction labeling on relationships and attributes among the entities to obtain entity associated data; constructing a multi-modal knowledge graph based on the image and drawing data and the entity association data; and performing model training on the pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and performing fusion storage on the large model and the multi-modal knowledge graph to obtain the electric power engineering knowledge base. According to the method, the query efficiency of the knowledge base in practical application and the effectiveness of answers are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power engineering information technology, and particularly to a method and system for constructing a power engineering knowledge base based on a large model. Background Art

[0002] With the continuous expansion of the complexity and scale of the power system, the knowledge in the field of power engineering has become increasingly rich and complex. Traditional knowledge management and dissemination methods are no longer sufficient to meet the needs of modern power engineering. The method for constructing a power engineering knowledge base based on a large model can effectively extract, organize, and store knowledge by using large-scale pre-trained models (such as BERT, GPT, etc.) to process and organize a large amount of text data such as power engineering literature, manuals, and technical reports, providing convenient knowledge retrieval and application services for engineers and technicians.

[0003] Currently, in the process of power engineering services, users mainly obtain the required knowledge through question-and-answer queries. Since power engineering texts often contain a large number of professional terms and complex sentence structures, it may cause the knowledge extraction process to fail to understand relevant data, resulting in data omission, and thus problems such as low query efficiency and inability to obtain effective answers during the knowledge query process. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art and propose a method and system for constructing a power engineering knowledge base based on a large model.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A method for constructing a power engineering knowledge base based on a large model includes the following steps:

[0007] Obtain the original power engineering knowledge data and the power engineering professional term library; the original knowledge data includes text data and image and drawing data;

[0008] Train a pre-trained language model based on the text data and the power engineering professional term library to obtain a language model;

[0009] Perform named entity recognition on the text data based on the language model to obtain multiple entities in power engineering, and extract and label the relationships and attributes between entities to obtain entity association data;

[0010] Construct a multi-modal knowledge graph based on the image and drawing data and the entity association data;

[0011] Training the pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and fusing and storing the large model with the multi-modal knowledge graph to obtain a power engineering knowledge base.

[0012] According to a method for constructing a power engineering knowledge base based on a large model provided by the present invention, the text data includes power equipment data, power system operation data, power engineering design data, power engineering construction data, and power engineering maintenance data.

[0013] According to a method for constructing a power engineering knowledge base based on a large model provided by the present invention, performing named entity recognition on the text data based on the language model to obtain multiple entities in power engineering, including:

[0014] Splitting based on the syntactic structure of the text data to obtain multiple sub-data; each sub-data is a sequence of words or phrases;

[0015] Converting each sub-data to obtain word vectors;

[0016] Encoding the word vectors based on the encoding layer of the language model to obtain encoded data;

[0017] Extracting features from the encoded data based on the language model to obtain encoded feature data;

[0018] Classifying the multiple sub-data based on the encoded feature data to obtain multiple entities in power engineering.

[0019] According to a method for constructing a power engineering knowledge base based on a large model provided by the present invention, extracting and annotating entity relationships and attributes to obtain entity association data, including:

[0020] Based on the association relationships between multiple entities in the power engineering, obtaining multiple entity pairs;

[0021] Extracting the text data based on the relationship features between the entity pairs to obtain relationship feature data;

[0022] Classifying and annotating the relationships between the entity pairs based on the relationship feature data to obtain entity relationship data;

[0023] Obtaining keywords or phrases related to entity attributes in the text data to obtain entity attribute data;

[0024] Based on the entity attribute data, determining entity attribute values, and associating and annotating the entity attribute values with the entity relationship data to obtain entity association data.

[0025] A method for constructing a power engineering knowledge base based on a large model provided by the present invention, which constructs a multi-modal knowledge graph based on image and drawing data and entity association data, includes:

[0026] Extract features from the image data to obtain first feature information;

[0027] Perform recognition and conversion processing on the drawing data to obtain second feature information;

[0028] Construct a basic framework of the knowledge graph based on entity association data;

[0029] Associate the first feature information, the second feature information and the basic framework of the knowledge graph to obtain a multi-modal knowledge graph.

[0030] A method for constructing a power engineering knowledge base based on a large model provided by the present invention, which trains a pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, includes:

[0031] Embed and fuse the multi-modal knowledge graph into the input layer of the pre-trained large model to obtain a fused large model framework;

[0032] Obtain power engineering knowledge answering tasks, knowledge reasoning tasks and knowledge completion tasks in the text data;

[0033] Use the power engineering knowledge answering tasks, knowledge reasoning tasks and knowledge completion tasks as training objectives to train the fused large model framework to obtain a large model.

[0034] A method for constructing a power engineering knowledge base based on a large model provided by the present invention, the loss function of the obtained large model is:

[0035] L ZO = L qa + L kr + L ki ;

[0036]

[0037] Among them, N qa represents the number of samples of the knowledge answering task; V represents the size of the answer vocabulary; y i,j represents the value of the one-hot encoding of the true answer at the j-th vocabulary position for the i-th knowledge answering sample; represents the value of the probability distribution of the model-predicted answer at the j-th vocabulary for the i-th knowledge answering sample; q i represents the question semantic vector generated by the model for the i-th knowledge answering question; denote the expected problem semantic vectors for the i-th problem obtained from knowledge graph analysis; γ denotes the weight coefficient; N kr denote the number of samples for the knowledge inference task; denote the relational representation form inferred by the model for the k-th knowledge inference sample; r k denote the true relational representation form for the k-th knowledge inference sample; N ki denote the number of samples for the knowledge completion task; denote the knowledge representation after completion by the model for the m-th knowledge completion task sample; k m denote the true and complete knowledge representation for the m-th knowledge completion task sample; R m denote the degree to which the knowledge completed by the model for the m-th knowledge completion task sample violates the rules or common sense of power engineering; v denotes the weight coefficient used to balance the two losses of completion accuracy loss and completion rationality loss in the knowledge completion task loss function.

[0038] A power engineering knowledge base construction system based on a large model, comprising:

[0039] An acquisition unit: used to acquire original power engineering knowledge data and a power engineering professional term library; the original knowledge data includes text data and image and drawing data;

[0040] A first training unit: used to perform model training on a pre-trained language model based on the text data and the power engineering professional term library to obtain a language model;

[0041] An entity extraction and annotation unit: used to perform named entity recognition on the text data based on the language model to obtain multiple entities in power engineering, and extract and annotate the relationships and attributes between entities to obtain entity association data;

[0042] A knowledge graph construction unit: used to construct a multi-modal knowledge graph based on the image and drawing data and the entity association data;

[0043] A storage unit: used to perform model training on a pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and fuse and store the large model with the multi-modal knowledge graph to obtain a power engineering knowledge base.

[0044] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the above-mentioned method for constructing a power engineering knowledge base based on a large model are implemented.

[0045] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for constructing a power engineering knowledge base based on a large model are implemented.

[0046] The present invention has the following advantages compared with the prior art:

[0047] A method and system for constructing a power engineering knowledge base based on a large model provided by the present invention first obtain a power engineering professional term library and train a first model with text data. By means of the professional term library, the first model is trained to accurately understand and process these complex text contents, improving the quality and integrity of knowledge extraction, and trying to avoid the problems that power engineering texts are rich in a large number of professional terms and have complex sentence structures, resulting in difficulties in understanding and data omission in the knowledge extraction process. In addition, by constructing a multi-modal knowledge graph from image and drawing data and integrating knowledge of different modalities, various types of knowledge in the power engineering field can be presented more completely, covering power engineering-related information in all aspects, from literal descriptions such as equipment parameters and operating principles to intuitive displays such as equipment appearance and system layout, solving the problem of single and incomplete knowledge representation. Finally, through multi-modal fusion and staged model training, etc., the constructed knowledge base can respond more accurately to natural language queries, reason and answer based on comprehensive and accurate knowledge, effectively improving the query efficiency and the effectiveness of answers of the knowledge base in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a schematic flowchart of the method for constructing a power engineering knowledge base based on a large model provided by the embodiment of the present invention;

[0050] Figure 2 It is a schematic structural diagram of the system for constructing a power engineering knowledge base based on a large model provided by the embodiment of the present invention;

[0051] Figure 3 It is a schematic structural diagram of the electronic device proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0053] The following will describe Figure 1 - Figure 3 a method and system for constructing a power engineering knowledge base based on a large model of the present invention.

[0054] Figure 1 is a schematic flowchart of a method for constructing a power engineering knowledge base based on a large model provided by the present invention. As Figure 1 shown, the method includes:

[0055] Step 101, obtaining original power engineering knowledge data and a power engineering professional term library; the original knowledge data includes text data, as well as image and drawing data. The text data includes power equipment data, power system operation data, power engineering design data, power engineering construction data, and power engineering maintenance data. By obtaining the original knowledge data and the power engineering professional term library, rich materials and a professional knowledge foundation can be provided for subsequent model training, ensuring that the model can understand and process specific terms and concepts in the power engineering field, and improving the accuracy of subsequent knowledge extraction and processing.

[0056] Specifically, for the acquisition of text data in power engineering, on the one hand, various project documents can be extracted from the internal document management system of power engineering enterprises, including feasibility study reports, preliminary design documents, detailed design drawing descriptions, construction plans, completion acceptance reports, equipment operation manuals, maintenance manuals, and accident analysis reports of power engineering. The above-mentioned documents cover information on all stages of power engineering from planning to operation, including professional knowledge, technical details, actual cases, and lessons learned.

[0057] On the other hand, relevant text materials can also be collected from professional literature databases, academic journals, and technical standard specification documents in the power industry. And public knowledge information can be collected from platforms such as professional websites, forums, and social media groups related to power engineering by web crawler technology.

[0058] For the acquisition of image and drawing data in power engineering, on the one hand, various design drawings of the power system can be obtained from power engineering design units, such as electrical main connection diagrams, substation layout plans, transmission line route maps, equipment installation drawings, etc. In a graphical way, it can accurately display key information such as the topological structure, equipment layout, and connection relationships of the power engineering system. On the other hand, image materials at the construction site of power engineering can also be collected, including pictures and video materials at different stages, etc.

[0059] In the process of constructing a power engineering terminology database, various power engineering text materials collected can be used for term extraction and sorting. The terms are classified separately to establish a preliminary term classification system. Then, referring to professional dictionaries, industry standards and specifications, and authoritative textbooks in power engineering, the definitions, explanations, and usages of the terms are sorted out and standardized in detail to ensure that each term has an accurate, clear, and unified definition and explanation, avoiding confusion caused by differences in the expressions of materials from different sources. Finally, the association relationships between terms are established to construct a term semantic network. Analyze the semantic logical relationships between terms, such as hyponymy relationships ("substation" is a hyponym of "power facility"), synonymy relationships ("circuit breaker" and "switch" have the same meaning in some contexts), antonymy relationships ("high voltage" and "low voltage"), and component relationships ("transformer" is composed of "iron core", "winding", "insulating oil", etc.), and store these relationships in the terminology database in a structured manner. Such a term semantic network helps to better understand and process the semantic connotations of power engineering knowledge during the construction and application of the knowledge base, improving the accuracy of knowledge reasoning and question answering.

[0060] Step 102: Train the pre-trained language model based on the text data and the power engineering terminology database to obtain a language model.

[0061] Specifically, before training the model, the text data also needs to be preprocessed and annotated. During the cleaning and preprocessing of the collected text data, noise information in the text is removed, such as redundant punctuation marks, special characters, repeated blank characters, etc. At the same time, the text is standardized, for example, unifying the representation of numbers, unit conversions, etc. When annotating the text data, the power engineering terminology database is used for annotation. According to the term list and definitions in the terminology database, the terms in the text are initially annotated.

[0062] In addition, a pre-trained language model based on the Transformer architecture (such as BERT, GPT, etc.) is selected. Since such models are pre-trained on large-scale general corpora and have powerful language understanding and generation capabilities, they can quickly adapt to text processing tasks in the field of power engineering. And according to the characteristics of power engineering text data and task requirements, the architecture of the pre-trained model is appropriately adjusted.

[0063] Furthermore, during the model training process, taking the power engineering knowledge extraction tasks as the goal, including tasks such as named entity recognition, relation extraction, and attribute extraction, the model is trained. In the named entity recognition task, the model is trained to recognize various entities in power engineering texts, such as substation names, transformer models, transmission line numbers, etc.; in the relation extraction task, the model is trained to recognize the relationships between entities, such as the "contains" relationship between a substation and a transformer, the "connection" relationship between a transmission line and a substation, etc.; in the attribute extraction task, the model is trained to extract the attribute information of entities, such as the rated capacity, voltage level, production date, etc. of a transformer. Thus, the language model can be applied in subsequent named entity recognition, relation extraction, and attribute extraction.

[0064] Step 103: Perform named recognition on the text data based on the language model to obtain multiple entities in power engineering, and extract and label the relationships and attributes between the entities to obtain entity association data.

[0065] Specifically, mainly by using the trained language model to process the text data, entities such as substations, transformers, and transmission lines are recognized, and the relationships (such as connection relationships, subordination relationships, etc.) and attributes (such as equipment parameters, production dates, etc.) between the entities are extracted to form entity association data. The structured processing of the text data is realized, and the unstructured text is transformed into an organized entity relationship network, which is convenient for subsequent knowledge graph construction and knowledge reasoning.

[0066] Correspondingly, perform named recognition on the text data based on the language model to obtain multiple entities in power engineering, including: splitting according to the grammatical structure of the text data to obtain multiple sub-data; each sub-data is a sequence of words or phrases; converting each sub-data to obtain word vectors; encoding the word vectors based on the encoding layer of the language model to obtain encoded data; extracting features from the encoded data based on the language model to obtain encoded feature data; classifying the multiple sub-data based on the encoded feature data to obtain multiple entities in power engineering.

[0067] Furthermore, extract and label the relationships and attributes between entities to obtain entity association data, including: obtaining multiple entity pairs based on the association relationships between multiple entities in the power engineering; extracting text data based on the relationship features between entity pairs to obtain relationship feature data; classifying and labeling the relationships between entity pairs based on the relationship feature data to obtain entity relationship data; obtaining keywords or phrases related to entity attributes in the text data to obtain entity attribute data; determining entity attribute values based on the entity attribute data, and associatively labeling the entity attribute values with the entity relationship data to obtain entity association data.

[0068] Step 104: Construct a multimodal knowledge graph based on the image and drawing data and the entity association data.

[0069] Specifically, use image and drawing data processing technologies (such as image recognition, graphic parsing, etc.) to extract visual information, combine it with the entity association data, and construct a multimodal knowledge graph with entities as nodes and relationships as edges, so that the knowledge graph contains text, image, and drawing information in multiple aspects. It can integrate multimodal information, provide a more comprehensive and intuitive knowledge representation, enhance the expression ability of the knowledge base for power engineering knowledge, and contribute to the understanding and application of complex knowledge.

[0070] In addition, it specifically includes: extracting feature information from the image data to obtain the first feature information; performing recognition and conversion processing on the drawing data to obtain the second feature information; constructing the basic framework of the knowledge graph based on the entity association data; associating the first feature information, the second feature information, and the basic framework of the knowledge graph to obtain a multimodal knowledge graph.

[0071] Step 105: Train a pre-trained large model based on the text data and the multimodal knowledge graph to obtain a large model, and fuse and store the large model with the multimodal knowledge graph to obtain a power engineering knowledge base.

[0072] Specifically, input the text data and the multimodal knowledge graph into the pre-trained large model, and use tasks such as power engineering knowledge answering, knowledge reasoning, and knowledge completion as training objectives to perform multiple rounds of training to optimize the model parameters, so that the large model can understand and process multimodal knowledge. Finally, fuse and store the large model with the multimodal knowledge graph to form a power engineering knowledge base, and knowledge queries and applications can be performed through the large model interface. Thus, an intelligent and efficient power engineering knowledge base is constructed, which can quickly respond to users' knowledge needs, provide accurate and comprehensive knowledge services, and can perform knowledge reasoning and innovation to assist power engineering decision-making and problem-solving.

[0073] Meanwhile, the specific steps to obtain the large model include: embedding and fusing the multi-modal knowledge graph into the input layer of the pre-trained large model to obtain the fused large model framework; obtaining the power engineering knowledge Q&A tasks, knowledge reasoning tasks, and knowledge completion tasks in the text data; and training the fused large model framework with the power engineering knowledge Q&A tasks, knowledge reasoning tasks, and knowledge completion tasks as the training objectives to obtain the large model.

[0074] Among them, the loss function for obtaining the large model is:

[0075] L ZO =L qa +L kr +L ki ;

[0076]

[0077] Among them, N qa represents the number of samples for the knowledge Q&A task; V represents the size of the answer vocabulary; y i,j represents the value of the one-hot encoding of the true answer at the j-th vocabulary position for the i-th knowledge Q&A sample; represents the value of the probability distribution of the model's predicted answer at the j-th vocabulary for the i-th knowledge Q&A sample; q i represents the question semantic vector generated by the model for the i-th knowledge Q&A question; represents the expected question semantic vector for the i-th question obtained from the knowledge graph analysis; γ represents the weight coefficient; N kr represents the number of samples for the knowledge reasoning task; represents the form of the relationship inferred by the model for the k-th knowledge reasoning sample; r k represents the true form of the relationship for the k-th knowledge reasoning sample; N ki represents the number of samples for the knowledge completion task; represents the knowledge representation after the model completes the m-th knowledge completion task sample; k m represents the true and complete knowledge representation for the m-th knowledge completion task sample; R m represents the degree to which the knowledge completed by the model for the m-th knowledge completion task sample violates the power engineering rules or common sense; v represents the weight coefficient used to balance the two losses of completion accuracy loss and completion rationality loss in the knowledge completion task loss function.

[0078] Furthermore, the loss function of the large model comprehensively considers three different types of task losses: knowledge question answering, knowledge reasoning, and knowledge completion, and can comprehensively measure the performance of the model in the application of the power engineering knowledge base. By optimizing multiple tasks simultaneously, it ensures that the model can not only accurately answer various power engineering knowledge questions, but also perform effective knowledge reasoning to solve complex problems, and reasonably complete the missing knowledge, so that the model has a wider range of knowledge processing capabilities and is suitable for a variety of practical demand scenarios in the field of power engineering, such as power system planning and design, equipment operation and maintenance, fault diagnosis and processing, etc.

[0079] The present invention relates to the field of electric power engineering information technology, and proposes a method and system for constructing an electric power engineering knowledge base based on a large model. In the present invention, the proposed method for constructing an electric power engineering knowledge base based on a large model first obtains an electric power engineering professional terminology library, and trains a first model with text data, and trains the first model with the help of the professional terminology library, so that it can accurately understand and process these complex text contents, improve the quality and integrity of knowledge extraction, and try to avoid the problem that the electric power engineering text is rich in a large number of professional terms and the sentence structure is complex, which makes the knowledge extraction process prone to understanding difficulties and data omissions; in addition, by constructing a multimodal knowledge graph with image and drawing data, the knowledge of different modes is integrated together, and various types of knowledge in the field of electric power engineering can be more completely presented, from text descriptions such as equipment parameters and operating principles to intuitive displays such as equipment appearance and system layout, covering all aspects of electric power engineering related information, and solving the problem of single and incomplete knowledge representation; finally, through multimodal fusion and staged model training, the constructed knowledge base can respond to natural language queries more accurately, and reason and answer based on comprehensive and accurate knowledge, effectively improving the query efficiency of the knowledge base in practical applications and the effectiveness of answers.

[0080] Step 103, in one embodiment, the text data is split based on the grammatical structure. The grammatical analysis technology in natural language processing, such as a rule-based grammatical parser or a grammatical analysis model trained by a statistical learning method, can be used to process the collected power engineering text data. For example, for the sentence "The rated capacity of transformer T1 of substation A is 50MVA, and it is connected to substation B through transmission line L1", according to the grammatical relationship between the words, it is split into "substation A", "of", "transformer T1", "rated capacity", "for"

[0081] “50MVA”, “,” through “transmission line L1”, “and”, “substation B”, and other sub-data.

[0082] At the same time, it should be noted that multiple grammatical structures should be considered, including but not limited to subject-predicate-object structure, attributive-predicate structure, adverbial-predicate structure, etc., to ensure that the text can be accurately decomposed according to reasonable grammatical rules, and try to avoid destroying the integrity of semantic information during the decomposition process.

[0083] In the process of word vector conversion, word vector generation models such as Word2Vec, GloVe, etc. can be used to convert each sub-data (word or word sequence) into a corresponding word vector representation. It should be noted that when selecting a word vector model, the characteristics of the power engineering text data should be considered. For example, for specific professional terms or abbreviations in the field of power engineering, it is necessary to fine-tune the pre-trained word vector model to ensure that professional vocabulary can be accurately represented by vectors, thereby better serving the subsequent entity recognition tasks.

[0084] When encoding word vectors, the encoding layer of the language model receives word vectors as input, such as the encoder in the common Transformer architecture. Since the encoder is composed of multiple self-attention layers and feedforward neural network layers, the encoder can capture long-distance dependencies and semantic features between word vectors. For example, for a set of input word vectors (corresponding to the word vectors of each sub-data after splitting), the self-attention mechanism calculates the degree of association between each word vector and other word vectors, and then performs weighted sum update on each word vector based on these association degrees, so that each word vector incorporates the semantic information of the context.

[0085] After being processed by multiple self-attention layers and feedforward neural network layers, the semantic representation of the word vector is further enriched and abstracted. The output encoded data not only contains the semantics of a single word itself, but also incorporates its contextual semantic information in the entire text sentence, providing a more comprehensive and in-depth semantic foundation for subsequent feature extraction.

[0086] Furthermore, based on the coded data, the language model further extracts features using its specific structure and parameters. For example, through the convolutional neural network (CNN) structure or pooling operation, the feature combinations that play a key role in distinguishing different power engineering entities are screened out from the coded data, and some noise information and redundant features are filtered out, so that the final coded feature data can be more focused on the feature dimensions related to entity recognition. In addition, the feature extraction process can also be targeted according to the category characteristics of the power engineering entity. For example, for power equipment entities, more attention may be paid to features related to equipment name, model, and function; for power engineering location entities, more emphasis is placed on the extraction of features related to geographical location, etc., in order to improve the accuracy and pertinence of entity recognition.

[0087] Furthermore, a classifier (using classification algorithms based on traditional machine learning, such as support vector machines, decision trees, etc.) can be used to perform classification operations on the extracted encoded feature data. The classifier is trained based on the already annotated power engineering text data to learn the mapping relationship between the encoded feature data and different entity categories. For example, the encoded feature data is F = [f1, f2, …, f δ (δ is the feature dimension). For the encoded feature vector corresponding to each sub-data, when determining which power engineering entity category it belongs to, assuming the total number of preset categories is B, the probability of belonging to each category is calculated through a logistic regression model as follows: where v represents the category variable, b represents the b-th category, θ b represents the learnable parameter vector corresponding to the b-th category, and T represents the transpose operation.

[0088] In obtaining the entity association data, first, based on multiple entities in the previously named recognized power engineering, all entity combinations are traversed, and entity pairs with an association relationship are screened out. During this process, the association relationship is judged and screened according to the syntactic structure, semantic logic in the text, and professional knowledge in the power engineering field. For example, in the text "The rated capacity of transformer T1 in substation A is 50 MVA and it is connected to substation B through transmission line L1", entity pairs ("substation A", "transformer T1"), ("transformer T1", "transmission line L1"), ("transmission line L1", "substation B"), etc. can be determined, and they have relationships such as inclusion and connection respectively.

[0089] After that, for each entity pair, information extraction is performed in the text data using its relationship features as clues. The relationship features can include conjunctions, demonstrative pronouns, specific power engineering terms, and the relative position information of the entity pair in the text, etc. Feature extraction techniques can also be used, such as window-based feature extraction methods. Taking the position where the entity pair first appears in the text as the center, a window range is set (such as 5 words before and after), and information such as words, part-of-speech, and word vectors within the window is extracted as relationship feature data.

[0090] When classifying and annotating the relationships between entity pairs, a classifier is used to classify and annotate the extracted relationship feature data. The training of the classifier is based on a dataset of power engineering texts with pre-annotated entity relationships, which contains various entity pairs and their corresponding relationship type annotations. For entity attribute data, keywords or phrases related to entity attributes are identified from the text data. The attribute values are determined based on the obtained entity attribute data. For text segments with explicitly given attribute values, numerical or text information is directly extracted as the attribute value. The determined entity attribute values are associated and annotated with the previously obtained entity relationship data to form entity association data. Centered around the entity, its attribute values and related entity relationship information are integrated together to construct a structured data unit.

[0091] In step 104, in one embodiment, computer vision technology is used to extract the features of image data. First, the image is preprocessed, including grayscale conversion (if the original image is in color) and normalization (adjusting the image size, brightness, contrast, etc. to make the image data within a suitable numerical range and with a unified specification) to reduce the data volume and highlight key information.

[0092] In the recognition and conversion of drawing data, first, the digitization of the drawing is carried out to convert paper drawings or electronic drawings in different formats into a unified editable format, such as SVG (Scalable Vector Graphics) format, for subsequent processing. For example, optical character recognition (OCR) technology and graphic recognition algorithms are used to identify and extract elements such as text annotations, device symbols, and line connections in the drawing. Semantic understanding and conversion processing are performed on the recognized drawing elements to convert them into structured second feature information. For example, for the device symbols in an electrical main wiring diagram, according to the power engineering standards and knowledge system, they are converted into the corresponding device entity names and type information; for the line connection relationship, a connection relationship matrix or graph structure is constructed to represent it.

[0093] Build the basic framework of the knowledge graph based on entity association data. Entity association data includes entities in power engineering (such as substations, transformers, transmission lines, etc.), relationships between entities (such as inclusion, connection, subordination, etc.), and attributes of entities (such as the rated capacity and production date of equipment). Use entities as nodes in the knowledge graph, relationships between entities as edges, and attributes of entities as attribute information of nodes to build a preliminary knowledge graph framework. For example, create an instance of a graph database (such as Neo4j), store each entity node as a node object, and the node object contains the entity name, type, and related attribute information; store the relationships between entities as edge objects, and the edge objects contain information such as relationship type, start node, and end node, thereby building a basic framework of the knowledge graph based on entity association data in power engineering, and this framework presents the basic structure and logical relationships of power engineering knowledge.

[0094] Associate the first feature information of the image data and the second feature information of the drawing data with the basic framework of the knowledge graph. For the image feature information, associate it with the corresponding entity node. For example, if an image is an external image of a certain substation, the extracted image feature information is associated with the entity node of this substation in the knowledge graph. The image feature vector can be stored as a special attribute of this node or a separate association table can be established to store the correspondence between the image features and the entity nodes, so that image information can be quickly obtained and integrated with other knowledge during query and reasoning. For the feature information of the drawing data, associate and integrate the converted structured data (such as equipment location coordinates, connection relationship diagrams, etc.) with the entity nodes and edges in the knowledge graph. Through the association operation, the multi-modal feature information of the image and drawing data is incorporated into the knowledge graph framework to form a complete multi-modal knowledge graph.

[0095] Step 105, in one embodiment, first preprocess and feature-encode the multi-modal knowledge graph. For the entities in the knowledge graph, convert their text attributes (such as name, type, attribute description, etc.) into vector representations through a word vector model (such as Word2Vec or BERT word embedding); perform normalization and dimension adjustment processing on the entity features associated with the image and drawing data (such as visual feature vectors extracted from the image, spatial layout and connection feature vectors in the drawing) so that they can be fused in the same feature space as the text vectors. And design an embedding fusion layer to embed the processed knowledge graph features into the input layer of the pre-trained large model. One way is to use the concatenation operation, and concatenate the text vector, image feature vector, and drawing feature vector of the entity in sequence to form a new comprehensive vector as part of the input layer. For example, the text vector is V t (with dimension d t ), the image feature vector is V i(with dimension d i ), the drawing feature vector is V d (with dimension d d ), and the combined total vector is V con = [V t ; V i ; V d (with dimension d t + d i + d d ).

[0096] For the power engineering Q&A task, various types of Q&A objects are sorted out from the power engineering text data, including equipment parameter queries, operation process inquiries, fault diagnosis Q&A, etc. For the knowledge reasoning task, the logical reasoning relationships contained in the text data are explored to construct an inference task dataset. For example, according to the topological structure of the power system and the operating principles of equipment, inference cases are created, such as "Given that transformer T1 fails in substation A, infer its impact on the downstream transmission lines and user power supply", "If a new transmission line is added to substation B, analyze the change in the power flow distribution of the entire power system", etc. For the knowledge completion task, some knowledge missing situations are artificially created in the text data to construct a knowledge completion task dataset. For example, some key parameter or relationship information is deleted from the text describing power equipment, and the model is allowed to predict and complete the missing knowledge content based on the context and other relevant information in the knowledge graph.

[0097] For the knowledge Q&A task, the question part of the Q&A pair is input into the integrated large model framework. The model is processed through multiple layers of neural networks (such as the encoder layer based on the Transformer architecture), and the answer is predicted at the output layer. For the knowledge reasoning task, the preconditions or known information of the reasoning task are input into the model, and the model outputs the reasoning result or conclusion. In the knowledge completion task, the text with missing knowledge is input into the model, and the model predicts and completes the missing part.

[0098] And by adjusting the weight coefficients of the loss function, according to the actual needs and the performance of the model on different tasks, the training focus of the model in aspects such as knowledge Q&A, knowledge reasoning, and knowledge completion can be flexibly controlled, so that the model can achieve better performance on multiple tasks, thereby obtaining a large model applicable to the power engineering field.

[0099] Figure 2 is a structural schematic diagram of a power engineering knowledge base construction system provided by the present invention, as Figure 2As shown in the figure, the device includes: an acquisition unit 10: configured to acquire original knowledge data of power engineering and a power engineering professional term library; the original knowledge data includes text data, and image and drawing data; a first training unit 20: configured to perform model training on a pre-trained language model based on the text data and the power engineering professional term library to obtain a language model; an entity extraction and annotation unit 30: configured to perform named entity recognition on the text data based on the language model to obtain multiple entities in power engineering, and extract and annotate the relationships and attributes between entities to obtain entity association data; a knowledge graph construction unit 40: configured to construct a multi-modal knowledge graph based on the image and drawing data and the entity association data; a storage unit 50: configured to perform model training on a pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and fuse and store the large model with the multi-modal knowledge graph to obtain a power engineering knowledge base.

[0100] Figure 3 is a schematic structural diagram of the electronic device provided by the present invention. As Figure 3 shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 1030 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute a method for constructing a power engineering knowledge base based on a large model. The method includes: acquiring original knowledge data of power engineering and a power engineering professional term library; the original knowledge data includes text data, and image and drawing data; performing model training on a pre-trained language model based on the text data and the power engineering professional term library to obtain a language model; performing named entity recognition on the text data based on the language model to obtain multiple entities in power engineering, and extracting and annotating the relationships and attributes between entities to obtain entity association data; constructing a multi-modal knowledge graph based on the image and drawing data and the entity association data; performing model training on a pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and fusing and storing the large model with the multi-modal knowledge graph to obtain a power engineering knowledge base.

[0101] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods according to the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0102] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a method for constructing a power engineering knowledge base based on a large model provided by the above-mentioned various methods. The method includes: obtaining original power engineering knowledge data and a power engineering professional term library; the original knowledge data includes text data and image and drawing data; training a pre-trained language model based on the text data and the power engineering professional term library to obtain a language model; performing named entity recognition on the text data based on the language model to obtain multiple entities in the power engineering, and extracting and annotating the relationships and attributes between the entities to obtain entity association data; constructing a multi-modal knowledge graph based on the image and drawing data and the entity association data; training a pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and fusing and storing the large model with the multi-modal knowledge graph to obtain a power engineering knowledge base.

[0103] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute a method for constructing a power engineering knowledge base based on a large model provided above. The method includes: obtaining original power engineering knowledge data and a power engineering terminology library; the original knowledge data includes text data, as well as image and drawing data; training a pre-trained language model based on the text data and the power engineering terminology library to obtain a language model; performing named entity recognition on the text data based on the language model to obtain multiple entities in power engineering, and extracting and annotating the relationships and attributes between the entities to obtain entity association data; constructing a multi-modal knowledge graph based on the image and drawing data and the entity association data; training a pre-trained large model based on the text data and the multi-modal knowledge graph to obtain a large model, and fusing and storing the large model with the multi-modal knowledge graph to obtain a power engineering knowledge base.

[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0106] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.

Claims

1. A method for constructing a power engineering knowledge base based on a large model, characterized in that: The following steps are involved: Acquire original knowledge data of electric power engineering and a professional terminology library of electric power engineering; the original knowledge data includes text data and image and drawing data; Performing model training on the pre-trained language model based on the text data and the electric power engineering professional terminology library to obtain a language model; Based on the language model, the text data is named and recognized to obtain multiple entities in the power engineering project, and the relationships and attributes between the entities are extracted and annotated to obtain entity-related data; Construct a multimodal knowledge graph based on image and drawing data and entity association data; The pre-trained large model is trained based on the text data and the multimodal knowledge graph to obtain a large model, and the large model is fused and stored with the multimodal knowledge graph to obtain a power engineering knowledge base.

2. The method for constructing a power engineering knowledge base based on a large model according to claim 1, characterized in that: The text data includes power equipment data, power system operation data, power engineering design data, power engineering construction data and power engineering maintenance data.

3. The method for constructing a power engineering knowledge base based on a large model according to claim 1, characterized in that: The naming and recognition of the text data based on the language model is performed to obtain multiple entities in the power engineering, including: Splitting the text data based on its grammatical structure to obtain a plurality of sub-data; each sub-data is a word or a word sequence; Transform each sub-data to obtain word vector; Encoding the word vector based on the encoding layer of the language model to obtain encoded data; Extracting features from the coded data based on the language model to obtain coded feature data; The plurality of sub-data are classified based on the coded feature data to obtain a plurality of entities in the electric power engineering project.

4. The method for constructing a power engineering knowledge base based on a large model according to claim 3 is characterized in that: The extraction and annotation of the relationships and attributes between entities to obtain entity association data includes: Based on the association relationship between the multiple entities in the power engineering, a plurality of entity pairs are obtained; Extracting text data based on the relationship features between the entity pairs to obtain relationship feature data; Classify and label the relationships between entity pairs based on the relationship feature data to obtain entity relationship data; Obtaining keywords or phrases related to entity attributes in the text data to obtain entity attribute data; Based on the entity attribute data, the entity attribute value is determined, and the entity attribute value is associated with the entity relationship data to obtain the entity association data.

5. The method for constructing a power engineering knowledge base based on a large model according to claim 1, characterized in that: The multimodal knowledge graph is constructed based on image and drawing data and entity association data, including: Extracting features from the image data to obtain first feature information; Performing identification and conversion processing on the drawing data to obtain second feature information; Build the basic framework of knowledge graph based on entity association data; The first feature information, the second feature information and the basic framework of the knowledge graph are associated to obtain a multimodal knowledge graph.

6. The method for constructing a power engineering knowledge base based on a large model according to claim 1, characterized in that: The pre-trained large model is trained based on the text data and the multimodal knowledge graph to obtain the large model, including: Embedding and fusing the multimodal knowledge graph into the input layer of the pre-trained large model to obtain a fused large model frame; Acquire power engineering knowledge question-answering tasks, knowledge reasoning tasks, and knowledge completion tasks in the text data; The fusion large model frame is trained with the electric power engineering knowledge question-answering task, knowledge reasoning task and knowledge completion task as training targets to obtain a large model.

7. A method for constructing a power engineering knowledge base based on a large model according to claim 6, characterized in that: The loss function of the large model is: L ZO =L qa +L kr +L ki ; Among them, N qa represents the number of samples in the knowledge question answering task; V represents the size of the answer vocabulary; y i,j represents the value of the one-hot encoding of the true answer at the jth vocabulary position for the i-th knowledge question and answer sample; represents the value of the probability distribution of the model predicting the answer for the i-th knowledge question and answer sample on the j-th word; q i represents the question semantic vector generated by the model for the i-th knowledge question answering question; represents the expected question semantic vector for the i-th question obtained by knowledge graph analysis; γ represents the weight coefficient; N kr Indicates the number of samples for the knowledge reasoning task; represents the relational representation obtained by model reasoning for the kth knowledge reasoning sample; r k Table 1 is the true relational representation of the kth knowledge reasoning sample; N ki Indicates the number of samples for the knowledge completion task; represents the knowledge representation after model completion for the mth knowledge completion task sample; k m represents the true and complete knowledge representation for the mth knowledge completion task sample; R m It represents the degree to which the knowledge completed by the model violates the rules or common sense of power engineering for the mth knowledge completion task sample; v represents the weight coefficient used to balance the loss of completion accuracy and the loss of completion rationality in the loss function of the knowledge completion task.

8. A large-model-based power engineering knowledge base construction system, characterized in that: Applied to the IoT device security protection method based on big data as claimed in any one of claims 1 to 7, the power engineering knowledge base construction system based on a large model includes: Acquisition unit: used to acquire original knowledge data of electric power engineering and a professional terminology library of electric power engineering; the original knowledge data includes text data and image and drawing data; A first training unit: used for performing model training on the pre-trained language model based on the text data and the electric power engineering professional terminology library to obtain a language model; Entity extraction and annotation unit: used for performing naming and recognition on the text data based on the language model to obtain multiple entities in the power engineering, and extracting and annotating the relationships and attributes between the entities to obtain entity association data; Graph construction unit: used to construct a multimodal knowledge graph based on image and drawing data and entity association data; Storage unit: used to perform model training on the pre-trained large model based on the text data and the multimodal knowledge graph to obtain the large model, fuse the large model with the multimodal knowledge graph and store it to obtain the power engineering knowledge base.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a method for constructing a large model-based electric power engineering knowledge base are implemented as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for constructing a large model-based electric power engineering knowledge base as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Knowledge base construction method, system and equipment based on large model

    CN120542552A

  • Large model processing method and device for multi-modal knowledge graph, equipment and medium

    CN120579617A

  • YOLO-OCR highway engineering quantity intelligent hooking system

    CN120580714A

  • Knowledge base processing method and system for power field

    CN120745783A

  • A knowledge base processing method and system for the power field

    CN120745783B