Power distribution network knowledge base generation method and device, equipment and storage medium
By acquiring and preprocessing distribution network data, using preset knowledge and relationship extraction models to determine entity information and relationships, and generating distribution network knowledge bases, the problem of low efficiency in generating distribution network knowledge bases in the existing technology is solved, and a more efficient and accurate knowledge base construction is achieved.
Patent Information
- Application Number
- CN202510124293.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the efficiency of generating a distribution network knowledge base is low, mainly relying on expert experience to build, and the distribution network data is large and the construction process is cumbersome.
By obtaining distribution network data, the data is processed based on the preset knowledge extraction model and the preset relationship extraction model, the entity information and entity relationship are obtained, and the knowledge base is determined. The method includes obtaining structured and unstructured data, performing preprocessing, determining entity information using embedded sub-models, coding sub-models and decoding sub-models, and determining entity relationships using feature fusion sub-models, cross-attention sub-models, etc.
It effectively improves the efficiency of generating the distribution network knowledge base. Through automated data processing and extraction of entity information and relationships, it reduces the complexity of manual intervention and construction, and improves the accuracy and efficiency of data processing.
Smart Images

Figure CN120069050A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of distribution networks, and particularly to a method, device, equipment, and storage medium for generating a distribution network knowledge base. Background Art
[0002] With the development of smart grids, the intelligent requirements for distribution networks are getting higher and higher. Establishing a knowledge base for distribution network production command and conducting reasoning, analysis, and decision-making have important theoretical significance and practical value for the evaluation and optimization of power grids, improving grid performance, and enhancing enterprise economic benefits.
[0003] In the prior art, when establishing a distribution network knowledge base based on distribution network data, it mainly relies on expert experience for construction. Due to the large amount of distribution network data and the cumbersome construction process, the efficiency of generating a distribution network knowledge base based on the prior art is relatively low. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, equipment, and storage medium for generating a distribution network knowledge base that can effectively improve the generation efficiency in view of the above technical problems.
[0005] In a first aspect, the present application provides a method for generating a distribution network knowledge base, including:
[0006] Obtain distribution network data, where the distribution network data includes distribution network operation data and distribution network production command materials, and the distribution network production command materials include various information resources and data for supporting distribution network operation, maintenance, fault handling, and dispatching command;
[0007] Process the distribution network data based on a preset knowledge extraction model and a preset relationship extraction model to obtain entity information and entity relationships;
[0008] Determine a knowledge base according to the entity information and entity relationships.
[0009] In one embodiment, the distribution network data includes structured data and unstructured data. Before processing based on the preset knowledge extraction model and the preset relationship extraction model, the method further includes:
[0010] Obtain the structured data in the distribution network data and store the structured data in a preset database; obtain the unstructured data in the distribution network data and preprocess the unstructured data to obtain target text information.
[0011] In one embodiment, processing the distribution network data based on the preset knowledge extraction model and the preset relationship extraction model to obtain entity information and entity relationships includes:
[0012] Input the target text information into a preset knowledge extraction model to determine entity information. The preset knowledge extraction model includes an embedding submodel, an encoding submodel, and a decoding submodel. Input the target text information into a preset relationship extraction model to determine entity relationships. The preset relationship extraction model includes a feature fusion submodel, a cross-attention submodel, a latent relationship prediction submodel, and a subject-object correspondence submodel.
[0013] In one embodiment, the encoding submodel includes a first encoding submodel based on the multi-head self-attention mechanism and a second encoding submodel based on the dilated convolutional neural network. Inputting the target text information into the preset knowledge extraction model to determine entity information includes:
[0014] Input the target text information into the BERT-based embedding submodel to convert each character of the target text information into an embedding vector. Input the embedding vector into the first encoding submodel to obtain a first encoding result, and input the embedding vector into the second encoding submodel to obtain a second encoding result. Fuse the first encoding result and the second encoding result to obtain a target encoding result. Input the target encoding result into the conditional random field-based decoding submodel to determine entity information according to the output of the decoding submodel.
[0015] In one embodiment, the feature fusion submodel includes a BiLSTM submodel and a CDIL-CNN convolutional neural network submodel. Inputting the target text information into the preset relationship extraction model to determine entity relationships includes:
[0016] Respectively perform feature extraction on the target text information according to the BiLSTM submodel and the CDIL-CNN convolutional neural network submodel to obtain different feature extraction results, and determine a first fusion feature according to each feature extraction result. Determine a second fusion feature according to the first fusion feature and perturbation data, and input the first fusion feature and the second fusion feature into the cross-attention submodel to obtain a first feature word vector and a second feature word vector. The perturbation data is determined according to the first fusion feature and a preset adversarial learning submodel. Determine various latent relationships included in the target text information according to the first feature word vector and the latent relationship prediction submodel. Input each latent relationship and the second feature word vector into the subject-object correspondence submodel to determine entity relationships.
[0017] In one embodiment, determining a knowledge base according to entity information and entity relationships includes:
[0018] Determine the abstract concept layer data and the concept instance layer data according to the entity information. The entity information in the abstract concept layer data includes distribution network generation command, equipment, line, operation data, equipment name, and equipment number. The entity information in the concept instance layer data is used to represent the specific content of the entity information in the abstract concept layer data; determine the capability layer data according to the entity relationship; determine the knowledge base according to the abstract concept layer data, the concept instance layer data, and the capability layer data.
[0019] In a second aspect, the present application also provides a device for determining a distribution knowledge base, including:
[0020] An acquisition module, configured to acquire distribution network data, where the distribution network data includes distribution network operation data and distribution network production command materials, and the distribution network production command materials include various information resources and data for supporting distribution network operation, maintenance, fault handling, and dispatching command.
[0021] An extraction module, configured to process the distribution network data based on a preset knowledge extraction model and a preset relationship extraction model to obtain entity information and entity relationships.
[0022] A generation module, configured to determine the knowledge base according to the entity information and the entity relationship.
[0023] In one embodiment, the distribution network data includes structured data and unstructured data. Before the preset knowledge extraction model and the preset relationship extraction model, a preprocessing module is further included, configured to acquire the structured data in the distribution network data and store the structured data in a preset database; acquire the unstructured data in the distribution network data and preprocess the unstructured data to obtain target text information.
[0024] In one embodiment, the extraction module is specifically configured to input the target text information into the preset knowledge extraction model to determine the entity information, and the preset knowledge extraction model includes an embedding submodel, an encoding submodel, and a decoding submodel; input the target text information into the preset relationship extraction model to determine the entity relationship, and the preset relationship extraction model includes a feature fusion submodel, a cross-attention submodel, a potential relationship prediction submodel, and a subject-object correspondence submodel.
[0025] In one embodiment, the encoding sub-model includes a first encoding sub-model based on the multi-head self-attention mechanism and a second encoding sub-model based on the dilated convolutional neural network. The extraction module is specifically configured to input the target text information into the embedding sub-model based on BERT to convert each character of the target text information into an embedding vector; input the embedding vector into the first encoding sub-model to obtain a first encoding result, input the embedding vector into the second encoding sub-model to obtain a second encoding result; fuse the first encoding result and the second encoding result to obtain a target encoding result; input the target encoding result into the decoding sub-model based on the conditional random field, and determine the entity information according to the output of the decoding sub-model.
[0026] In one embodiment, the feature fusion sub-model includes a BiLSTM sub-model and a CDIL-CNN convolutional neural network sub-model. The extraction module is specifically configured to perform feature extraction on the target text information according to the BiLSTM sub-model and the CDIL-CNN convolutional neural network sub-model respectively to obtain different feature extraction results, and determine a first fusion feature according to each feature extraction result; determine a second fusion feature according to the first fusion feature and the perturbation data, and input the first fusion feature and the second fusion feature into the cross-attention sub-model to obtain a first feature word vector and a second feature word vector, and the perturbation data is determined according to the first fusion feature and the preset adversarial learning sub-model; determine various potential relationships included in the target text information according to the first feature word vector and the potential relationship prediction sub-model; input each potential relationship and the second feature word vector into the subject-object correspondence sub-model to determine the entity relationship.
[0027] In one embodiment, the generation module is specifically configured to determine the abstract concept layer data and the concept instance layer data according to the entity information. The entity information in the abstract concept layer data includes distribution network generation command, equipment, line, operation data, equipment name, equipment number, and the entity information in the concept instance layer data is used to represent the specific content of the entity information in the abstract concept layer data; determine the capability layer data according to the entity relationship; determine the knowledge base according to the abstract concept layer data, the concept instance layer data and the capability layer data.
[0028] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in any one of the first aspects above is implemented.
[0029] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the first aspects above is implemented.
[0030] In a fifth aspect, the present application also provides a computer program product, including a computer program which, when executed by a processor, implements the method described in any one of the first aspects above.
[0031] For the above method, device, equipment, and storage medium for generating a distribution network knowledge base, first, distribution network data is obtained. The distribution network data includes distribution network operation data and distribution network production command materials, and the distribution network production command materials include various information resources and data for supporting distribution network operation, maintenance, fault handling, and dispatching command. Then, based on a preset knowledge extraction model and a preset relationship extraction model, the distribution network data is processed to obtain entity information and entity relationships. Then, a knowledge base is determined according to the entity information and entity relationships. By inputting the distribution network data into the corresponding models, entity information extraction and entity relationship extraction are performed, and then a knowledge base is determined based on the extracted entity information and entity relationships, which can effectively improve the processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic flowchart of the method for generating a distribution network knowledge base in one embodiment;
[0034] Figure 2 It is a schematic flowchart of the method for generating a distribution network knowledge base in another embodiment;
[0035] Figure 3 It is a schematic flowchart of the steps for determining entity information and entity relationships in one embodiment;
[0036] Figure 4 It is a schematic flowchart of the steps for determining entity information in one embodiment;
[0037] Figure 5 It is a schematic flowchart of the steps for determining entity relationships in one embodiment;
[0038] Figure 6 It is a schematic flowchart of the steps for determining a knowledge base according to the entity information and the entity relationships in one embodiment;
[0039] Figure 7 It is a schematic flowchart of the method for generating a distribution network knowledge base in another embodiment;
[0040] Figure 8The structural block diagram of the distribution network knowledge base generation device in an embodiment;
[0041] Figure 9 The internal structure diagram of a computer device in an embodiment. Specific embodiments
[0042] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] With the development of smart grids, the requirements for the intelligence of distribution networks are getting higher and higher. Establishing a knowledge base for distribution network production command and conducting reasoning, analysis and decision-making have important theoretical significance and practical value for the evaluation and optimization of power grids, improving grid performance and enhancing enterprise economic benefits.
[0044] In the prior art, when establishing a distribution network knowledge base based on distribution network data, it mainly relies on expert experience for construction. Due to the large amount of distribution network data and the cumbersome construction process, the efficiency of generating the distribution network knowledge base based on the prior art is relatively low.
[0045] In view of this, the present application provides a method for generating a distribution network knowledge base that can improve the generation efficiency. The method for generating a distribution network knowledge base provided by the embodiments of the present application may be executed by a distribution network knowledge base generation device. The distribution network knowledge base generation device may be implemented in a software, hardware or a combination of software and hardware manner. It may be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form. In the following method embodiments, the computer device is taken as an example of the execution subject for description. The computer device may be a server or a computer. The specific type of the computer device is not limited in the embodiments of the present application.
[0046] In an exemplary embodiment, as Figure 1 shown, a method for generating a distribution network knowledge base is provided, including the following steps 101 to 103. Among them:
[0047] S101, obtain distribution network data.
[0048] Among them, the distribution network data includes distribution network operation data and distribution network production command materials. The distribution network production command materials include various information resources and data for supporting the operation, maintenance, fault handling and dispatching command of the distribution network.
[0049] Optionally, the operation data of the distribution network in the distribution network data can be obtained in real time or can be the historical operation data of the distribution network. The distribution network production command materials in the distribution network data can be obtained from a preset file or obtained from the distribution network system through crawler analysis. The embodiments of the present application do not limit the manner of obtaining the distribution network data.
[0050] Optionally, the distribution network data may further include data in the field of distribution network production command, including data such as equipment names, equipment numbers, equipment power curves, standby power supply status, load curves, etc., as well as important terms in the field of distribution network production command listed in advance, relevant knowledge and configurations involved in distribution network production command, which are classified into equipment, lines, operation data, etc.
[0051] S102. Process the distribution network data based on a preset knowledge extraction model and a preset relationship extraction model to obtain entity information and entity relationships.
[0052] Optionally, the preset knowledge extraction model can be a named entity recognition model for identifying entities and entity types in the distribution network data, such as equipment names, equipment numbers, etc. The preset relationship extraction model can be used to identify the semantic relationships between entities, such as causality, subordination, or time.
[0053] In a possible implementation manner, the preset knowledge extraction model and the preset relationship extraction model can be two independent tasks, that is, the distribution network data is respectively input into the preset knowledge extraction model and the preset relationship extraction model to obtain entity information and entity data respectively.
[0054] Optionally, the preset knowledge extraction model can include a hidden Markov model, a conditional random field, a recurrent neural network model and its variants, a pre-trained model based on the Transformer architecture, etc.; the preset relationship extraction model can include a support vector machine, maximum entropy, a convolutional neural network, a recurrent neural network, or a pre-trained language model, etc. The embodiments of the present application do not limit this.
[0055] In another possible implementation manner, the preset knowledge extraction model and the preset relationship extraction model can be a joint extraction model. The knowledge extraction model and the relationship extraction model are respectively subtasks in the joint extraction model, and the two subtasks interact with each other. When the distribution network data is input into the joint extraction model, entity information and entity relationships can be fully identified through the interaction of the two subtasks.
[0056] Optionally, the joint extraction model can include joint extraction based on feature engineering, such as integer linear programming, card pyramid parsing, or structured prediction, etc., and can also include a joint extraction model based on deep learning and a joint extraction model based on reinforcement learning, etc.
[0057] S103. Determine the knowledge base according to the entity information and entity relationships.
[0058] Optionally, the entity information and relationship information can be fused and aligned through a preset knowledge fusion method, and the fused entity information and entity relationships are stored to obtain the knowledge base. The storage methods can include graph databases and ontology representations.
[0059] Optionally, a graph database can store entity information and entity relationships using a graph structure, with entity information as nodes and entity relationships as edges.
[0060] Optionally, an ontology representation can use the Ontology Language (Web Ontology Language, OWL) to describe entity information and entity relationships. Specifically, OWL includes classes, instances, properties, and logical reasoning. Classes can be a collection of entities, such as device names, lines, etc., which can be defined as classes. Instances can be specific objects of classes, such as specific device numbers starting with the number "K". Properties can be entity relationships. At the same time, OWL supports logical reasoning and can derive new knowledge from existing knowledge.
[0061] Exemplarily, the Protégé software can be used to define classes, properties, and relationships.
[0062] Optionally, the knowledge base can be continuously updated and expanded through knowledge reasoning, knowledge updating, etc. to maintain the timeliness and accuracy of the knowledge base.
[0063] Optionally, the determined knowledge base can be used in scenarios such as intelligent question - answering systems, knowledge graph visualization, or semantic search.
[0064] For the above - mentioned method for generating a distribution network knowledge base, first, obtain distribution network data. The distribution network data includes distribution network operation data and distribution network production command materials. The distribution network production command materials include various information resources and data for supporting distribution network operation, maintenance, fault handling, and dispatching command. Then, based on a preset knowledge extraction model and a preset relationship extraction model, process the distribution network data to obtain entity information and entity relationships. Then, determine the knowledge base according to the entity information and entity relationships. By inputting the distribution network data into the corresponding models, entity information extraction and entity relationship extraction are performed, and then the knowledge base is determined based on the extracted entity information and entity relationships, which can effectively improve the processing efficiency.
[0065] In an exemplary embodiment, as Figure 2 shown, optionally, the distribution network data includes structured data and unstructured data. Before the preset knowledge extraction model and preset relationship extraction model, the method for generating a distribution network knowledge base further includes the following steps 201 to step 202. Among them:
[0066] S201. Obtain the structured data in the distribution network data and store the structured data in a preset database.
[0067] Optionally, the structured data may refer to data stored in a tabular form, usually having clear fields and data types, which can be determined by querying a relational database, obtained from a CSV or Excel file, or structured data obtained through an API interface.
[0068] Optionally, after obtaining the structured data, it can be stored in a preset database. It should be noted that entity information can be determined based on the table structure and fields of the database. For example, the table name can represent an entity information, that is, a class, and the fields can contain the attributes of the entity. At the same time, based on the primary key and foreign key relationships in the database, the entity relationships between entity information can be determined, or an ER diagram tool can be used to visualize the database structure to intuitively identify the entity relationships.
[0069] S202. Obtain the unstructured data in the distribution network data, preprocess the unstructured data, and obtain the target text information.
[0070] Optionally, the unstructured data is text data without a fixed format and structure. When preprocessing the unstructured data, it can include sentence splitting, paragraph splitting, data cleaning, and abnormal data correction processing, etc.
[0071] The above steps of obtaining the structured data in the distribution network data, storing the structured data in a preset database, obtaining the unstructured data in the distribution network data, preprocessing the unstructured data, and obtaining the target text information. Since the distribution network data is obtained through various methods and is multi-source heterogeneous data, the distribution network data can be divided into different types according to the structure and different processing methods can be adopted, which can improve the efficiency and accuracy of knowledge base production.
[0072] In an exemplary embodiment, as Figure 3 shown, optionally, based on a preset knowledge extraction model and a preset relationship extraction model, process the distribution network data to obtain entity information and entity relationships, including the following steps 301 to 302. Among them:
[0073] S301. Input the target text information into the preset knowledge extraction model to determine the entity information.
[0074] Among them, the preset knowledge extraction model includes an embedding sub-model, an encoding sub-model, and a decoding sub-model.
[0075] Optionally, an embedding sub-model can be used to convert target text information into embedding vectors, which can be implemented based on Word2Vec, BERT, or other Transformer models; an encoding sub-model can be used to further extract feature information from the embedding vectors, which can be implemented based on a recurrent neural network, for example, a Long Short-Term Memory (LSTM), a Gated Recurrent Unit (GRU), or a Bidirectional Long Short-Term Memory (BiLSTM), etc.; a decoding sub-model can be used to process the feature information to obtain entity labels, which can be implemented based on a conditional random field or based on softmax.
[0076] Optionally, the encoding sub-model includes a first encoding sub-model based on a multi-head self-attention mechanism and a second encoding sub-model based on a dilated convolutional neural network. As Figure 4 shown, inputting the target text information into a preset knowledge extraction model to determine entity information includes the following steps 401 to 404. Among them:
[0077] S401, input the target text information into the BERT-based embedding sub-model to convert each character of the target text information into an embedding vector.
[0078] Optionally, is a single character in the target text information, is the input target text information, E represents the embedding vector obtained from the BERT-based embedding sub-model, represents a single character embedding, represents the embedding sub-model, and the following expression can be obtained:
[0079]
[0080]
[0081] S402, input the embedding vector into the first encoding sub-model to obtain a first encoding result, and input the embedding vector into the second encoding sub-model to obtain a second encoding result.
[0082] Optionally, before inputting the embedding vector into the first encoding sub-model, the position embedding of the target text information can be determined and incorporated into the embedding vector to obtain a fused embedding vector. It can be understood that the embedding vector determined in step 401 is a character embedding.
[0083] Optionally, when fusing character embeddings and positional embeddings, they can be directly added together, or the character embeddings and relative positional embeddings can be concatenated and then fused through a fully connected layer (or linear layer). The embodiments of the present application do not limit the process of determining the fused embedding vectors.
[0084] Optionally, the positional embeddings can be determined based on relative position information or absolute position information. In the embodiments of the present application, taking relative position as an example, the positional embeddings can be determined based on Transformer-XL. For each element in the target text information, the relative positions between this element and other elements can be calculated. For the characters and characters
[0085]
[0086] The relative positions can be determined in the following way:
[0087] If i is even, ;
[0088] If i is odd, 。
[0089] where i is the index dimension, is the preset model dimension.
[0090] Optionally, the fused embedding vectors can be input into the first encoding sub-model. Under the multi-head self-attention mechanism, the fused embedding vectors containing positional embeddings and character embeddings can obtain the correlations between each character in the sequence and capture the hidden features of the entire sequence (the sequence length is l, and the input is d). Three learnable parameter matrices with dimensions are trained and projected into different spaces to obtain the query vector Q, the key vector K, and the value vector V. The expressions are as follows:
[0091]
[0092] where the single-head self-attention head can calculate with the help of the three vectors Q, K, and V through feature scaling and the softmax function. The expression is as follows:
[0093]
[0094] where is the dimension of the key vector, and the scaling factor Used to avoid the vanishing gradient caused by an overly large dot product.
[0095] Optionally, the results of multiple heads can be concatenated together and projected back to the original dimension through a linear transformation layer, which is expressed as follows:
[0096]
[0097] Among them, the number of attention heads is denoted as n; the attention combinations of different heads are denoted as ; the learnable weight matrix in the linear transformation layer is denoted as and its dimension is .
[0098] The output of the multi-head attention can be determined as the first encoding result, or the output of the multi-head attention can be fused with the embedding vector to determine the first encoding result.
[0099] In a possible implementation, the first encoding sub-model also includes an adaptive module that can fuse deep and shallow features. After obtaining the output of the multi-head attention, the first encoding result can be determined in the following way: the adaptive module will fuse the output features of the current layer with the output E of the previous layer. This method can better combine deep and shallow features to obtain the first fused feature compared to the direct addition residual method. Subsequently, after normalization processing, then, the first fused feature is input into the feed-forward neural network to obtain the second fused feature, and then the first fused feature and the second fused feature are processed by the adaptive module to obtain the first encoding result.
[0100] Optionally, the second encoding sub-model IDCNN network is formed by iteratively expanding a convolutional neural network. The dilated convolutional neural network adds a dilation width to the CNN network. In this way, with the same convolutional kernel size, different convolutional layers have different dilation rates, enabling a larger receptive field and thus obtaining more context information.
[0101] Optionally, the dilation coefficient of the first layer of IDCNN is 1. After the embedding vector E undergoes a convolutional operation, an output sequence is obtained, and the expression is as follows:
[0102]
[0103] represents the output of the dilated convolution of each layer, and the expression is as follows:
[0104]
[0105] Among them, the dilation coefficient of the j-th layer is denoted as , the dilated convolution block B is obtained by stacking multiple convolutional layers. To avoid overfitting caused by deep networks and ensure a wider range of solutions, the convolutional block B will be iteratively used. After the output result is obtained after the l-th iteration, it is mapped and transformed through a preset matrix to obtain the second encoded result.
[0106] S403. Fuse the first encoded result and the second encoded result to obtain the target encoded result.
[0107] Optionally, the first encoded result and the second encoded result can be concatenated to obtain the target encoded result.
[0108] S404. Input the target encoded result into the decoding sub-model based on the conditional random field, and determine the entity information according to the output of the decoding sub-model.
[0109] Optionally, the conditional random field outputs the optimal output label sequence Y given the input sequence X. Its conditional probability formula is determined by the normalization factor, state features, and transition features. Moreover, in the conditional random field, the Viterbi algorithm can be used to find the label sequence Y that maximizes the conditional probability.
[0110] Optionally, the parameter learning of the conditional random field is determined by using maximum likelihood estimation.
[0111] S302. Input the target text information into the preset relation extraction model to determine the entity relationship.
[0112] Among them, the preset relation extraction model includes a feature fusion sub-model, a cross-attention sub-model, a latent relation prediction sub-model, and a subject-object correspondence sub-model.
[0113] Optionally, as Figure 5 shown, the feature fusion sub-model includes a BiLSTM sub-model and a CDIL-CNN convolutional neural network sub-model. Inputting the target text information into the preset relation extraction model to determine the entity relationship includes the following steps 501 to 502. Among them:
[0114] S501. Respectively extract features from the target text information according to the BiLSTM sub-model and the CDIL-CNN convolutional neural network sub-model to obtain different feature extraction results, and determine the first fusion feature according to each feature extraction result.
[0115] Optionally, optionally, each word is converted into a semantic embedding vector with the help of the pre-trained model GLM-4-9B to generate a word feature vector of length n , and .
[0116] Optionally, before extracting features from the target text information according to the BiLSTM sub-model, the following processing can be performed on the target text information:
[0117] First, tokenize and perform part-of-speech tagging on the target text information to obtain the tokenization and part-of-speech feature results S, where each entry is a tuple containing the token and the corresponding part-of-speech tag Subsequently, each part-of-speech tag is mapped to a part-of-speech ID through a part-of-speech lookup table to obtain the corresponding part-of-speech embedding representation , where the part-of-speech lookup table is provided by a pre-trained part-of-speech embedding matrix PE, and the expression is as follows:
[0118]
[0119] where map represents mapping the part-of-speech tag to the index of the embedding matrix, and the vector retrieved from the part-of-speech embedding matrix PE is , and the dimension of d is the same as the input dimension of the downstream model.
[0120] Optionally, to fully learn the part-of-speech feature information, through a part-of-speech-aware fusion strategy, the activation probability is dynamically adjusted, and the part-of-speech embedding is input into a non-linear activation function to obtain the dynamically adjusted result , which can be specifically represented by the following formula:
[0121]
[0122] where , are trainable weight matrices, , are bias matrices.
[0123] Optionally, the dynamically adjusted result can be concatenated with the word feature vector and then input into the BiLSTM sub-model to generate bidirectional hidden layer vectors and , and then concatenated to obtain the part-of-speech embedding feature corresponding to BiLSTM , which can be represented by the following process:
[0124]
[0125]
[0126] Optionally, there are multi-level dependency relationships in the information of the distribution network production command. Important entity and relationship information may be at the text boundary position. To improve the model's ability to capture complex features and long-range dependencies and reduce boundary effects, by adjusting the dilation rate and padding size of CDIL-CNN, convolutional layers with different dilation rates are used to extract multi-scale feature information, and local patterns and long-range dependencies of the distribution network production command text are captured when the number of layers is small. The expression is as follows:
[0127]
[0128]
[0129] Among them, the word feature vector is output from the pre-trained model GLM-4-9B. The dilated convolution operation is expressed as Conv1d( ), and the residual connection is expressed as Res( ). C( ) is a single-layer symmetric dilated convolution block. C is the number of output channels of the convolution. The multi-layer symmetric dilated convolution is stacked by multiple dilated convolution blocks. The convolution parameter is represented by L, and the dilation rate is represented by d. CP( ) performs the multi-layer symmetric dilated convolution operation. Through circular padding, the boundary information can be completely retained, thereby improving the processing ability of the key information at the boundary position of the distribution network production command.
[0130] Optionally, the part-of-speech embedding feature and the output after semantic enhancement by CDIL-CNN are concatenated to form the first fusion feature.
[0131] Optionally, this feature fusion method not only fuses the grammatical information of part-of-speech tagging and the deep semantic representation, but also integrates the wider context dependencies captured by the dilated convolution network, thereby providing richer and more detailed information for multi-task learning of downstream tasks.
[0132] S502. Determine the second fusion feature according to the first fusion feature and the perturbation data, and input the first fusion feature and the second fusion feature into the cross-attention sub-model to obtain the first feature word vector and the second feature word vector.
[0133] Among them, the perturbation data is determined according to the first fusion feature and the preset adversarial learning sub-model.
[0134] Optionally, there are often a large number of dense entities and complex technical descriptions in the production command text of the distribution network. When processing such texts, the model must have strong robustness and generalization ability to handle various unseen data changes. Therefore, to improve the generalization ability of the model, an adversarial learning method is introduced, and the Fast Gradient Sign Method (FGSM) is used to generate adversarial samples. This method creates perturbations by leveraging the gradient information of the input data, and the perturbation can be represented as follows:
[0135]
[0136] Among them, the relational triple label corresponding to the production command text of the distribution network is y, the first fusion feature is , the parameters of the adversarial learning sub-model are , the loss function is L, the gradient of the loss function with respect to the input is represented by , the sign function is sign, and there is also a preset small constant used to control the perturbation amplitude.
[0137] Optionally, a perturbation can be added to the original input data based on the first fusion feature, , and the adversarial sample second fusion feature is generated.
[0138] Optionally, the initial relational feature word vector and the initial entity feature word vector can be mapped based on the first fusion feature, the second fusion feature, the trainable weight matrix, and the bias matrix.
[0139] Optionally, to improve the accuracy of relation extraction, it is necessary to ensure that the relational features can capture the key information in the entity features. The initial relational feature word vector can be used as the query (Q), and the initial entity feature word vector can be used as the key (K) and value (V) to obtain the multi-head attention scores, and the cross-attention mechanism is used to strengthen the connection between entities, thereby realizing the information interaction between different task feature word vectors.
[0140]
[0141]
[0142] Among them, the input initial relational feature is mapped to the query space by the learnable parameter matrix , while the learnable parameter matrices and are responsible for mapping the initial entity features to the key space and the value space., the cross-attention mechanism uses these mappings to calculate the attention scores between the relationship features and the entity features, and then performs a weighted sum of the value vectors to obtain , and combines with the initial relationship feature word vector to generate the first feature word vector incorporating entity features .
[0143] Similarly, using the initial entity feature word vector as the query (Q), the initial relationship feature word vector as the key (K) and value (V), through the interaction enhancement mechanism, the entity feature word vector extracts information from the relationship feature word vector to generate the second feature word vector incorporating relationship features .
[0144] Optionally, this information interaction enhancement strategy based on the cross-attention mechanism ensures the independence of each task and the flow and complementarity of entity and relationship information.
[0145] S503. Determine multiple potential relationships included in the target text information according to the first feature word vector and the potential relationship prediction sub-model.
[0146] Optionally, there may be multiple relationships in a sentence of the target text information, so the potential relationship prediction can be regarded as a multi-label classification problem. Assuming that the sentence contains n tokens, the first feature word vector is obtained through the pre-trained model GLM-4-9B encoder and the cross-attention mechanism, and its potential relationship prediction process is as follows:
[0147]
[0148]
[0149] Among them, the average pooling result of the first feature word vector is , which can condense the overall sentence information into a one-dimensional vector. The weight parameter matrix is trainable, is the Sigmoid function, which converts the weighted pooled vector into the existence probability of each relationship.
[0150] Optionally, the potential relationship prediction sub-model can determine which potential relationships exist through a preset threshold, and only retain the high-confidence relationships.
[0151] S504. Input each potential relationship and the second feature word vector into the subject-object correspondence sub-model to determine the entity relationship.
[0152] Optionally, after determining the potential relationships, the subject and object entities related to these relationships are extracted. There is an overlap problem among the entities in the production command text of the distribution network. To solve this problem, BIO annotations are set for each determined relationship, and these annotations are used to extract the relevant subject and object entities.
[0153]
[0154]
[0155] Among them, the second feature word vector The encoding representation of the i-th word in it is , and the embedding representation of the j-th relationship is . The trainable weight matrix has . The length of the sequence annotation dictionary {B, I, O} is 3. The probability distributions of the word position i as the subject and object labels are represented by and respectively.
[0156] Optionally, after the relationships and the subjects and objects are determined, the subject-object correspondence module will screen out the correct pairings between the subject and object entities according to the constructed subject-object correspondence matrix, and then integrate them into triples. The subject-object correspondence matrix needs to evaluate all possible subject-object combinations in the text and assign a score to each pair. This score reflects the possibility of forming a valid triple. Suppose the subject-object correspondence matrix is M, and the first feature word vector has n input words, then M . The calculation formula for the score of the entity pair is as follows:
[0157]
[0158] Among them, the word embeddings of the i-th word and the j-th word in the sequence are . The trainable weights and biases used to calculate the subject-object correspondence score are .
[0159] Optionally, during the training process of the preset relationship extraction model, the loss function of the PRGC model can be used. This loss function is composed of the sum of the potential relationship prediction loss, the entity extraction loss, and the subject-object correspondence loss.
[0160] In an exemplary embodiment, as Figure 6 shown, optionally, a knowledge base is determined according to the entity information and entity relationships, including the following steps 601 to step 603. Among them:
[0161] S601, determine the abstract concept layer data and the concept instance layer data according to the entity information.
[0162] Among them, the entity information in the abstract concept layer data includes distribution network generation command, equipment, lines, operation data, equipment name, and equipment number. The entity information in the concept instance layer data is used to represent the specific content of the entity information in the abstract concept layer data.
[0163] Exemplarily, the "distribution network production command" node is a high-level concept, nodes such as "equipment", "lines", and "operation data" are secondary concepts, and "equipment name", "equipment number", "equipment power curve", "standby power supply status", and "load curve" belong to abstract capabilities. The above content is the abstract concept layer data; the entire set of distribution network production command data is a high-level instance, equipment and line instances are secondary instances, the specific name of the equipment is the equipment instance, the specific name of the line is the line instance, and the specific name and data, etc. are the concept instance layer data.
[0164] S602. Determine the data in the capability layer according to the entity relationship.
[0165] Optionally, use the entity relationship as the number of layers in the capability layer.
[0166] S603. Determine the knowledge base according to the abstract concept layer data, concept instance layer data, and capability layer data.
[0167] Optionally, the abstract concept layer data, concept instance layer data, and capability layer data can all be stored in the Neo4j graph database, thereby generating a distribution network knowledge base containing nodes and relationships.
[0168] In another possible implementation manner, the distribution network can also be generated in combination with the structured data stored in the database.
[0169] As an optional implementation manner, as Figure 7 shown, the method for generating a distribution network knowledge base provided in the embodiments of the present application may include the following specific steps:
[0170] S701. Obtain distribution network data.
[0171] Among them, the distribution network data includes distribution network operation data and distribution network production command materials. The distribution network production command materials include various information resources and data for supporting distribution network operation, maintenance, fault handling, and dispatching command. The distribution network data includes structured data and unstructured data.
[0172] S702. Obtain the structured data in the distribution network data and store the structured data in a preset database.
[0173] S703. Obtain the unstructured data in the distribution network data, preprocess the unstructured data, and obtain the target text information.
[0174] S704. Input the target text information into the BERT-based embedding sub-model, and convert each character of the target text information into an embedding vector.
[0175] S705. Input the embedding vector into the first encoding sub-model to obtain the first encoding result, and input the embedding vector into the second encoding sub-model to obtain the second encoding result.
[0176] S706. Fuse the first encoding result and the second encoding result to obtain the target encoding result.
[0177] S707. Input the target encoding result into the decoding sub-model based on the conditional random field, and determine the entity information according to the output of the decoding sub-model.
[0178] S708. Extract features from the target text information according to the BiLSTM sub-model and the CDIL-CNN convolutional neural network sub-model respectively to obtain different feature extraction results, and determine the first fusion feature according to each feature extraction result.
[0179] S709. Determine the second fusion feature according to the first fusion feature and the perturbation data, and input the first fusion feature and the second fusion feature into the cross-attention sub-model to obtain the first feature word vector and the second feature word vector.
[0180] Among them, the perturbation data is determined according to the first fusion feature and the preset adversarial learning sub-model.
[0181] S710. Determine various potential relationships included in the target text information according to the first feature word vector and the potential relationship prediction sub-model.
[0182] S711. Input each potential relationship and the second feature word vector into the subject-object correspondence sub-model to determine the entity relationship.
[0183] S712. Determine the abstract concept layer data and the concept instance layer data according to the entity information.
[0184] Among them, the entity information in the abstract concept layer data includes distribution network generation command, equipment, line, operation data, equipment name, equipment number, and the entity information in the concept instance layer data is used to represent the specific content of the entity information in the abstract concept layer data.
[0185] S713. Determine the capability layer data according to the entity relationship.
[0186] S714. Determine the knowledge base according to the abstract concept layer data, the concept instance layer data and the capability layer data.
[0187] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0188] Based on the same inventive concept, an embodiment of the present application further provides a distribution network knowledge base generation device for implementing the above-mentioned distribution network knowledge base generation method. The implementation solution provided by this device to solve problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the distribution network knowledge base generation device provided below can refer to the limitations on the distribution network knowledge base generation method in the above text, and will not be repeated here.
[0189] In an exemplary embodiment, as Figure 8 shown, a distribution network knowledge base generation device 800 is provided, including: an acquisition module 801, an extraction module 802, and a generation module 803, where:
[0190] The acquisition module 801 is configured to acquire distribution network data, where the distribution network data includes distribution network operation data and distribution network production command materials, and the distribution network production command materials include various information resources and data for supporting distribution network operation, maintenance, fault handling, and dispatching command;
[0191] The extraction module 802 is configured to process the distribution network data based on a preset knowledge extraction model and a preset relationship extraction model to obtain entity information and entity relationships;
[0192] The generation module 803 is configured to determine a knowledge base according to the entity information and the entity relationships.
[0193] In one of the embodiments, the distribution network data includes structured data and unstructured data. Before the preset knowledge extraction model and the preset relationship extraction model, a preprocessing module is further included, which is configured to acquire the structured data in the distribution network data and store the structured data in a preset database; acquire the unstructured data in the distribution network data, and preprocess the unstructured data to obtain target text information.
[0194] In one embodiment, the extraction module 802 is specifically configured to input the target text information into a preset knowledge extraction model to determine entity information. The preset knowledge extraction model includes an embedding sub-model, an encoding sub-model, and a decoding sub-model; input the target text information into a preset relationship extraction model to determine entity relationships. The preset relationship extraction model includes a feature fusion sub-model, a cross-attention sub-model, a latent relationship prediction sub-model, and a subject-object correspondence sub-model.
[0195] In one embodiment, the encoding sub-model includes a first encoding sub-model based on a multi-head self-attention mechanism and a second encoding sub-model based on a dilated convolutional neural network. The extraction module 802 is specifically configured to input the target text information into the BERT-based embedding sub-model to convert each character of the target text information into an embedding vector; input the embedding vector into the first encoding sub-model to obtain a first encoding result, input the embedding vector into the second encoding sub-model to obtain a second encoding result; fuse the first encoding result and the second encoding result to obtain a target encoding result; input the target encoding result into the decoding sub-model based on a conditional random field, and determine entity information according to the output of the decoding sub-model.
[0196] In one embodiment, the feature fusion sub-model includes a BiLSTM sub-model and a CDIL-CNN convolutional neural network sub-model. The extraction module 802 is specifically configured to perform feature extraction on the target text information according to the BiLSTM sub-model and the CDIL-CNN convolutional neural network sub-model respectively to obtain different feature extraction results, and determine a first fusion feature according to each feature extraction result; determine a second fusion feature according to the first fusion feature and perturbation data, and input the first fusion feature and the second fusion feature into the cross-attention sub-model to obtain a first feature word vector and a second feature word vector. The perturbation data is determined according to the first fusion feature and a preset adversarial learning sub-model; determine multiple latent relationships included in the target text information according to the first feature word vector and the latent relationship prediction sub-model; input each latent relationship and the second feature word vector into the subject-object correspondence sub-model to determine entity relationships.
[0197] In one embodiment, the generation module 803 is specifically configured to determine abstract concept layer data and concept instance layer data according to entity information. The entity information in the abstract concept layer data includes distribution network generation command, equipment, line, operation data, equipment name, and equipment number. The entity information in the concept instance layer data is used to represent the specific content of the entity information in the abstract concept layer data; determine capability layer data according to entity relationships; determine a knowledge base according to the abstract concept layer data, the concept instance layer data, and the capability layer data.
[0198] Each module in the above-mentioned power distribution network knowledge base generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in the form of hardware or independent of it, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0199] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for generating a power distribution network knowledge base.
[0200] Those skilled in the art can understand that Figure 9 the structure shown in
[0201] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0202] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps described in any one of the above method embodiments.
[0202] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps described in any one of the above method embodiments.
[0203] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps described in any one of the above method embodiments.
[0204] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0205] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0206] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for generating a distribution network knowledge base, characterized in that: The method comprises: Acquire distribution network data, wherein the distribution network data includes distribution network operation data and distribution network production command information, wherein the distribution network production command information includes various information resources and data used to support distribution network operation, maintenance, fault handling and dispatching command; Based on a preset knowledge extraction model and a preset relationship extraction model, the distribution network data is processed to obtain entity information and entity relationships; A knowledge base is determined according to the entity information and the entity relationship.
2. The method according to claim 1, characterized in that The distribution network data includes structured data and unstructured data. Before the extraction model based on the preset knowledge and the extraction model based on the preset relationship, the method further includes: Acquire structured data in the distribution network data, and store the structured data in a preset database; Unstructured data in the distribution network data is acquired, and the unstructured data is preprocessed to obtain target text information.
3. The method according to claim 2, characterized in that The method of processing the distribution network data based on the preset knowledge extraction model and the preset relationship extraction model to obtain entity information and entity relationships includes: Inputting the target text information into the preset knowledge extraction model to determine the entity information, wherein the preset knowledge extraction model includes an embedding sub-model, an encoding sub-model and a decoding sub-model; The target text information is input into the preset relationship extraction model to determine the entity relationship. The preset relationship extraction model includes a feature fusion sub-model, a cross-attention sub-model, a potential relationship prediction sub-model and a subject-object correspondence sub-model.
4. The method according to claim 3, characterized in that The encoding sub-model includes a first encoding sub-model based on a multi-head self-attention mechanism and a second encoding sub-model based on an expanded convolutional neural network, and the inputting of the target text information into the preset knowledge extraction model to determine the entity information includes: Inputting the target text information into a BERT-based embedding sub-model, and converting each character of the target text information into an embedding vector; Input the embedding vector into the first encoding sub-model to obtain a first encoding result, and input the embedding vector into the second encoding sub-model to obtain a second encoding result; Merging the first encoding result and the second encoding result to obtain a target encoding result; The target encoding result is input into a decoding sub-model based on a conditional random field, and the entity information is determined according to an output of the decoding sub-model.
5. The method according to claim 3, characterized in that: The feature fusion sub-model includes a BiLSTM sub-model and a CDIL-CNN convolutional neural network sub-model, inputs the target text information into the preset relationship extraction model, and determines the entity relationship, including: Performing feature extraction on the target text information according to the BiLSTM sub-model and the CDIL-CNN convolutional neural network sub-model respectively to obtain different feature extraction results, and determining the first fusion feature according to each of the feature extraction results; Determine a second fused feature according to the first fused feature and the perturbation data, and input the first fused feature and the second fused feature into a cross attention sub-model to obtain a first feature word vector and a second feature word vector, wherein the perturbation data is determined according to the first fused feature and a preset adversarial learning sub-model; Determining a plurality of potential relationships included in the target text information according to the first feature word vector and the potential relationship prediction sub-model; Each of the potential relationships and the second feature word vector is input into the subject-object corresponding sub-model to determine the entity relationship.
6. The method according to claim 1, characterized in that Determining a knowledge base according to the entity information and the entity relationship includes: Determine the abstract concept layer data and the concept instance layer data according to the entity information, wherein the entity information in the abstract concept layer data includes the distribution network generation command, equipment, line, operation data, equipment name, and equipment number, and the entity information in the concept instance layer data is used to represent the specific content of the entity information in the abstract concept layer data; Determine capability layer data according to the entity relationship; A knowledge base is determined according to the abstract concept layer data, the concept instance layer data and the capability layer data.
7. A distribution network knowledge base generation device, characterized in that: The device comprises: An acquisition module is used to acquire distribution network data, wherein the distribution network data includes distribution network operation data and distribution network production command information, wherein the distribution network production command information includes various information resources and data used to support distribution network operation, maintenance, fault handling and dispatching command; An extraction module, used to process the distribution network data based on a preset knowledge extraction model and a preset relationship extraction model to obtain entity information and entity relationships; A generation module is used to determine a knowledge base according to the entity information and the entity relationship.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.