Knowledge graph-based text processing method, device and equipment in field of power dispatching
By segmenting entity fragments, obtaining semantic vector features and relation types in the field of power dispatching, and constructing a knowledge graph, the inaccuracy of existing methods is solved, and accurate analysis and response to question texts are achieved.
Patent Information
- Application Number
- CN202511011700.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-11
AI Technical Summary
Existing text processing methods based on knowledge graphs in the field of power dispatching are not accurate enough.
By acquiring text from the power dispatching domain, dividing it into multiple entity segments, obtaining the semantic vector features of each entity segment, detecting the associated entity segments and their relationship types associated with each entity segment, updating the semantic vector features of each entity segment, and constructing a knowledge graph based on these features, we can perform question text analysis to obtain accurate answers.
The constructed knowledge graph is more accurate, enabling it to accurately analyze question texts and provide precise answers.
Smart Images

Figure CN120930741A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text processing technology, and in particular to a text processing method, apparatus and equipment based on knowledge graph in the field of power dispatching. Background Technology
[0002] Knowledge graphs in the power dispatching field are a knowledge representation method that uses a graph structure to describe entities and relationships in a power system. They formally represent entities such as substations, transmission lines, generator units, and loads, as well as their topological connections, electrical parameter relationships, and dispatch control relationships, through nodes and edges. This representation method can intuitively show the physical structure, operating status, and dispatch rules of a power system. Therefore, knowledge graphs are commonly used for text processing in the power dispatching field.
[0003] Currently, the main approach used is based on triples to construct knowledge graphs in the power dispatching field. The main steps are as follows: Knowledge Extraction: 1. Extract entities and relationships from historical data, technical documents, regulations, and standards in the power dispatching field. For example: Entity: "Main Transformer No. 1" Type: Equipment; Entity: "220 kV Transmission Line" Type: Line; 2. Relationship Modeling: Represent the relationships between the extracted entities as triples. For example: Triple: <Main Transformer No. 1, Connection, 220 kV Transmission Line>, Triple: <220 kV Transmission Line, Transmission Power, 200 MW>, etc. Then, the constructed power dispatching knowledge graph is used to answer question texts.
[0004] However, current text processing methods based on knowledge graphs in the field of power dispatching suffer from inaccuracies. Summary of the Invention
[0005] Therefore, it is necessary to provide an accurate knowledge graph-based text processing method, device, computer equipment, computer-readable storage medium, and computer program product for the field of power dispatching, addressing the aforementioned technical problems.
[0006] Firstly, this application provides a knowledge graph-based text processing method in the field of power dispatching, including:
[0007] Obtain text related to power dispatching and the corresponding question text.
[0008] Multiple entity fragments are segmented from texts in the power dispatching domain, and the semantic vector features of each entity fragment are obtained;
[0009] Detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and its associated entity fragments;
[0010] The semantic vector features of each entity fragment are updated based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, as well as the relationship type between the entity fragments and the associated entity fragments.
[0011] Based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments, a knowledge graph of text in the power dispatching domain is constructed.
[0012] Based on the knowledge graph, the question text is analyzed to obtain the question answer text.
[0013] Secondly, this application also provides a knowledge graph-based text processing device in the field of power dispatching, comprising:
[0014] The data acquisition module is used to acquire text in the field of power dispatching and the corresponding question text in the field of power dispatching.
[0015] The entity segmentation module is used to segment multiple entity fragments from power dispatching domain text and obtain the semantic vector features of each entity fragment.
[0016] The relationship identification module is used to detect the associated entity fragments of each entity fragment, as well as the relationship type between each entity fragment and its associated entity fragments;
[0017] The semantic update module is used to update the semantic vector features of each entity fragment based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, as well as the relationship type between the entity fragments and the associated entity fragments.
[0018] The graph construction module is used to construct a knowledge graph of text in the power dispatching domain based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments.
[0019] The question reasoning module is used to analyze the question text based on the knowledge graph and obtain the question answer text.
[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0021] Obtain text related to power dispatching and the corresponding question text.
[0022] Multiple entity fragments are segmented from texts in the power dispatching domain, and the semantic vector features of each entity fragment are obtained;
[0023] Detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and its associated entity fragments;
[0024] The semantic vector features of each entity fragment are updated based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, as well as the relationship type between the entity fragments and the associated entity fragments.
[0025] Based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments, a knowledge graph of text in the power dispatching domain is constructed.
[0026] Based on the knowledge graph, the question text is analyzed to obtain the question answer text.
[0027] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0028] Obtain text related to power dispatching and the corresponding question text.
[0029] Multiple entity fragments are segmented from texts in the power dispatching domain, and the semantic vector features of each entity fragment are obtained;
[0030] Detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and its associated entity fragments;
[0031] The semantic vector features of each entity fragment are updated based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, as well as the relationship type between the entity fragments and the associated entity fragments.
[0032] Based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments, a knowledge graph of text in the power dispatching domain is constructed.
[0033] Based on the knowledge graph, the question text is analyzed to obtain the question answer text.
[0034] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0035] Obtain text related to power dispatching and the corresponding question text.
[0036] Multiple entity fragments are segmented from texts in the power dispatching domain, and the semantic vector features of each entity fragment are obtained;
[0037] Detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and its associated entity fragments;
[0038] The semantic vector features of each entity fragment are updated based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, as well as the relationship type between the entity fragments and the associated entity fragments.
[0039] Based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments, a knowledge graph of text in the power dispatching domain is constructed.
[0040] Based on the knowledge graph, the question text is analyzed to obtain the question answer text.
[0041] The aforementioned knowledge graph-based text processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product in the power dispatching field acquires power dispatching text and corresponding question text; divides the power dispatching text into multiple entity segments and obtains the semantic vector features of each entity segment; detects the associated entity segments of each entity segment and the relationship types between each entity segment and its associated entity segments; updates the semantic vector features of each entity segment based on its corresponding associated entity segments and the relationship types; constructs a knowledge graph of the power dispatching text based on the updated semantic vector features of all entity segments and the relationship types between different entity segments; and analyzes the question text using the knowledge graph to obtain the question answer text. Throughout this process, the semantic vector features of each entity segment are integrated with the semantic vector features of associated entity segments of different relationship types. Therefore, the constructed knowledge graph is more accurate, allowing for accurate analysis of the question text and the generation of the question answer text. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is an application environment diagram of a knowledge graph-based text processing method in the field of power dispatching, as shown in one embodiment.
[0044] Figure 2 This is a flowchart illustrating a knowledge graph-based text processing method in the field of power dispatching, as shown in one embodiment.
[0045] Figure 3 This is a flowchart illustrating a knowledge graph-based text processing method in the field of power dispatching, as shown in another embodiment.
[0046] Figure 4 This is a structural block diagram of a knowledge graph-based text processing device in the field of power dispatching, as shown in one embodiment.
[0047] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.
[0049] The knowledge graph-based text processing method in the field of power dispatching provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server.
[0050] The user enters the question text corresponding to the power dispatch field text on the text processing interface of terminal 102 and triggers the text processing control. Terminal 102 responds to the trigger request of the text processing control, generates a text processing request, carries the question text corresponding to the power dispatch field text, and sends the text processing request to server 104.
[0051] Server 104 receives a text processing request, retrieves the question text corresponding to the power dispatching domain text, and obtains the power dispatching domain text from the database. It then divides the power dispatching domain text into multiple entity fragments and obtains the semantic vector features of each entity fragment. It detects the associated entity fragments of each entity fragment and the relationship types between each entity fragment and its associated entity fragments. Based on the semantic vector features of each entity fragment's corresponding associated entity fragments and the relationship types, it updates the semantic vector features of each entity fragment. Based on the updated semantic vector features of all entity fragments and the relationship types between different entity fragments, it constructs a knowledge graph of the power dispatching domain text. Based on the knowledge graph, it analyzes the question text to obtain the question answer text. Furthermore, server 104 pushes the question answer text to terminal 102 for display to the user.
[0052] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0053] In one exemplary embodiment, such as Figure 2 As shown, a knowledge graph-based text processing method for power dispatching is provided, which is then applied to... Figure 1 Let's take server 104 as an example for explanation. Specifically:
[0054] S100, obtain the power dispatching domain text and the corresponding question text.
[0055] The texts in the power dispatching field include at least several professional texts such as dispatching regulations of power grid companies and power industry standard documents. The corresponding question texts in the power dispatching field refer to questions raised in response to these texts, such as: fault diagnosis questions: "What measures should be taken when a transformer is overloaded?", operation verification questions: "Is this operation sequence safe?", and risk assessment questions: "Which loads will be affected by shutting down a certain line?", etc.
[0056] Specifically, the user enters the question text corresponding to the power dispatching domain text on the terminal's text processing interface and triggers the text processing control. The terminal responds to the trigger request of the text processing control, generates a text processing request, carries the question text corresponding to the power dispatching domain text, and sends the text processing request to the server. The server receives the text processing request, obtains the question text corresponding to the power dispatching domain text, and retrieves the power dispatching domain text from the database.
[0057] For the acquired texts in the field of power dispatching, this application designs a complete multi-source heterogeneous data management system. The system on the server can simultaneously access multiple data sources, including power dispatching procedure documents (approximately 100,000 words), equipment manuals (approximately 2,000 copies), historical fault cases (approximately 10,000 entries), and expert knowledge bases (approximately 10,000 entries). Considering the heterogeneity and scale of the data, the system adopts a distributed storage architecture, storing document-type data in a document database, time-series data in a time-series database, and knowledge data in a graph database.
[0058] In one embodiment, after acquiring power dispatching domain text, it is necessary to preprocess the text, including basic processing such as special character handling, number standardization, and unit unification. Regarding data standardization, the system adopts the Z-score standardization method to unify data with different dimensions to the same scale space, facilitating subsequent model processing. During standardization, the system retains the unit information and timestamps of the original data to ensure data traceability. To guarantee data quality, the system also designs a rule-based and statistical anomaly detection mechanism to mark and filter abnormal data.
[0059] The rule-based anomaly detection mainly includes: 1. Physical constraint rules for electrical parameters. For example, voltage levels must conform to standard values, such as 110 kV, 220 kV, and 500 kV. If a voltage value of 118 kV or a negative value appears, the system will immediately mark it as an anomaly. Current and power values cannot exceed the rated capacity range of the equipment. For example, a generator set with a rated capacity of 100 MW should not have its output power displayed as 150 MW. 2. Logical consistency rules for equipment status. The status of a circuit breaker can only be "closed" or "open," and cannot be in an intermediate state such as "partially closed." At the same time, if a circuit breaker on a certain line is in the open state, the current value of that line should be zero. If the data shows that the circuit breaker is open but current is still flowing, it is judged as a logical anomaly. 3. Reasonableness of the timing sequence of operations. For example, a power outage operation requires disconnecting the load-side switch first, and then disconnecting the power-side switch. If the order is reversed, it is an abnormal operation sequence.
[0060] Statistical anomaly detection is determined by establishing a statistical model of historical data and calculating the normal fluctuation range of various parameters. If a data point deviates from the historical mean by more than three standard deviations, it will be marked as a statistical anomaly. For example, the daily load of a substation typically fluctuates between 50 and 80 megawatts. If a value of 10 megawatts or 200 megawatts suddenly appears on a certain day, and there is no corresponding record of operational adjustments, it will be identified as an anomaly.
[0061] S200 divides multiple entity segments from power dispatching domain text and obtains the semantic vector features of each entity segment.
[0062] Specifically, this application employs a pre-trained language model combined with a multi-task learning strategy to process power dispatching texts from different sources, thereby achieving accurate entity identification and extraction from the power dispatching texts and obtaining multiple entity fragments. Generally, considering the specialized and complex nature of power dispatching texts, RoBERTa-large is selected as the basic pre-trained language model. Furthermore, this application also requires semantic feature vector extraction for each entity text fragment to obtain the semantic vector features of each entity fragment.
[0063] S300, detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and its associated entity fragments.
[0064] Specifically, the process involves detecting the associated entity fragments of each entity fragment. Associated entity fragments refer to neighboring entity fragments that are directly related to the entity fragment. Further, it's necessary to detect the relationship type between each entity fragment and its associated entity fragments. In practical applications, relationship types are mainly categorized as follows: physical connection relationships, parameter constraint relationships, operational dependency relationships, causal influence relationships, functional composition relationships, temporal sequence relationships, state transition relationships, and other relationships. For example, if A is an entity fragment and B is an associated entity fragment of A, then A-connection-B is a physical connection relationship, A-control-B is an operational control relationship, and A-influence-B is a parameter influence relationship, etc. The proportion of different relationship types varies.
[0065] S400, based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, and the relationship type between them, update the semantic vector features of each entity fragment.
[0066] Specifically, for each associated entity fragment, the associated entity fragments can be classified according to the relationship type between them. Based on the semantic vector features of each associated entity fragment, the semantic vector features of the associated entity fragments corresponding to different relationship types after classification are obtained. Then, the semantic vector features of the entity fragments are updated by combining the semantic vector features of the associated entity fragments corresponding to different relationship types after classification, as well as the semantic vector features of the entity fragment itself. In other words, the semantic vector features of the updated entity fragments not only contain its own semantic vector features, but also the semantic vector features of the associated entity fragments corresponding to different relationship types after classification.
[0067] S500 constructs a knowledge graph of text in the power dispatching domain based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments.
[0068] Specifically, each entity fragment is used as a node in the knowledge graph, the relationship type between different entity fragments is used as the connection relationship between different nodes, and the updated semantic vector features corresponding to each entity fragment are used as the attributes of the nodes. Based on all nodes, the attributes of all nodes, and the connection relationships between different nodes, a knowledge graph for text in the power dispatching domain is constructed.
[0069] S600 analyzes the question text based on the knowledge graph to obtain the question answer text.
[0070] Specifically, when analyzing question text based on a knowledge graph, the optimal solution path related to the question text is searched from the knowledge graph. Then, based on the optimal solution path, the question answer text is determined. For example, if the question text is "Where is device A purchased?", the target entity fragment associated with the entity fragment "device A" through the relation type "purchase location" can be found in the knowledge graph, and the question answer text of the question text can be obtained from the target entity fragment.
[0071] The aforementioned knowledge graph-based text processing method in the power dispatching field involves: obtaining the power dispatching text and its corresponding question text; dividing the power dispatching text into multiple entity segments and obtaining the semantic vector features of each entity segment; detecting the associated entity segments of each entity segment and the relationship types between each entity segment and its associated entity segments; updating the semantic vector features of each entity segment based on its corresponding associated entity segments and the relationship types; constructing a knowledge graph of the power dispatching text based on the updated semantic vector features of all entity segments and the relationship types between different entity segments; and analyzing the question text using the knowledge graph to obtain the question answer text. Throughout this process, the semantic vector features of each entity segment are integrated with the semantic vector features of associated entity segments of different relationship types. Therefore, the constructed knowledge graph is more accurate, allowing for accurate analysis of the question text and the resulting question answer text.
[0072] In one exemplary embodiment, multiple entity fragments are segmented from power dispatch domain text, including:
[0073] The system acquires multiple entity boundary sequences from text in the power dispatching domain. Each entity boundary sequence contains boundary location information and label information of multiple entities. It detects the sequence conditional probability corresponding to each entity boundary sequence and determines the target entity boundary sequence with the highest sequence conditional probability from multiple entity boundary sequences. The target entity boundary sequence contains multiple label information that satisfies preset label constraints. The system then detects multiple entity fragments corresponding to the target entity boundary sequence from the text in the power dispatching domain.
[0074] Among them, label information refers to the category information of an entity. For example, label information can be categories such as equipment or power parameters.
[0075] Specifically, this application can use a pre-trained language model combined with a multi-task learning strategy to achieve accurate entity recognition and extraction in texts in the power industry.
[0076] In the pre-training process of pre-trained language models, dynamic masking strategies can be used to process training data in the power dispatching field. Dynamic masking is a core technology of pre-trained models such as BERT, enabling the model to learn deep semantic representations and contextual understanding. Specifically, dynamic masking randomly masks parts of the input text, then trains the model to predict the masked words based on contextual information. This training method forces the model to understand the semantic relationships and grammatical structures between words, thereby gaining strong language understanding capabilities. For example, when the text contains "500 kV substation main transformer [MASK] operation," the model needs to infer from the context that the masked words might be professional terms such as "commissioning," "maintenance," or "outage." Through this training, the model can master the language patterns and professional expression habits in the power field. In this example, the masking rate is set to 15%, with 80% replaced by [MASK] tags, 10% replaced by random words, and 10% remaining unchanged.
[0077] Furthermore, the pre-training tasks also incorporate power-specific tasks, such as power equipment name prediction and electrical parameter value prediction. The power equipment name prediction task is specifically designed for entity recognition needs in the power sector. Equipment naming in power systems has strict norms and hierarchical structures. For example, a name like "500 kV Jia Substation No. 1 Main Transformer" includes multi-dimensional information such as voltage level, geographical location, equipment type, and serial number. Through equipment name prediction training, the model can learn the inherent rules of power equipment naming. When the model encounters text like "220 kV Jia Substation No. 2 [MASK]", it can infer the possible equipment type based on the voltage level and substation environment, such as "main transformer," "circuit breaker," or "disconnector." Electrical parameter prediction... The parameter prediction task allows the model to learn the physical constraints and numerical patterns between various parameters in the power system, such as the relationship between power, voltage, and current, or the transformation ratio relationship between different voltage levels. Through parameter prediction training, the model can understand these physical constraints. For example, when the power dispatching field describes "the voltage on the high-voltage side of the main transformer is 220 kV, and the transformation ratio is 220 / [MASK]", the model can infer from common transformer transformation ratios that the low-voltage side may be "10 kV" or "35 kV". This training enables the model to have basic power system knowledge and to perform rationality verification in the subsequent relationship extraction and reasoning process.
[0078] In addition, the pre-training process uses the Adam optimizer. Generally, the learning rate is set to 2e-5, the batch size is 32, and the number of training rounds is 10, which is not limited here.
[0079] During the training process of the pre-trained model, the training data in the power dispatching domain needs to be manually labeled, covering multiple major categories such as equipment entities (e.g., "circuit breaker," "transformer"), parameter entities (e.g., "voltage level," "rated current"), and operation entities (e.g., "switching operation," "voltage adjustment"). Each power dispatching domain text label undergoes multiple rounds of manual review to ensure labeling quality. Generally, the labeling adopts the BIO (Begin-Inside-Outside, entity start position - entity middle or end position - non-entity position) tagging scheme. Furthermore, this application also sets up a validation set and a test set, each containing 10,000 data points, maintaining the same category distribution as the training set. In practical applications, the training data dataset format can be: whether it is a key entity / ordinary entity, entity name, entity category.
[0080] The BIO tagging scheme is based on three core tags used to mark the position of the token within the entity:
[0081] B-XXX: Indicates the starting position of the entity, where XXX is the entity type (e.g., person name, place name); I-XXX: Indicates the internal position of the entity, belonging to the subsequent part of the current entity; O: Indicates the non-entity part, not belonging to any target entity. Here, the token is the product of text segmentation, which can be understood as the "smallest processing unit of text." It can be a word, a subword, a character, or even a punctuation mark. During the training phase, the model learns how to assign high probabilities to correct BIO tagging schemes and low probabilities to incorrect BIO tagging schemes.
[0082] The pre-trained language model is trained using training data to obtain a fully trained pre-trained language model. This pre-trained language model is then used to perform entity recognition on text in the power dispatching domain. In practical applications, the pre-trained language model is typically a BiLSTM-CRF (Bidirectional Long Short-Term Memory-Conditional Random Field) model. This is a deep learning model commonly used in sequence labeling tasks in natural language processing, combining the advantages of BiLSTM and CRF.
[0083] Therefore, the entity recognition process includes: First, preliminary boundary identification is performed using a pre-trained language model. Based on the BIO annotation scheme used during training, the pre-trained language model acquires multiple entity boundary sequences from the power dispatching domain text. Then, a conditional probability is assigned to each of these sequences to find the optimal entity boundary annotation scheme. This probability represents the likelihood of a specific entity boundary sequence Y appearing given the input power dispatching domain text X. The BiLSTM-CRF model evaluates the rationality of different entity boundary annotation schemes using conditional probability and selects the entity boundary sequence with the highest conditional probability from among the various possible sequences as the target entity boundary sequence. Here, the entity boundary sequence Y is a BIO sequence of the same length as the power dispatching domain text X. The method for selecting the entity boundary sequence with the highest conditional probability can be a dynamic programming method such as the Viterbi algorithm.
[0084] Specifically, if a text X in the power dispatch domain has n entity tokens, then the entity boundary sequence Y = {y1, y2, ..., y...} n}, where each y i All of these are BIO tags, and possible values include: B-device, I-device; B-parameter, I-parameter; B-operation, I-operation.
[0085] Secondly, during the detection of target entity boundary sequences, a CRF layer is needed to determine whether the multiple label information in each entity boundary sequence meets preset label constraints. These preset constraints include, for example, that I-device labels cannot directly follow O labels; B-device labels must precede them to obtain a more accurate target entity boundary sequence. Finally, multiple entity fragments corresponding to the target entity boundary sequence are detected from power dispatching domain text.
[0086] Furthermore, for the text X={x1, x2, ..., x...} in the field of power dispatching... n The conditional probability calculation for the BiLSTM-CRF layer is: P(Y|X)=exp(Score(X,Y)) / Z(X), where P(Y|X) is the sequence conditional probability corresponding to each entity boundary sequence, Y is the entity boundary sequence, and Score(X,Y) is the path score. Z(X) is the normalization factor: Z(X) = ∑yexp(Score(X,y)), where y is the label sequence in the entity boundary sequence.
[0087] It should be explained that the path score considers both observed features (the association between words and labels) and constraints between labels (such as BIO rules), enabling the model to more accurately identify entity boundaries. Here, Emit(i) represents the relationship between the observed value (e.g., word) at position i and the label information y.i The degree of matching, for example, "Beijing" as a word is more likely to correspond to the "B-LOC" tag; Trans(y i-1 y i ) indicates that the label sequence starts from y i-1 Transfer to y i The probability (transfer score) is such that, for example, "B-LOC" can only be followed by "I-LOC" or "O", not "B-PER".
[0088] In one embodiment, the label information in the entity boundary sequence is initial information. The label information for each entity segment can be further updated after detecting multiple entity segments corresponding to the target entity boundary sequence. This step can be implemented using a span classifier or other classifiers, which are not limited here. In this application, by decoupling boundary detection and category classification, each stage can focus on solving its own problem; secondly, classification implemented using a span classifier or similar methods can utilize global information of the entire entity segment, rather than local information of a single token; finally, for complex and long entity names existing in the power industry, this method can better handle their internal hierarchical structures.
[0089] The method for updating the label information of each entity segment includes: using a classifier to detect the class probabilities of multiple possible label information for each entity segment, and selecting the label information with the highest class probability as the updated target label information. In practical applications, the class probability directly reflects the classifier's confidence in the classification result. If the probability of a segment being classified as "transformer" is 0.95, while the probability of being classified as "circuit breaker" is 0.03, then the classifier is confident that "transformer" is the target label information. However, if the probabilities of each class are relatively close, such as "transformer" 0.4, "circuit breaker" 0.35, and "busbar" 0.25, it indicates that the classifier is not certain enough about the target label information of the entity segment, and manual review may be required.
[0090] The expression for the class probability of a certain label information c of each entity segment span can be: P(c|span) = softmax(W·h) span +b), where W and b are the weight matrix and bias of the classifier, respectively; softmax: converts the scores into a probability distribution, ensuring that the sum of the probabilities of all classes is 1, h span These are the semantic vector features of the span, calculated using an attention mechanism: ,in, Let W be the hidden state at position i, W be the learnable weight matrix, and v be the learnable weight vector used to calculate the attention score. Let be the attention weight of the i-th token, representing its importance to Span, and tanh be the activation function that compresses the vector values to the interval [-1, 1].
[0091] The final loss function L incorporates the sequence labeling loss L. seq And span classification loss L span :L=λ1L seq +λ2L span λ1 and λ2 are balancing factors, and their optimal values are obtained through validation set tuning: λ1 = 0.6 and λ2 = 0.4. The loss function L can be used to further update the target entity boundary sequence to obtain more accurate multiple entity fragments.
[0092] In the above embodiments, by determining the target entity boundary sequence with the highest sequence condition probability from multiple entity boundary sequences and whose multiple label information satisfies the preset label constraint conditions, multiple entity fragments corresponding to the target entity boundary sequence can be accurately detected from power dispatch domain text.
[0093] In an exemplary embodiment, after detecting multiple entity fragments corresponding to a target entity boundary sequence from power dispatch domain text, the method includes:
[0094] For each entity fragment, search the preset standard library for standard text fragments similar to the entity fragment and the corresponding standard associated entity fragments; verify the accuracy of the entity fragment based on the standard text fragments and standard associated entity fragments; if the accuracy verification of the entity fragment fails, update the entity fragment based on the standard text fragments and standard associated entity fragments.
[0095] Specifically, to improve the accuracy of entity extraction, this application also introduces a domain knowledge constraint mechanism. Multiple entity fragments corresponding to the target entity boundary sequence identified by the model are processed using a pre-built power equipment ontology library (containing over 10,000 standard equipment names) and parameter specification library (containing over 500 standard parameter definitions). This method, combining domain knowledge, significantly improves the accuracy of entity recognition, especially for the recognition of technical terms and complex equipment names.
[0096] The specific processing flow includes: 1. Fuzzy matching stage: Using algorithms such as edit distance and cosine similarity, standard text fragments similar to entity fragments and corresponding standard associated entity fragments are searched from a preset standard library; 2. Rule verification stage: Applying professional rules in the power field, the accuracy of entity fragments is verified based on standard text fragments and standard associated entity fragments; 3. Confidence recalculation: The final confidence score is recalculated by combining the matching degree between standard text fragments and standard associated entity fragments; 4. Standardized output: The verified entity fragments are converted into standard format output. When the accuracy verification of an entity fragment fails, the entity fragment is updated based on the standard text fragments and standard associated entity fragments.
[0097] Furthermore, this application can also employ more methods to detect entity fragments, such as:
[0098] The first layer: Standardized matching and error correction: Match the entity fragments identified by the model with the standard library. The matching involves processing such as case normalization (kilovolt → kilovolt) and abbreviation expansion (main transformer → main transformer).
[0099] The second layer: Semantic consistency verification: Verifying the semantic rationality of the identification results. The power equipment ontology library not only contains standard names but also equipment attribute information, such as rated voltage and applicable scenarios. When the model identifies a statement like "a 10 kV circuit breaker connected to a 500 kV busbar," the system will detect a voltage level mismatch, because 10 kV equipment is usually not directly connected to a 500 kV system. This may be an identification error or requires further confirmation.
[0100] The third layer: Context constraint verification: Verifies the rationality of entity recognition based on contextual information. For example, in a document describing substation equipment, if the model identifies "generator," the system will determine whether this is reasonable based on the context. This is because generators are usually found in power plants, not substations. If it does appear in a substation environment, it might be short for "generator truck."
[0101] Fourth layer: Hierarchical structure completion: Power equipment usually has a hierarchical structure. When the model only recognizes part of the hierarchical information, the system will try to complete the hierarchical structure based on the standard ontology library.
[0102] Fifth layer: Parameter constraint verification: The parameter specification library contains standard definitions and reasonable ranges for various electrical parameters. When the model identifies a parameter entity, the system verifies the reasonableness of its value.
[0103] In the above embodiments, the accuracy of entity fragments is verified through various methods, which can further improve the detection accuracy of entity fragments.
[0104] In an exemplary embodiment, the semantic vector features of each entity fragment are updated based on the semantic vector features of its corresponding associated entity fragments and the relationship type with the associated entity fragments, including:
[0105] For each entity fragment, detect at least one target associated entity fragment belonging to the same relation type from the associated entity fragments; detect the category association semantic vector features of the entity fragment based on the semantic vector features of at least one target associated entity fragment of the same relation type; update the semantic vector features of the entity fragment based on the category association semantic vector features corresponding to different relation types, wherein the updated semantic vector features contain both the category association semantic vector features corresponding to different relation types and the semantic vector features of the entity fragment.
[0106] Specifically, the relation extraction model design adopts a multi-head interaction architecture based on graph neural networks. Considering the various types of relationships between entities in the power dispatching domain (such as topological relationships between equipment, constraint relationships between parameters, and sequential relationships between operations), the relation extraction model uses a relation graph convolutional network as its basic structure and introduces an attention mechanism to enhance the capture capability of relation features. In each relation graph convolutional layer, the update formula for the semantic vector features of entity fragments is:
[0107]
[0108] Among them, h i (l) represents the semantic vector feature of entity fragment i in the l-th layer of the relation graph convolutional network, h j (l) represents the semantic vector feature of entity fragment j in the l-th layer of the relational graph convolutional network, h i (l+1) represents the semantic vector feature of entity fragment i at the (l+1)th layer in the relational graph convolutional network, R is the set of all relation types, and N is the set of all relation types. i r represents the target associated entity fragment directly connected to entity fragment i through relation type r, c (i,r) W is the normalization constant. r (l) is the specific transformation matrix corresponding to the relation type, W0(l) is the learnable weight matrix of the self-loop connection, which preserves the node's own information, and σ is the activation function. The relation extraction model consists of 3 layers of R-GCN, with a hidden layer dimension of 512.
[0109] The representation grouped all associated entity fragments of entity fragment i according to relation type r, obtaining the target associated entity fragment corresponding to each relation type r, and the semantic vector feature h of each target associated entity fragment. j (l) First, use the specific weight matrix W. r(l) Transform to obtain the category association semantic vector features corresponding to each relation type r, and then perform a weighted summation according to all relation types r, with a weight of 1 / c. (i,r) ,W0(l)h i (l) refers to the transformation of the entity fragment i's own representation through the weight matrix W0(l), preserving the inherent features of entity fragment i. Finally, the two results are non-linearly activated, so that the updated semantic vector features simultaneously contain the category association semantic vector features corresponding to different relation types and the semantic vector features of entity fragments. In one embodiment, the performance of updating the semantic vector features of each entity fragment according to the associated entity fragments of different relation types is shown in Table 1 below, where the F1 value is the combined precision and recall value.
[0110]
[0111] In the above embodiments, by aggregating associated entity fragments of different relation types, the semantic vector features of each entity fragment not only include its own semantic vector features, but also incorporate the semantic information of associated entity fragments, forming a richer semantic representation. Furthermore, in the process of aggregating associated entity fragments of different relation types, the different degrees of influence of different relation types on the semantic vector features of entity fragments can be learned through the specific transformation matrix corresponding to the relation type.
[0112] In an exemplary embodiment, before analyzing the question text according to a knowledge graph to obtain the question answer text, the method includes:
[0113] Based on preset relational rules, the relationship types between multiple entity fragments in the knowledge graph are initially screened. Then, based on these rules, a probability prediction is made to determine the correctness of each initially screened relationship type, yielding a rule-predicted probability. Next, based on a preset classification model, a model-predicted probability is made to determine the correctness of each initially screened relationship type. Finally, the confidence assessment results for each initially screened relationship type are checked based on the rule-predicted and model-predicted probabilities. Based on the confidence assessment results, a second screening is performed on the relationships between the initially screened relationship types. Finally, the consistency assessment results for the second screening are checked, and based on the consistency assessment results, a third screening is performed to obtain the updated knowledge graph.
[0114] Specifically, to ensure the accuracy of relation identification before the actual problem reasoning process, a multi-level relation verification algorithm was designed. The verification process includes three levels:
[0115] The first layer is rule-based validation, which performs preliminary screening of the relationship types between multiple entity fragments in the knowledge graph based on preset relationship rule constraints. The rules include at least: equipment type compatibility rules, electrical parameter validity rules, operation sequence rationality rules, and topology connection feasibility rules.
[0116] The second layer is probability-based validation, calculating the confidence score for each relation type. This involves obtaining the rule-predicted probability of each initially filtered relation type being correct, based on pre-defined relation rule constraints. For example, each relation rule constraint corresponds to a probability value; if an initially filtered relation type violates a rule constraint, the corresponding probability value is subtracted from the initial rule-predicted probability. Next, based on a pre-defined classification model, the model-predicted probability of each initially filtered relation type is obtained. For instance, the model-predicted probability comes directly from the softmax output of the pre-defined classification model, providing a model-predicted probability between 0 and 1 for each relation type, with the sum of all model-predicted probabilities being 1. Finally, based on the rule-predicted probability and the model-predicted probability, the confidence evaluation result for each initially filtered relation type is determined, expressed as: Score(r) = α·P model(r) +β·P rule(r) , where: P model(r) For the model to predict probabilities, P rule(r) The probability is predicted for the rule, and α, β, and γ are the weight coefficients. The optimal values are determined by grid search. For example, α=0.5, β=0.3, and γ=0.2.
[0117] The third layer is a verification based on global consistency, constructing a relational reasoning graph G=(V, E), where V is the entity set and E is the relation set. For each relation type r after secondary filtering, its global consistency evaluation result Consistency(r)=∑ p∈P w(p)·valid(p), where P is the set of all inference paths containing relation type r in the relation inference graph, w(p) is the weight of the p-th inference path, w(p) = 1 / path length. The shorter the path, the higher the weight of the inference path. Therefore, the inference of a short path is more reliable. valid(p) is the path validity index of the p-th inference path, valid(p) = 1 (valid) or 0 (invalid). If any relation type in the inference path violates the basic power rule, then valid(p) = 0, otherwise valid(p) = 1.
[0118] Finally, after three rounds of filtering, the relationship types that failed the verification were filtered out, resulting in the updated knowledge graph.
[0119] In the above embodiments, erroneous or unreasonable relation types in the initial knowledge graph are filtered to ensure that the relations stored in the knowledge graph are credible, providing a reliable knowledge foundation for subsequent reasoning, reducing the path space that needs to be searched during reasoning, and improving reasoning efficiency. Furthermore, the multi-level verification mechanism used in the filtering makes the reasoning results based on the updated knowledge graph more credible, especially in fields such as power dispatching where security requirements are extremely high.
[0120] In one exemplary embodiment, such as Figure 3 As shown, S600 includes:
[0121] S610: Obtain the meta-path template of the question text, and in the knowledge graph, query the initial entity fragment corresponding to the question text based on the question text.
[0122] S620, Filtering step: Based on the metapath template, detect multiple next candidate entity fragments associated with the initial entity fragment.
[0123] S630, based on the semantic vector features of the initial entity fragment and the semantic vector features of each next candidate entity fragment, detect the importance quantification value of each next candidate entity fragment, and determine the next candidate entity fragment with the largest importance quantification value as the next entity fragment.
[0124] S640, determine the next entity fragment as the initial entity fragment, return to the filtering step, until no next entity fragment is detected.
[0125] S650: Based on the initial entity fragment, all detected next entity fragments, and the relationship type between all entity fragments, obtain the solution path corresponding to the problem text, and based on the solution path, obtain the problem response text of the problem text.
[0126] Specifically, when analyzing question text based on knowledge graphs, the solution path related to the question text is searched from the knowledge graph, and then the answer text of the question text is determined based on the solution path.
[0127] To improve search efficiency, we designed a multi-level search strategy:
[0128] 1. First level: Fast search based on metapaths: Retrieves the metapath template of the question text. The metapath template is a set. Each metapath template defines a valid relation sequence pattern. During the search, solutions conforming to the metapath template are given priority. For example, for a fault impact analysis problem, metapath templates like "device-connection-device-impact-device" are preferred; for an operation sequence verification problem, metapath templates like "operation-dependency-operation-sequence-operation" are preferred. Based on the metapath template, multiple next candidate entity fragments associated with the initial entity fragment are detected. Furthermore, when obtaining metapath templates for the problem text, metapath templates are filtered using an importance score. The importance score is I(m) = α·F(m) + β·P(m), where F(m) is the frequency of the metapath template, P(m) is the historical accuracy of the metapath template, and α and β are the weights corresponding to the frequency and historical accuracy of the metapath template, respectively.
[0129] F(m) = Number of times the metapath template appears in the historical graph data / Total number of times all different metapath templates appear in the historical graph data. For example, the metapath template "Device-Connection-Device-Power Supply-Load" appears 200 times in the historical graph data, and all metapath templates appear a total of 1000 times, so F(m) = 200 / 1000 = 0.2. Historical accuracy P(m) is calculated as follows: P(m) = Number of times the metapath template was used correctly in reasoning / Total number of times the metapath template was used in reasoning. For example, the metapath template "Device-Connection-Device" was used for reasoning 100 times, and 85 of those results were correct, so P(m) = 85 / 100 = 0.85.
[0130] 2. Second Level: Attention-Based Path Expansion. When expanding nodes, an attention mechanism is used to calculate the importance quantification value of the next hop: at(v next |v current = softmax(W·[h current ; h next ] + b), where h current and h next Here, W represents the semantic vector features of the initial entity fragment and the next candidate entity fragment, respectively. W is a learnable vector matrix, b is a bias matrix, and softmax is the activation function. Therefore, based on the semantic vector features of the initial entity fragment and the semantic vector features of each next candidate entity fragment, the importance quantization value of each next candidate entity fragment can be detected, and the next candidate entity fragment with the highest importance quantization value can be determined as the next entity fragment.
[0131] The next entity fragment is designated as the initial entity fragment, and the above two-level search strategy is repeated until no next entity fragment is detected. At this point, the relationship types of all entity fragments between the initial entity fragment and all detected next entity fragments are obtained. Based on the initial entity fragment, all detected next entity fragments, and the relationship types of all entity fragments, the solution path corresponding to the question text is obtained. Based on the solution path, the question answer text is obtained.
[0132] In one embodiment, the application further includes setting a third-level path search strategy, namely setting an adaptive threshold θ to prune paths in the knowledge graph whose scores are below the threshold: θ = μ + σ·log(d), where μ is the average historical path score, σ is the standard deviation, and d is the depth of path search.
[0133] In the above embodiments, the path expansion method using the meta-path template method and the attention mechanism can select the corresponding solution path at the macro-level strategy level through the meta-path template method, and specifically select the next entity fragment in the solution path at the micro-level node selection level through the attention mechanism path expansion method, so as to select the most valuable expansion direction and obtain the solution path corresponding to the problem text, thereby accurately obtaining the problem answer text.
[0134] In one exemplary embodiment, the problem text includes types such as fault diagnosis reasoning, operation verification reasoning, risk assessment reasoning, and knowledge completion reasoning. The performance of the solution path determined by this method for different types of problem text is shown in Table 2 below:
[0135]
[0136] In one exemplary embodiment, before obtaining the question answer text based on the resolution path, the process includes:
[0137] The system detects the confidence level of the relationships between multiple entity segments in the solution path, the path length of the solution path, and the frequency of occurrence of the solution path in a pre-defined path library. Based on multiple confidence levels, it detects the path reliability of the solution path, the path length penalty information of the solution path based on the path length, and the path novelty of the solution path based on the frequency of occurrence. Based on the path reliability, path length penalty information, and path novelty, it detects the path evaluation result of the solution path. Based on the path evaluation result, it verifies whether the solution path is accurate.
[0138] Specifically, after determining the solution path corresponding to the problem text, the determined solution path can be verified, that is, the accuracy of the solution path can be verified. Specifically, the accuracy of the solution path can be verified from three aspects: path novelty, path reliability, and path length penalty information.
[0139] Specifically, the path novelty of the solution path is obtained by the frequency of its occurrence in the preset path library, and the path reliability of the solution path is obtained by the confidence of the relationship between multiple entity segments in the solution path. In addition, the path length penalty information of the solution path can be detected based on the path length. Finally, the path evaluation result of the solution path is detected based on the path reliability, path length penalty information and path novelty, and the accuracy of the solution path is verified by the path evaluation result.
[0140] Its expression can be: S(p) = λ1·R(p) + λ2·L(p) + λ3·N(p)
[0141] Where: R(p) represents path reliability; L(p) represents path length penalty, which uses an exponential decay function, and the longer the path length, the lower the path length penalty; N(p) represents path novelty; λ1, λ2, and λ3 are weighting coefficients, which are generally set to 0.5, 0.3, and 0.2 according to experience, but are not limited here, and can also be other values.
[0142] The specific formula for calculating path reliability R(p) is as follows: ,in, To resolve the confidence level of the i-th relation type in the path, n is the path length of the path being resolved.
[0143] In the above embodiments, by considering the novelty, reliability, and length penalty information of the solution path, the accuracy of the solution path can be accurately verified, thereby accurately determining the question response text.
[0144] In one embodiment, this application constructs a power dispatching domain knowledge graph framework embedding dynamic relationship strength for accurate text processing. Traditional power dispatching knowledge graphs mainly rely on simple triple structures, which cannot effectively express the dynamic changes and strengths of relationships between entities. This application proposes an innovative ToG (Thought-on-Graph) framework, which combines a large language model with a knowledge graph. It aims to allow the large language model to "think" along reasoning paths on the knowledge graph, achieving more efficient and reliable reasoning. Through a multi-level architectural design, it achieves dynamic calculation and updating of relationship strength. The framework first extracts entities based on a pre-trained language model and a multi-task learning strategy, then uses a multi-head interaction architecture based on graph neural networks for relationship type identification, and finally achieves dynamic fusion and updating of knowledge through a multi-path reasoning mechanism. By introducing a scoring function for relationship types and path reliability calculation, the knowledge graph can accurately reflect the changes in the strength of relationships between entities in the power system, significantly improving the effectiveness of practical applications such as fault diagnosis and operation verification.
[0145] This framework adopts a layered architecture design, which includes five core layers: data access layer, preprocessing layer, knowledge processing layer, knowledge reasoning layer, and application interface layer. Through progressive processing, it ensures the accuracy and completeness of knowledge construction.
[0146] At the data access level, the framework is designed with a complete multi-source heterogeneous data management system. The system can simultaneously access multiple data sources, such as power dispatching regulations, equipment manuals, historical fault cases, and expert knowledge bases, to obtain texts related to power dispatching.
[0147] At the preprocessing level, text in the power dispatching field is preprocessed, including basic processing such as special character processing, number standardization, unit unification, and anomaly detection.
[0148] At the knowledge processing level, the framework employs a deep learning model trained on historical power dispatching text to extract entities based on boundary information and update relationships based on neighbor nodes from new power dispatching text, resulting in a knowledge graph. The core model uses an improved BERT structure, enhancing its sequence processing capabilities by incorporating bidirectional BiLSTM layers and a multi-head attention mechanism. The model has an input dimension of 768, employs 512-dimensional hidden layers in its intermediate layers, and utilizes a bidirectional structure to improve feature extraction capabilities. The attention mechanism uses an 8-head design, with each attention head having a dimension of 128, effectively capturing long-distance dependencies in the text.
[0149] At the knowledge reasoning level, within the constructed knowledge graph, the reasoning process evaluates the reliability of different reasoning paths through a three-level path search strategy and a path evaluation function. The three-level path search strategy includes fast search based on meta-path templates, path expansion based on attention mechanisms, and dynamic pruning. The path evaluation function comprehensively considers three dimensions: path reliability, path length penalty, and path novelty, ensuring the accuracy and innovativeness of the reasoning results.
[0150] At the application interface level, the knowledge reasoning layer is connected with various application scenarios to perform path reasoning on different question texts in each application scenario and obtain the question answer text.
[0151] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0152] Based on the same inventive concept, this application also provides a knowledge graph-based text processing device for implementing the knowledge graph-based text processing method in the field of power dispatching as described above. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the knowledge graph-based text processing device in the field of power dispatching provided below can be found in the limitations of the knowledge graph-based text processing method in the field of power dispatching described above, and will not be repeated here.
[0153] In one exemplary embodiment, such as Figure 4 As shown, a knowledge graph-based text processing device for power dispatching is provided, comprising: a data acquisition module 100, an entity segmentation module 200, a relationship recognition module 300, a semantic update module 400, a knowledge graph construction module 500, and a question reasoning module 600, wherein:
[0154] Data acquisition module 100 is used to acquire text in the field of power dispatching and the corresponding question text in the field of power dispatching;
[0155] The entity segmentation module 200 is used to segment multiple entity fragments from power dispatching domain text and obtain the semantic vector features of each entity fragment.
[0156] The relationship identification module 300 is used to detect the associated entity fragments of each entity fragment, as well as the relationship type between each entity fragment and the associated entity fragments;
[0157] The semantic update module 400 is used to update the semantic vector features of each entity fragment based on the semantic vector features of the associated entity fragments corresponding to each entity fragment, as well as the relationship type between the entity fragments and the associated entity fragments.
[0158] The graph construction module 500 is used to construct a knowledge graph of text in the power dispatching domain based on the updated semantic vector features corresponding to all entity fragments and the relationship types between different entity fragments.
[0159] The question reasoning module 600 is used to analyze the question text based on the knowledge graph to obtain the question answer text.
[0160] In one embodiment, the entity segmentation module 200 is further configured to acquire multiple entity boundary sequences in power dispatching domain text, wherein the entity boundary sequences contain boundary location information and label information of multiple entities; detect the sequence conditional probability corresponding to each entity boundary sequence, and determine the target entity boundary sequence with the highest sequence conditional probability from multiple entity boundary sequences, wherein multiple label information in the target entity boundary sequence satisfies preset label constraint conditions; and detect multiple entity fragments corresponding to the target entity boundary sequence from the power dispatching domain text.
[0161] In one embodiment, the knowledge graph-based text processing device in the power dispatching field further includes an entity fragment standardization module. The entity fragment standardization module is used to search for standard text fragments similar to the entity fragment and standard associated entity fragments corresponding to the standard text fragments from a preset standard library for each entity fragment; to verify the accuracy of the entity fragment based on the standard text fragments and standard associated entity fragments; and to update the entity fragment based on the standard text fragments and standard associated entity fragments when the accuracy verification of the entity fragment fails.
[0162] In one embodiment, the semantic update module 400 is further configured to, for each entity fragment, detect at least one target associated entity fragment belonging to the same relation type from the associated entity fragments; detect the category association semantic vector features of the entity fragment based on the semantic vector features of the at least one target associated entity fragment of the same relation type; and update the semantic vector features of the entity fragment based on the category association semantic vector features corresponding to different relation types, wherein the updated semantic vector features contain both the category association semantic vector features corresponding to different relation types and the semantic vector features of the entity fragment.
[0163] In one embodiment, the knowledge graph-based text processing device in the power dispatching field further includes a relation type filtering module. This module performs preliminary filtering of the relation types between multiple entity fragments in the knowledge graph according to preset relation rule constraints. It then performs probability prediction on the correctness of each pre-filtered relation type, obtaining a rule-predicted probability for each pre-filtered relation type. Next, it performs probability prediction on the correctness of each pre-filtered relation type according to a preset classification model, obtaining a model-predicted probability for each pre-filtered relation type. Finally, it detects the confidence assessment result of each pre-filtered relation type based on the rule-predicted probability and the model-predicted probability. Based on the confidence assessment result, it performs a secondary filtering of the relation types between the pre-filtered relation types. Finally, it detects the consistency assessment result of the second-filtered relation types and performs a third-stage filtering based on the consistency assessment result to obtain an updated knowledge graph.
[0164] In one embodiment, the question reasoning module 600 is further configured to obtain the meta-path template of the question text and, in the knowledge graph, query the initial entity fragment corresponding to the question text; the filtering step is as follows: based on the meta-path template, detect multiple next candidate entity fragments associated with the initial entity fragment; based on the semantic vector features of the initial entity fragment and the semantic vector features of each next candidate entity fragment, detect the importance quantification value of each next candidate entity fragment, and determine the next candidate entity fragment with the largest importance quantification value as the next entity fragment; determine the next entity fragment as the initial entity fragment, return to the filtering step, until no next entity fragment is detected; based on the initial entity fragment, all detected next entity fragments, and the relationship type between all entity fragments, obtain the solution path corresponding to the question text, and obtain the question answer text based on the solution path.
[0165] In one embodiment, the knowledge graph-based text processing device in the power dispatching field further includes a path evaluation module. The path evaluation module is also used to detect the confidence level of the relationships between multiple entity segments in the solution path, the path length of the solution path, and the frequency of occurrence of the solution path in a preset path library; based on multiple confidence levels, it detects the path reliability of the solution path; based on the path length, it detects the path length penalty information of the solution path; and based on the frequency of occurrence, it detects the path novelty of the solution path; based on the path reliability, path length penalty information, and path novelty, it detects the path evaluation result of the solution path; and based on the path evaluation result, it verifies whether the solution path is accurate.
[0166] In the aforementioned power dispatching field, the modules in the knowledge graph-based text processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can invoke and execute the operations corresponding to each module.
[0167] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores textual data related to power dispatching. The I / O interfaces facilitate information exchange between the processor and external devices. The communication interface allows communication with external terminals via a network connection. When executed by the processor, the computer program implements a knowledge graph-based text processing method for power dispatching.
[0168] Those skilled in the art will understand that Figure 5 The structure shown is a block diagram of a partial structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0169] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0170] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0171] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0173] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0174] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0175] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A text processing method based on knowledge graphs in the field of power dispatching, characterized in that, The method includes: Obtain the power dispatching domain text and the corresponding question text; Multiple entity fragments are segmented from the power dispatch domain text, and the semantic vector features of each entity fragment are obtained; Detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and the associated entity fragment; The semantic vector features of each entity fragment are updated based on the semantic vector features of the associated entity fragment corresponding to each entity fragment, and the relationship type between the entity fragment and the associated entity fragment. Based on the updated semantic vector features corresponding to all the entity fragments and the relationship types between different entity fragments, a knowledge graph of the power dispatching domain text is constructed. Based on the knowledge graph, the question text is analyzed to obtain the question answer text.
2. The method according to claim 1, characterized in that, The process of dividing multiple entity fragments from the power dispatch domain text includes: Obtain multiple entity boundary sequences from the power dispatch domain text, wherein the entity boundary sequences contain boundary location information and label information of multiple entities; Detect the sequence condition probability corresponding to each entity boundary sequence, and determine the target entity boundary sequence with the highest sequence condition probability from the plurality of entity boundary sequences, wherein multiple label information in the target entity boundary sequence satisfies preset label constraint conditions; From the power dispatch domain text, detect multiple entity fragments corresponding to the target entity boundary sequence.
3. The method according to claim 2, characterized in that, After detecting multiple entity fragments corresponding to the target entity boundary sequence from the power dispatch domain text, the method includes: For each entity fragment, a standard text fragment similar to the entity fragment and a standard associated entity fragment corresponding to the standard text fragment are searched from a preset standard library. The accuracy of the entity fragment is verified based on the standard text fragment and the standard associated entity fragment; When the accuracy verification of the entity fragment fails, the entity fragment is updated based on the standard text fragment and the standard associated entity fragment.
4. The method according to claim 1, characterized in that, The step of updating the semantic vector features of each entity fragment based on the semantic vector features of the associated entity fragment corresponding to each entity fragment, and the relationship type with the associated entity fragment, includes: For each of the entity fragments, detect at least one target associated entity fragment belonging to the same relation type from the associated entity fragments; Based on the semantic vector features of at least one target associated entity fragment of the same relation type, detect the category-related semantic vector features of the entity fragment; Based on the category association semantic vector features corresponding to different relation types, the semantic vector features of the entity fragment are updated, wherein the updated semantic vector features contain both the category association semantic vector features corresponding to the different relation types and the semantic vector features of the entity fragment.
5. The method according to claim 1, characterized in that, Before analyzing the question text based on the knowledge graph to obtain the question answer text, the method includes: Based on preset relational rule constraints, the relational types between multiple entity fragments in the knowledge graph are initially screened, and based on preset relational rule constraints, the probability of whether each initially screened relational type is correct is predicted, thus obtaining the rule prediction probability of each initially screened relational type. Based on the preset classification model, the probability of whether each initially screened relation type is correct is predicted, and the model prediction probability of each initially screened relation type is obtained. Based on the predicted probabilities of the rules and the predicted probabilities of the model, the confidence evaluation results of each of the pre-screened relationship types are detected. Based on the confidence assessment results, a second screening is performed on the relationship types among the initially screened relationship types; The consistency assessment results of the relation types after the second screening are detected. Based on the consistency assessment results, the relation types after the second screening are screened a third time to obtain the updated knowledge graph.
6. The method according to claim 1, characterized in that, The step of analyzing the question text based on the knowledge graph to obtain the question answer text includes: Obtain the meta-path template of the question text, and in the knowledge graph, query the initial entity fragment corresponding to the question text based on the question text; Filtering step: Based on the meta-path template, detect multiple next candidate entity fragments associated with the initial entity fragment; Based on the semantic vector features of the initial entity fragment and the semantic vector features of each next candidate entity fragment, the importance quantization value of each next candidate entity fragment is detected, and the next candidate entity fragment with the largest importance quantization value is determined as the next entity fragment. The next entity fragment is identified as the initial entity fragment, and the process returns to the filtering step until no next entity fragment is detected. Based on the initial entity fragment, all detected next entity fragments, and the relationship type between all entity fragments, a solution path corresponding to the question text is obtained, and based on the solution path, a question answer text is obtained.
7. The method according to claim 6, characterized in that, Before obtaining the question answer text based on the solution path, the process includes: The confidence level of the relationships between multiple entity segments in the solution path, the path length of the solution path, and the frequency of occurrence of the solution path in a preset path library are detected. The path reliability of the solution path is detected based on multiple confidence levels, the path length penalty information of the solution path is detected based on the path length, and the path novelty of the solution path is detected based on the occurrence frequency. The path evaluation result of the solution path is detected based on the path reliability, the path length penalty information, and the path novelty. Based on the path evaluation results, verify whether the solution path is accurate.
8. A text processing device based on knowledge graphs in the field of power dispatching, characterized in that, The device includes: The data acquisition module is used to acquire power dispatching domain text and the corresponding question text; An entity segmentation module is used to segment multiple entity fragments from the power dispatch domain text and obtain the semantic vector features of each entity fragment. The relationship identification module is used to detect the associated entity fragments of each entity fragment, and the relationship type between each entity fragment and the associated entity fragment; The semantic update module is used to update the semantic vector features of each entity fragment according to the semantic vector features of the associated entity fragment corresponding to each entity fragment, and the relationship type between the entity fragment and the associated entity fragment; The graph construction module is used to construct a knowledge graph of the power dispatching domain text based on the updated semantic vector features corresponding to all the entity fragments and the relationship types between different entity fragments. The question reasoning module is used to analyze the question text based on the knowledge graph to obtain the question answer text.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Domain knowledge graph question answering method and system based on query path sorting
CN115982338A
RGCN-based operator knowledge graph completion method and apparatus, and medium
CN117763164A
Construction method, system and equipment of power dispatching knowledge question-answering system and storage medium
CN119311826A