Knowledge Graph Reasoning Method and Device Based on Text Mapping and Vector Library Retrieval
By constructing text mapping models and vector library retrieval methods, the problems of high storage cost of power knowledge graphs and inefficient inference are solved, and fast and accurate fault knowledge graph inference under limited resource conditions are achieved, and the efficiency and quality of power maintenance are improved.
Patent Information
- Application Number
- CN202510629037.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing power knowledge graph inference storage is expensive and inefficient in reasoning, resulting in reduced power maintenance efficiency and quality.
A text mapping model is constructed based on neural networks, a BGE model is used to search text embedding and vector library, and a vector database is constructed. Through the text embedding and splicing of fault phenomena and query relationships, the answer text is directly queried in the vector database.
It reduces storage requirements, improves inference efficiency and accuracy, and reduces hardware costs. It is suitable for rapid and accurate power failure knowledge graph inference in resource-constrained environments.
Smart Images

Figure CN120181243B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graphs, and in particular to a knowledge graph reasoning method and device based on text mapping and vector library retrieval. Background Art
[0002] Existing power knowledge graphs are derived from historical data tables exported from power system operation and maintenance systems. First, the ontology structure of the knowledge graph is designed, and then the table items therein are used as entity, attribute and other information to instantiate the knowledge graph. However, due to the huge amount of data in the historical data tables of the power system and the large number of missing table items, for example, the relevant information on the fault classification and fault location is missing for a certain fault, which will bring problems of semantic loss and high storage cost. Moreover, as the data in the power operation and maintenance system increases, the power knowledge graph constructed based on the historical data tables of the power operation and maintenance system will also require more and more hardware resources to meet the needs of reasoning. Therefore, after establishing the power knowledge graph, when trying to query information such as its corresponding technical reasons, responsibility reasons, maintenance opinions, fault classifications, and classification bases in the power knowledge graph according to the fault phenomenon, if no effective index is established for the power knowledge graph for reasoning, two problems will occur: firstly, a large amount of hardware resources are required to store the power knowledge graph, making it difficult to apply the reasoning of the power knowledge graph in actual engineering scenarios; secondly, during reasoning, it is necessary to traverse the existing power knowledge graph database with a huge storage capacity one by one, which will take a lot of time and resources, resulting in decision-makers being unable to timely master comprehensive information, affecting the efficiency and quality of power maintenance, and further affecting the reliability of power supply. Summary of the Invention
[0003] Therefore, the technical problem to be solved by the present invention is to overcome the problems of high storage cost and low reasoning efficiency during the reasoning of the power knowledge graph in the prior art, which lead to the reduction of the efficiency and quality of power maintenance.
[0004] To solve the above technical problem, the present invention provides a knowledge graph reasoning method based on text mapping and vector library retrieval, including:
[0005] Constructing a text mapping model based on a neural network, using the concatenation of the text embedding of the fault phenomenon and the text embedding of the query relationship in the power fault knowledge graph as the input, and using the vector of the answer text corresponding to the fault phenomenon and the query relationship as the label to train the text mapping model;
[0006] Obtaining the answer text in the power fault knowledge graph, and constructing a vector database using the answer text and its corresponding vector;
[0007] After the user inputs the query text, extracting the key information of the query text, where the key information includes the fault phenomenon text and the query relationship text;
[0008] Obtain the text embeddings of the fault phenomenon and the query relationship according to the fault phenomenon text and the query relationship text. After splicing, input them into the trained text mapping model to obtain the target query vector; use the target query vector to query in the vector database and output the target answer text.
[0009] Preferably, the query relationship includes technical reasons, responsible reasons, maintenance opinions, fault classifications, and classification bases.
[0010] Preferably, the text mapping model includes a first linear layer, a first relu activation function, a second linear layer, a second relu activation function, and a third linear layer connected in sequence.
[0011] Preferably, the loss function of the text mapping model is:
[0012] ;
[0013] where is the loss function, represents the total number of samples, represents the i-th output vector of the model, represents the vector of the i-th positive sample, represents the vector of the negative sample corresponding to the i-th positive sample, represents the vector of the negative sample, and margin is a fixed value.
[0014] Preferably, use the BGE model to obtain the text embeddings of the fault phenomenon, the text embeddings of the query relationship, and the vectors of the answer text.
[0015] Preferably, use the answer text and its corresponding vector to construct a vector database using chromaDB, including:
[0016] Use the BGE model to obtain the vector of the answer text, add the answer text and its corresponding vector to the set respectively, and use chromaDB to construct the corresponding persistent vector database so as to use the vector of the answer text as an index to query its corresponding answer text.
[0017] Preferably, when there are multiple query relationship texts of key information, obtain the text embeddings of each query relationship, splice them with the text embeddings of the fault phenomenon respectively, and input them into the trained text mapping model to obtain the target query vectors corresponding to each query relationship.
[0018] Preferably, when constructing a persistent vector database using the answer text and its corresponding vector, classify the answer text according to the query relationship, and use metadata to label the category to which the answer text belongs to obtain the answer set corresponding to each query relationship; when using the target query vector to query in the vector database, first obtain the answer set corresponding to the query relationship, and then obtain the answer text corresponding to the query relationship in the answer set.
[0019] Preferably, an offline Qwen7B language model is used to extract the key information of the query text.
[0020] The present invention also provides a knowledge graph reasoning device based on text mapping and vector library retrieval, including:
[0021] A mapping model training module, configured to construct a text mapping model based on a neural network, using the text embedding of the fault phenomenon and the text embedding of the query relationship in the power fault knowledge graph after splicing as input, and using the vector of the answer text corresponding to the fault phenomenon and the query relationship as a label to train the text mapping model;
[0022] A vector database construction module, configured to obtain the answer text in the power fault knowledge graph and construct a vector database using the answer text and its corresponding vector;
[0023] A key information extraction module, configured to extract the key information of the query text after the user inputs the query text, where the key information includes the fault phenomenon text and the query relationship text;
[0024] An answer text output module, configured to obtain the text embedding of the fault phenomenon and the text embedding of the query relationship according to the fault phenomenon text and the query relationship text, splice them and input them into the trained text mapping model to obtain a target query vector; use the target query vector to query in the vector database and output the target answer text.
[0025] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0026] A knowledge graph reasoning method based on text mapping and vector library retrieval according to the present invention, in view of the characteristics of a large number of long-tail entities and clear reasoning purposes in the power knowledge graph, constructs a text mapping model. The input is the concatenation of the text embeddings of the fault phenomena and the text embeddings of the query relationships in the power fault knowledge graph, and a target query vector is obtained. A vector database is constructed using the answer texts in the power fault knowledge graph and their corresponding vectors. The target query vector is directly used to query the target answer text in the vector database. This not only fully considers the text information in the knowledge graph, but also saves the storage of a large number of long-tail entities or attributes and only stores a small amount of answer text data. While reducing the storage pressure of the system, it improves the reasoning efficiency of the power fault knowledge graph, enables fast and accurate reasoning under limited hardware resource conditions, improves the quality and efficiency of power maintenance, and promotes the application of the power fault knowledge graph in the actual production environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the accompanying drawings, where:
[0028] Figure 1 is a flowchart of a knowledge graph reasoning method based on text mapping and vector library retrieval according to the present invention;
[0029] Figure 2 is a structure diagram of a power fault knowledge graph;
[0030] Figure 3 is a structure diagram of a text mapping model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following further illustrates the present invention in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention. Embodiment 1
[0032] Referring to Figure 1 as shown, the present invention provides a knowledge graph reasoning method based on text mapping and vector library retrieval, including:
[0033] S1: Construct a power fault knowledge graph based on historical fault information;
[0034] S2: Construct a text mapping model based on a neural network. The input is the concatenation of the text embeddings of the fault phenomena and the text embeddings of the query relationships in the power fault knowledge graph, and the vectors of the answer texts corresponding to the fault phenomena and the query relationships are used as labels to train the text mapping model;
[0035] S3: Obtain the answer text in the power failure knowledge graph, and construct a vector database using the answer text and its corresponding vector;
[0036] S4: After the user inputs the query text, extract the key information of the query text, where the key information includes the fault phenomenon text and the query relationship text;
[0037] S5: Obtain the text embeddings of the fault phenomenon and the query relationship, splice them and input them into the trained text mapping model to obtain the target query vector; use the target query vector to query in the vector database and output the target answer text.
[0038] In S1, first obtain the historical fault information of the power system operation and maintenance system, extract the corresponding entities and attribute descriptions from the historical fault information according to the structure of the knowledge graph, establish the corresponding relationship between texts, and construct the power failure knowledge graph. The power failure knowledge graph constructed by the present invention focuses on establishing the mapping relationship between information such as fault phenomena and their corresponding technical reasons, responsible reasons, maintenance opinions, fault classifications, and classification bases. Therefore, when constructing the power failure knowledge graph, first determine the entities in the power failure knowledge graph according to the historical fault information of the power system operation and maintenance system, such as fault phenomena, technical reasons, responsible reasons, etc., and define the relationships between entities, such as "causing", "belonging to", "corresponding", etc. Then, use a knowledge graph construction tool or programming language to convert the labeled data into the nodes and edges of the knowledge graph, forming a visual knowledge graph structure. The structure of the power failure knowledge graph constructed by the present invention refers to Figure 2 as shown.
[0039] Based on the constructed power failure knowledge graph, obtain information such as each fault phenomenon and its corresponding technical reasons, responsible reasons, maintenance opinions, fault classifications, and classification bases, as the original data set for the subsequent training of the text mapping model, and realize the mapping matching of key information on the knowledge base and the reasoning of the knowledge graph quickly. Among them, all child nodes under technical reasons, responsible reasons, maintenance opinions, fault classifications, and classification bases are used as the answer text.
[0040] Text embedding technology can convert text into a vector representation in a high-dimensional space, so as to better capture and retain the rich semantic information of the text. In the field of knowledge graph embedding, TransE and TransR are two representative methods. TransE realizes embedding by calculating the distance between entities and relationships in the vector space. Its advantage is that the model structure is simple and easy to implement, but there may be certain limitations when dealing with complex relationships. TransR, on the other hand, innovatively embeds entities and relationships into different vector spaces separately, and equips each relationship with a mapping matrix to map entities into the relationship-specific space, so that the complex interaction between entities and relationships can be characterized more accurately.
[0041] Another text embedding technique, the Language Models as Knowledge Embeddings (LMKE) method, utilizes pre-trained language models to capture the deep semantic features of text and generate text embeddings with rich semantic information. This method is particularly suitable for handling long-tail entities, can represent these entities well using text information, and effectively solves the problems of existing text-based knowledge embedding methods in terms of efficiency and performance.
[0042] However, despite the significant progress made by existing text embedding techniques and knowledge graph embedding methods in their respective application scenarios, they still have certain limitations when dealing with text mapping and retrieval tasks. Deep learning-based text embedding techniques perform well in processing unstructured text, being able to deeply understand the semantic and context relationships of text, but may be less effective than some graph structure or relation mapping-based methods when dealing with structured data. Knowledge graph embedding methods such as TransE and TransR, although having advantages in dealing with structured data and being able to capture the structural relationships between entities and relations well, may face challenges when dealing with unstructured text and are difficult to directly apply to text semantic understanding and retrieval tasks.
[0043] Therefore, based on the characteristics of the power failure knowledge graph, this embodiment selects the BGE (BAAI General Embedding) model as the text embedding model and establishes a mapping relationship between the embedding of the query text and the vector of the answer text through a text mapping model.
[0044] The BGE model learns the semantic representation of text through two stages: pre-training and fine-tuning. It adopts a variety of advanced techniques and algorithms in aspects such as multilingual support, long text processing, and embedding optimization. In the pre-training stage, the BGE model uses a large-scale multilingual corpus to learn the common semantic features between different languages, and in the fine-tuning stage, it further optimizes for specific tasks to improve the model's performance on that task. In addition, the BGE model also generates a more reasonable similarity distribution through techniques such as instruction fine-tuning and hard negative sample mining, thus significantly improving the retrieval performance.
[0045] In S2, a text mapping model is constructed to map the relevant text information obtained from the power failure knowledge graph in S1. The goal of model training is to map the description text of the fault phenomenon and the description text of the query relationship together as much as possible, so that the output vector of the model is as close as possible to the vector of the answer text corresponding to the fault phenomenon and the query relationship. For example, the fault phenomenon is: "The busbar support insulator of cylinder No. 36 beside the B-phase of capacitor 343 in the 500kV***35kV #6 capacitor bank is broken", and the query relationship is: "Responsible cause", then the target of mapping is the text embedding of "Poor material quality". The query relationships include information such as technical reasons, responsible causes, maintenance opinions, fault classifications, and classification bases.
[0046] S2 specifically includes the following steps:
[0047] S21: Construct a text mapping model and determine the input and output.
[0048] In this embodiment, the BGE model is used to obtain the text embedding of the fault phenomenon, the text embedding of the query relationship, and the vector of the answer text. The description text of the fault-related information in the power failure knowledge graph, such as entity names, attribute names, and relationship names, etc., is processed by the language embedding model BGE to obtain its corresponding embedding.
[0049] The embedding of the fault phenomenon and the text embedding of the query relationship are concatenated as the model input, and then several fully connected feedforward neural networks are used for several rounds of linear transformation. Finally, a vector with the same size as the text embedding is obtained as the result of the mapping. The label of the model is the vector of the answer text corresponding to the input fault phenomenon and query relationship. The structure of the text mapping model refers to Figure 3 As shown, the text mapping model includes a first linear layer, a first relu activation function, a second linear layer, a second relu activation function, and a third linear layer connected in sequence. Where hidden_size is the size of the embedding vector, which is 1024 when using the BGE model as the text embedding model.
[0050] S22: Determine the loss function and train the text mapping model.
[0051] In this embodiment, the method of contrastive learning is used for training. And since in this embodiment, the vector library uses chromaDB as a lightweight vector database, where the L2 loss is used to calculate the embedding similarity. In order to be consistent with the actual application scenario, a variant of Triplet Loss is also used as the loss function in training, and its formula is as follows:
[0052] ;
[0053] Where, is the loss function, represents the total number of samples, represents the i-th output vector of the model, represents the vector of the i-th positive sample, represents the vector of the negative sample corresponding to the i-th positive sample, represents the vector of the negative sample, || represents calculating its L2 distance, margin represents a fixed value used to control the repulsion degree of the negative sample and suppress overfitting and overly stretching the distance of the negative sample; a positive sample may correspond to different numbers of negative samples, so the distances of the negative samples are averaged.
[0054] S23: Evaluation of the model training results.
[0055] In this embodiment, the HIT@k (Hit Rate at k) and MRR (Mean Reciprocal Rank) metrics are used to evaluate the mapping results. They are common metrics in the recommendation system, where HIT@k indicates whether the correct answer is included in the top k results recommended by the model. And MRR is a metric based on the ranking position, used to measure the position of the correct answer in the recommendation list, and the higher the position, the higher the score. For each user or query, find the ranking r of the correct answer and calculate its reciprocal (Reciprocal Rank, 1 / r). If there is no correct answer, the score is 0. The results of the model training in this embodiment are shown in Table 1.
[0056] Table 1. Evaluation Results of Text Mapping Model Training
[0057] Index Score HIT@1 59.2 HIT@3 73.1 HIT@10 80.6 MRR 0.459
[0058] In step S3, based on the structure of the power failure knowledge graph and the answer domain related to the question, use the BGE model to obtain the vector of the answer text, add the answer text and its corresponding vector to the set respectively, and use chromaDB to build the corresponding persistent vector database, so as to use the vector of the answer text as an index to query its corresponding answer text. The data stored in the database is used as the retrieval data source in the inference stage.
[0059] Preferably, in the vector database, in this embodiment, the answer texts are classified according to the query relationship, and the metadata is used to label the specific answer domain, that is, the category to which the answer text belongs, for determining the candidate answer set according to the target domain corresponding to the query relationship. Since the number of answer texts is significantly smaller than the number of questions (fault phenomena), the storage space required for storing data is also smaller.
[0060] In step S4, extract the key information of the query text, which specifically includes the following steps:
[0061] S41: User question classification, which is used to determine whether the query text of the user belongs to the scope solved by the method of the present invention, that is, given a fault phenomenon, ask for information such as its corresponding technical reason, responsibility reason, maintenance opinion, fault classification, and classification basis.
[0062] In this embodiment, the offline Qwen7B language model is used to extract the key information of the query text. If the query text of the user belongs to the scope solved by the method of the present invention, then enter the next step; if not, other models need to be used for processing.
[0063] Exemplarily, the user gives the query text of the following two questions:
[0064] Question 1: "I found that the connection between the neutral line and the fuse of phase C of the 319 capacitor bank of the 35kV #1 capacitor in the 220kV ** substation is heated to 54.4 degrees, the normal phase is 32.2 degrees, the ambient temperature is 22 degrees, the relative temperature difference is 68%, and the load current is 240. How serious is this fault? What could be the responsibility reasons for this fault?";
[0065] Question 2: "What are the top three responsibility reasons for serious and above faults in oil-immersed transformers and how many times have they occurred in total?"
[0066] Among them, Question 1 is within the scope solved by the method of the present invention and should enter the next step, while Question 2 involves statistical information and does not belong to the scope solved by the method of the present invention, so other methods need to be called to solve it.
[0067] S42: Extract the key information of the query text, including the fault phenomenon text and the query relationship text. The query relationship can be one or more.
[0068] Taking the above Question 1 as an example, the extracted fault phenomenon text is "The connection between the neutral line and the fuse of phase C of the 319 capacitor bank of the 35kV #1 capacitor in the 220kV ** substation is heated to 54.4 degrees, the normal phase is 32.2 degrees, the ambient temperature is 22 degrees, the relative temperature difference is 68%, and the load current is 240", and the extracted query relationship texts are "fault classification" and "responsibility reason".
[0069] In step S5, the extracted fault phenomenon text and query relationship text are mapped to obtain the target answer text, which specifically includes the following steps:
[0070] S51: Use the BGE model to obtain the text embedding of the fault phenomenon and the text embedding of the query relationship.
[0071] Embed the fault phenomenon text and the query relationship text using the same BGE model as in the previous steps to obtain the text embedding of the fault phenomenon and the text embedding of the query relationship.
[0072] When there are multiple query relation texts for key information, obtain the text embeddings of each query relation. After concatenating the text embeddings of each query relation with the text embedding of the fault phenomenon respectively, input them into the trained text mapping model to obtain the target query vectors corresponding to each query relation text.
[0073] Taking the above problem 1 as an example, in this scenario, the text mapping model needs to be called twice. One is to concatenate the text embedding of the fault phenomenon and the text embedding of the query relation "fault classification" and then call the model to obtain the target query vector of "fault classification"; the other is to concatenate the text embedding of the fault phenomenon and the text embedding of the query relation "responsibility cause" and then call the model to obtain the target query vector of "responsibility cause".
[0074] S52: Use the target query vector to query in the persistent vector database and output the target answer text.
[0075] Use the obtained target query vector to query the vector database, and use the query relation text to restrict the candidate answer texts. First, obtain the answer set corresponding to the query relation text, and then query the text corresponding to the vector most similar to the target query vector in the answer set as the target answer text.
[0076] Taking the above problem 1 as an example, taking the fault phenomenon and "fault classification" as the input, restricting the answer set in the vector database to the answer texts corresponding to "fault classification", and the text content of the vector closest to the target query vector output in this answer set is "general", so the fault level corresponding to this fault phenomenon is "general". Similarly, it is queried that the responsibility cause corresponding to this fault phenomenon is "component aging". These target answer texts, as the results of knowledge graph reasoning, can be used as answers or information relied on in subsequent other steps.
[0077] A knowledge graph reasoning method based on text mapping and vector library retrieval according to the present invention has the following beneficial effects:
[0078] 1. The present invention provides a fast reasoning solution based on text information: The present invention fully considers the important text information in the power field, constructs a text mapping model based on text embeddings, and avoids the huge overhead caused by fine-tuning the entire text embedding model by directly mapping the embeddings. The text mapping model has a small number of parameters and a simple structure, so the training cost is low, the reasoning speed is faster, and the matching accuracy is also maintained at a high level, improving the efficiency of knowledge graph reasoning and being conducive to real-time deployment in actual application scenarios.
[0079] 2. The hardware cost requirement of the present invention is low, and the data storage volume is small: the text mapping model used in the present invention has a small number of parameters and requires few hardware resources, which is suitable for deployment in resource-constrained environments. At the same time, the constructed vector database does not need to store a large amount of fault phenomenon information, but only needs to store the corresponding reasoning target description, that is, the answer text involving technical reasons, responsibility reasons, maintenance opinions, fault classification and classification basis, etc. This part of information usually has a certain degree of versatility, so the data volume is much smaller than the fault phenomenon information, which effectively reduces storage resources and reduces storage costs.
[0080] 3. The present invention is closer to the application scenario: The training goal of the text mapping model for text embedding mapping is to make the mapped text as close as possible to the matching text vector data, so the output of the text mapping model can be used to directly search the vector database without the need for additional steps. In addition, the persistent vector library search scheme and the loss function of the text mapping model training also match, so the training scenario and the application scenario are not much different, which facilitates the direct application of the text mapping model in actual scenarios. At the same time, the method of the present invention is scalable, and the various vectors used are completely derived from the text description, which makes it easy to expand the system. Even if there is a newly added target text, it can be directly added to the vector database without the need to retrain the entire model like the traditional knowledge graph method.
[0081] In summary, the knowledge graph reasoning method based on text mapping and vector library retrieval described in the present invention, according to the characteristics of the power knowledge graph with many long-tail entities and clear reasoning purpose, constructs a text mapping model, takes the text embedding of the fault phenomenon in the power fault knowledge graph and the text embedding of the query relationship as input, obtains the target query vector, and uses the answer text in the power fault knowledge graph and its corresponding vector to construct a vector database, and uses the target query vector to directly query the target answer text in the vector database. It not only fully considers the text information in the knowledge graph, but also saves the storage of a large number of long-tail entities or attributes and only needs to store a small amount of answer text data. While reducing the system storage pressure, it improves the efficiency of power fault knowledge graph reasoning, and can achieve fast and accurate reasoning under limited hardware resources, which promotes the application of power fault knowledge graph in actual production environment. Embodiment 2
[0082] Based on the knowledge graph reasoning method based on text mapping and vector library retrieval described in Example 1, this embodiment provides a knowledge graph reasoning device based on text mapping and vector library retrieval, including:
[0083] A mapping model training module, which is used to construct a text mapping model based on a neural network, take the concatenation of the text embedding of the fault phenomenon and the text embedding of the query relationship in the power fault knowledge graph as the input, and take the vector of the answer text corresponding to the fault phenomenon and the query relationship as the label to train the text mapping model;
[0084] A vector database construction module, which is used to obtain the answer text in the power fault knowledge graph and construct a vector database by using the answer text and its corresponding vector;
[0085] A key information extraction module, which is used to extract the key information of the query text after the user inputs the query text, and the key information includes the fault phenomenon text and the query relationship text;
[0086] An answer text output module, which is used to obtain the text embedding of the fault phenomenon and the text embedding of the query relationship according to the fault phenomenon text and the query relationship text, concatenate them and input them into the trained text mapping model to obtain a target query vector; use the target query vector to query in the vector database and output the target answer text.
[0087] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or blocks Figure 1 a device for the function specified in one block or multiple blocks.
[0089] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements in the process Figure 1 a process or multiple processes and / or blocksFigure 1 The functions specified in one or more boxes.
[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one or more processes and / or boxes Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0091] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A knowledge graph reasoning method based on text mapping and vector library retrieval, characterized in that, Including: Construct a text mapping model based on a neural network, using the concatenation of the text embeddings of the fault phenomena and the text embeddings of the query relationships in the power fault knowledge graph as the input, and the vectors of the answer texts corresponding to the fault phenomena and the query relationships as the labels to train the text mapping model; the query relationships include technical reasons, responsibility reasons, maintenance opinions, fault classifications, and classification bases; the text mapping model includes a first linear layer, a first relu activation function, a second linear layer, a second relu activation function, and a third linear layer connected in sequence; Obtain the answer texts in the power fault knowledge graph, and use the answer texts and their corresponding vectors to construct a vector database; After the user inputs a query text, extract the key information of the query text, where the key information includes the fault phenomenon text and the query relationship text; Obtain the text embedding of the fault phenomenon and the text embedding of the query relationship according to the fault phenomenon text and the query relationship text, concatenate them and input them into the trained text mapping model to obtain the target query vector; use the target query vector to query in the vector database and output the target answer text.
2. The knowledge graph reasoning method based on text mapping and vector library retrieval according to claim 1, wherein The loss function of the text mapping model is: ; Among them, is the loss function, represents the total number of samples, represents the i-th output vector of the model, represents the vector of the i-th positive sample, represents the vector of the negative sample corresponding to the i-th positive sample, represents the vector of the negative sample, and margin is a fixed value.
3. A knowledge graph reasoning method based on text mapping and vector library retrieval according to claim 1, characterized in that Use the BGE model to obtain the text embedding of the fault phenomenon, the text embedding of the query relationship, and the vector of the answer text.
4. A knowledge graph reasoning method based on text mapping and vector library retrieval according to claim 1, characterized in that, Use the answer text and its corresponding vector to construct a vector database using chromaDB, including: Use the BGE model to obtain the vector of the answer text, add the answer text and its corresponding vector to the set respectively, and use chromaDB to construct the corresponding persistent vector database, so as to use the vector of the answer text as an index to query its corresponding answer text.
5. A knowledge graph reasoning method based on text mapping and vector library retrieval according to claim 1, characterized in that When there are multiple query relationship texts in the key information, obtain the text embedding of each query relationship, concatenate them with the text embedding of the fault phenomenon respectively, and input them into the trained text mapping model to obtain the target query vector corresponding to each query relationship.
6. A knowledge graph reasoning method based on text mapping and vector library retrieval according to claim 1, characterized in that, When constructing a persistent vector database using the answer text and its corresponding vector, classify the answer text according to the query relationship, and use metadata to label the category to which the answer text belongs to obtain the answer set corresponding to each query relationship; when using the target query vector to query in the vector database, first obtain the answer set corresponding to the query relationship, and then obtain the answer text corresponding to the query relationship in the answer set.
7. A knowledge graph reasoning method based on text mapping and vector library retrieval according to claim 1, characterized in that Use the offline Qwen7B language model to extract the key information of the query text.
8. A knowledge graph reasoning device based on text mapping and vector library retrieval, characterized in that Including: A mapping model training module, which is used to construct a text mapping model based on a neural network, using the concatenation of the text embeddings of the fault phenomena and the text embeddings of the query relationships in the power fault knowledge graph as the input, and the vectors of the answer texts corresponding to the fault phenomena and the query relationships as the labels to train the text mapping model; the query relationships include technical reasons, responsibility reasons, maintenance opinions, fault classifications, and classification bases; the text mapping model includes a first linear layer, a first relu activation function, a second linear layer, a second relu activation function, and a third linear layer connected in sequence; A vector database construction module, which is used to obtain the answer text in the power failure knowledge graph and construct a vector database by using the answer text and its corresponding vectors; A key information extraction module, which is used to extract the key information of the query text after the user inputs the query text, and the key information includes the fault phenomenon text and the query relationship text; An answer text output module, which is used to obtain the text embedding of the fault phenomenon and the text embedding of the query relationship according to the fault phenomenon text and the query relationship text, splice them and input them into the trained text mapping model to obtain the target query vector; use the target query vector to query in the vector database and output the target answer text.
Citation Information
Patent Citations
Metallurgical knowledge question-answering method and system based on knowledge graph
CN114238595A
Common question answering method and system, device and medium
WO2023246093A1