Power document index extraction method and device, equipment and storage medium

By segmenting the power document and extracting keywords, combining the semantic expansion and knowledge injection of the big model, the problem of low efficiency in extracting indicators of power engineering documents is solved, and efficient and accurate indicator review is achieved.

CN120256604APending Publication Date: 2025-07-04HUBEI ZHENGXIN ELECTRIC POWER ENGINEERING CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478058.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the index extraction efficiency of power engineering documents is inefficient, resulting in a long review cycle and prone to omissions, affecting the speed of project progress.

Method used

By segmenting the power document, extracting keywords and building a bag of words index, using the big model for semantic expansion and knowledge injection, reasoning according to the expert thinking chain logic, and extracting index values.

Benefits of technology

It improves the efficiency and accuracy of power indicator review, reduces errors caused by human factors, and shortens the review cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256604A_ABST
    Figure CN120256604A_ABST
Patent Text Reader

Abstract

The invention discloses an index extraction method and device for a power document, equipment and a storage medium, and the method comprises the steps: obtaining the power document and a to-be-queried index, segmenting the power document to obtain a plurality of chapter document blocks, extracting keywords in the chapter document blocks, determining the word frequency, weight and appearance position of each keyword to construct a word bag index, semantic extension is carried out on the to-be-queried index, retrieval is carried out on the bag-of-word index of the power document, a plurality of target chapter document blocks are determined and input into the large model, so that the large model loads knowledge regulations and outputs index information abstracts, and all the index information abstracts are input into the large model; the large model conducts reasoning according to knowledge regulations and logic of the thinking chain, and index values are output. Therefore, through semantic extension and knowledge injection, the large model is subjected to reasoning according to the standardized logic chain, it is ensured that the large model can be subjected to reasoning according to the expert thinking mode, the index extraction accuracy is improved, and therefore the power index evaluation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent processing of technical documents in the electric power industry, and more specifically, to an indicator extraction method, device, equipment and storage medium for electric power documents. Background Art

[0002] In the field of power engineering, feasibility study reports and preliminary design reports are important bases for project review. Usually, a feasibility study and preliminary design report for a power transmission and transformation project is a 300-500 page word or pdf document, in which various indicators are scattered in different locations such as the main text and tables.

[0003] At present, when reviewing submitted materials, review experts need to rely on their own professional knowledge and experience to find indicators one by one in lengthy documents, then judge their rationality and compare them with historical experience values. This process is not only time-consuming and laborious, but also prone to omissions, resulting in a long review cycle. With the continuous expansion of the scale of power engineering construction and the increasing number of projects, the inefficiency of traditional manual review methods has become increasingly prominent, seriously affecting the speed of project advancement.

[0004] With the rapid development of artificial intelligence technology, its application in power engineering document processing has become a key way to improve review efficiency. Through the intelligent extraction of indicators by AI models, and comparison and early warning with standard empirical values ​​of similar projects, abnormal indicators can be quickly and accurately screened out, providing strong support for review experts. This can not only significantly shorten the review cycle, but also improve the accuracy and scientificity of the review and reduce errors caused by human factors.

[0005] Therefore, how to efficiently extract the index values ​​of power documents to improve the efficiency of power index review is an issue that needs attention. Summary of the invention

[0006] In view of the above problems, the present application provides a method, device, equipment and storage medium for extracting indicators from power documents, so as to efficiently extract the indicator values ​​of power documents and improve the efficiency of power indicator review.

[0007] In order to achieve the above objectives, the specific plan is proposed as follows:

[0008] A method for extracting indicators from power documents, comprising:

[0009] Acquire an electric power document and an index to be queried, and segment the electric power document according to chapters of the electric power document to obtain a plurality of chapter document blocks;

[0010] For each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword;

[0011] Combine the bag-of-words indexes of each chapter to obtain a bag-of-words index for the power document;

[0012] Semantically expand the to-be-query index to obtain to-be-query index expansion information, and retrieve the bag-of-words index of the power document based on the to-be-query index expansion information to determine several target chapter document blocks;

[0013] Input each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base. The large model is deployed with the power engineering professional knowledge base and a pre-constructed thought chain example library;

[0014] Input the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and the logic of the thought chain in the thought chain example library, and obtains and outputs the index value of the to-be-query index.

[0015] Optionally, the step of inputting each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, includes:

[0016] Input each target chapter text block into a pre-established large model, so that after the large model extracts the target keywords in the to-be-query index expansion information, it determines the knowledge regulations corresponding to the target keywords in the pre-constructed power engineering professional knowledge base, and after the large model loads the knowledge regulations, it outputs an index information summary of the target chapter text block according to the knowledge regulations.

[0017] Optionally, the step of inputting the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and the logic of the thought chain in the thought chain example library, and obtains and outputs the index value of the to-be-query index, includes:

[0018] Input the index information summaries of each target chapter text block into the large model, so that the large model determines the target thought chain matching the to-be-query index in the thought chain example library, and after the large model loads the target thought chain, it reasons according to the knowledge regulations and the logic of the target thought chain, and obtains and outputs the index value of the to-be-query index.

[0019] Optionally, the large model is also loaded with the prompt requirements input by the user;

[0020] Input each target chapter text block into a pre-established large model, so that the large model outputs a summary of the index information of the target chapter text block according to a pre-constructed power engineering professional knowledge base, including:

[0021] Input each target chapter text block into a pre-established large model, so that the large model outputs a summary of the index information of the target chapter text block according to a pre-constructed power engineering professional knowledge base and the prompt requirements.

[0022] Optionally, for each chapter document block, extract each keyword in the chapter document block, including:

[0023] For each chapter document block, delete the pause words in the chapter document block to obtain a streamlined chapter document block;

[0024] For each streamlined chapter document block, match the power engineering professional terms in the streamlined chapter document block through a pre-established power engineering professional term library to obtain keywords for each power engineering specialty.

[0025] Optionally, retrieve the power document bag-of-words index based on the to-be-query index expansion information to determine a number of target chapter document blocks, including:

[0026] Use the BM25 algorithm to retrieve the power document bag-of-words index based on the to-be-query index expansion information to determine a number of target chapter document blocks.

[0027] An index extraction device for power documents, including:

[0028] A document segmentation unit, configured to obtain a power document and a to-be-query index, and segment the power document according to the chapters of the power document to obtain a plurality of chapter document blocks;

[0029] An index construction unit, configured to, for each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword;

[0030] An index combination unit, configured to combine the bag-of-words indexes of each chapter to obtain a power document bag-of-words index;

[0031] A target chapter document determination unit for semantically expanding the to-be-query index to obtain to-be-query index expansion information, and retrieving the power document bag-of-words index based on the to-be-query index expansion information to determine a number of target chapter document blocks;

[0032] An index information summary output unit for inputting each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, and the large model is deployed with the power engineering professional knowledge base and a pre-constructed thinking chain example library;

[0033] A thinking chain reasoning unit for inputting the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain example library to obtain and output the index value of the to-be-query index.

[0034] Optionally, the index information summary output unit includes:

[0035] A knowledge regulation dynamic loading unit for inputting each target chapter text block into a pre-established large model, so that after the large model extracts the target keywords in the to-be-query index expansion information, it determines the knowledge regulations corresponding to the target keywords in the pre-constructed power engineering professional knowledge base, and after the large model loads the knowledge regulations, it outputs the index information summary of the target chapter text block according to the knowledge regulations.

[0036] Optionally, the thinking chain reasoning unit includes:

[0037] A thinking chain dynamic loading unit for inputting the index information summaries of each target chapter text block into the large model, so that the large model determines the target thinking chain matching the to-be-query index in the thinking chain example library, and after the large model loads the target thinking chain, it reasons according to the knowledge regulations and in accordance with the logic of the target thinking chain to obtain and output the index value of the to-be-query index.

[0038] Optionally, the large model also loads the prompt word requirements input by the user;

[0039] The index information summary output unit includes:

[0040] A prompt word requirement guiding unit for inputting each target chapter text block into a pre-established large model, so that the large model outputs the index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base and the prompt word requirements.

[0041] Optionally, the index construction unit includes:

[0042] A stop word deletion unit, configured to delete stop words in each chapter document block for each chapter document block, so as to obtain a refined chapter document block;

[0043] A keyword extraction unit, configured to match power engineering professional terms in each refined chapter document block through a pre-established power engineering professional term library for each refined chapter document block, obtain keywords for each power engineering specialty, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index of the chapter document block based on the word frequency, weight, and occurrence position of each keyword.

[0044] Optionally, the target chapter document determination unit includes:

[0045] A BM25 algorithm retrieval unit, configured to use the BM25 algorithm to retrieve the power document bag-of-words index based on the to-be-query index expansion information, and determine a plurality of target chapter document blocks.

[0046] An index extraction device for a power document, including a memory and a processor;

[0047] The memory is configured to store a program;

[0048] The processor is configured to execute the program to implement each step of the index extraction method for the power document as described above.

[0049] A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each step of the index extraction method for the power document as described above is implemented.

[0050] With the above technical solution, the present application obtains power documents and query indicators to be retrieved, splits the power documents into multiple chapter document blocks according to the chapters of the power documents, extracts each keyword in the chapter document blocks for each chapter document block, determines the word frequency, weight, and occurrence position of each keyword, constructs a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword, combines the bag-of-words indexes of each chapter to obtain a power document bag-of-words index, semantically expands the query indicators to be retrieved to obtain query indicator expansion information, retrieves the power document bag-of-words index based on the query indicator expansion information, determines several target chapter document blocks, inputs each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base. The large model is deployed with a power engineering professional knowledge base and a pre-constructed thinking chain example library. The index information summaries of each target chapter text block are input into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain example library, and obtains and outputs the index value of the query indicator to be retrieved. Thus, through semantic expansion and knowledge injection, and enabling the large model to reason according to a standardized logical chain, it is ensured that the large model can reason in the way of an expert's thinking, improving the accuracy of index extraction, and thus enhancing the efficiency of power index review. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Also, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0052] Figure 1 FIG. is a schematic flow chart for implementing index extraction of power documents provided by an embodiment of the present application;

[0053] Figure 2 FIG. is a schematic processing logic diagram for implementing index extraction of power documents provided by an embodiment of the present application;

[0054] Figure 3 FIG. is a schematic device structure diagram for implementing index extraction of power documents provided by an embodiment of the present application;

[0055] Figure 4 FIG. is a schematic structure diagram of a device for implementing index extraction of a power document provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0057] The solution of the present application can be implemented based on a terminal with data processing capabilities, and the terminal can be a computer, the cloud, a server, etc.

[0058] Next, in combination with Figure 1 As described above, the method for extracting indicators of the power document of the present application may include the following steps:

[0059] Step S110: Obtain a power document and the indicators to be queried, and split the power document according to the chapters of the power document to obtain a plurality of chapter document blocks.

[0060] It can be understood that splitting the large power document into multiple chapters facilitates subsequent screening of chapter document blocks that meet the requirements of the indicators to be queried, thereby avoiding inputting the large power document into the large model at one time, reducing the processing pressure on the large model, and improving the efficiency of indicator extraction.

[0061] Step S120: For each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword.

[0062] Specifically, the keyword can represent a technical term in the power engineering field, such as "transformer", "rated voltage", etc. The word frequency of the keyword can represent the number of times the keyword appears in the chapter document block. The weight of the keyword can represent the importance of the role of the keyword in the sentence where it appears (for example, the weight of the subject is 2, the weight of the object is 1, and the weight of the common noun is 1). The occurrence position of the keyword can represent the chapter information where the keyword is located. For example, if the keyword "rated voltage" appears in the third chapter, then the occurrence position of the keyword "rated voltage" can be recorded as "3". The bag-of-words index of each chapter document block can include the word frequency information, weight information, and occurrence position information of all the keywords in the chapter document block.

[0063] Step S130: Combine the bag-of-words indexes of each chapter to obtain the bag-of-words index of the power document.

[0064] Specifically, the bag-of-words index of the power document can be a higher-level index of the bag-of-words index of each chapter. The bag-of-words index of the power document can include the indexes of each chapter. When indexing to a certain chapter, the bag-of-words index of that chapter can be obtained.

[0065] Step S140: Semantically expand the query target metric to obtain the extended information of the query target metric, and retrieve the power document bag-of-words index based on the extended information of the query target metric to determine several target chapter document blocks.

[0066] Specifically, after semantically expanding the query target metric, the obtained extended information of the query target metric may contain more keywords than the query target metric. Therefore, after obtaining the extended information of the query target metric, the power document bag-of-words index can be retrieved based on the keywords included in the extended information of the query target metric, and the chapter document blocks where the retrieved keywords are located can be sorted in reverse order, so that the chapter document blocks ranked higher contain more retrieved keywords. Finally, several chapter document blocks ranked at the top are selected as the target chapter document blocks.

[0067] It can be understood that after semantically expanding the query target metric, the retrieval accuracy can be improved, making the determined target chapter document blocks more matched with the query target metric.

[0068] More specifically, the BM25 algorithm can be used to retrieve the power document bag-of-words index based on the extended information of the query target metric to determine several target chapter document blocks.

[0069] Step S150: Input each target chapter text block into a pre-established large model, so that the large model outputs the index information summary of this target chapter text block according to the pre-constructed power engineering professional knowledge base.

[0070] Among them, the large model is deployed with a power engineering professional knowledge base and a pre-constructed thought chain example library.

[0071] Specifically, the power engineering professional knowledge base can contain multiple knowledge regulations, each knowledge regulation can include several data pairs, and the form of each data pair can be <term, explanation>. When establishing the power engineering professional knowledge base, multiple metrics can be obtained, and the corresponding knowledge regulations can be constructed according to the keywords in the metric names.

[0072] The thought chain example library can contain multiple thought chains. When constructing each thought chain, the inference logic chain format corresponding to the metric can be provided to the large model, and the knowledge regulations corresponding to the metric can be attached, and each output of the large model during the training process can be sampled, so as to screen the correct examples as the thought chain corresponding to the metric. Then, the correct thought chain examples corresponding to multiple metrics can form the thought chain example library.

[0073] Among them, the logical format of the thought chain can be "Think about the user's intention -> Locate relevant text blocks -> Deeply explore information related to the text blocks -> Think step by step -> Finally confirm -> Consider the output unit -> JSON list".

[0074] Step S160: Input the index information summary of each target chapter text block into the large model, so that the large model can reason according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain sample library, and obtain and output the index value of the index to be queried.

[0075] Among them, the processing logic for extracting the indexes of the power document according to the index to be queried is as Figure 2 shown.

[0076] The index extraction method for power documents provided in this embodiment obtains a power document and an index to be queried, and cuts the power document into multiple chapter document blocks according to the chapters of the power document. For each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword. Combine the bag-of-words indexes of each chapter to obtain a power document bag-of-words index. Perform semantic expansion on the index to be queried to obtain extended information of the index to be queried, and retrieve the power document bag-of-words index based on the extended information of the index to be queried to determine several target chapter document blocks. Input each target chapter text block into a pre-established large model, so that the large model can output the index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base. The large model is deployed with a power engineering professional knowledge base and a pre-constructed thinking chain sample library. Input the index information summaries of each target chapter text block into the large model, so that the large model can reason according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain sample library, and obtain and output the index value of the index to be queried. Thus, through semantic expansion and knowledge injection, and enabling the large model to reason according to a standardized logical chain, it is ensured that the large model can reason in the way of expert thinking, improve the accuracy of index extraction, and thus improve the efficiency of power index review.

[0077] In some embodiments of the present application, the process of the above step S150: Input each target chapter text block into a pre-established large model, so that the large model can output the index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base is introduced. This process may include:

[0078] Input each target chapter text block into a pre-established large model, so that after the large model extracts the target keywords in the extended information of the index to be queried, determine the knowledge regulations corresponding to the target keywords in the pre-constructed power engineering professional knowledge base, and after the large model loads the knowledge regulations, output the index information summary of the target chapter text block according to the knowledge regulations.

[0079] It can be understood that while the power engineering professional knowledge base contains knowledge regulations corresponding to the target keywords, it also contains knowledge regulations unrelated to the target keywords. To reduce the pressure on the large model to load knowledge regulation data, only the knowledge regulations corresponding to the target keywords can be loaded. Then, when the large model responds to other query indicators, it can also extract the keywords of the extended information of other query indicators, and then determine and load the knowledge regulations corresponding to the keywords to achieve dynamic loading of knowledge regulations.

[0080] In some embodiments of the present application, the process of the above step S160, inputting the index information summaries of each target chapter text block into the large model, so that the large model performs reasoning according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain sample library, and obtains and outputs the index value of the query indicator is introduced. This process may include:

[0081] Input the index information summaries of each target chapter text block into the large model, so that the large model determines the target thinking chain matching the query indicator in the thinking chain sample library, and after the large model loads the target thinking chain, according to the knowledge regulations, and performs reasoning in accordance with the logic of the target thinking chain, to obtain and output the index value of the query indicator.

[0082] It can be understood that while the thinking chain sample library contains thinking chains matching the query indicator, it also contains thinking chains unrelated to the query indicator, and the thinking chains unrelated to the query indicator are not applicable to the logical reasoning of the query indicator. To reduce the pressure on the large model to load thinking chain sample data, only the target thinking chain corresponding to the query indicator can be loaded. Then, when the large model responds to other query indicators, it can also determine and load the thinking chain matching the other query indicators, so as to achieve dynamic loading of the thinking chain.

[0083] In some embodiments of the present application, considering that when the large model reasons according to the thinking chain corresponding to the query indicator, it also needs to meet the requirements of the user's prompt words. Specifically, the large model can respond to the instruction of the user's prompt word requirements, load the prompt word requirements input by the user, and based on this, introduce the process of the above step S150, inputting each target chapter text block into the pre-established large model, so that the large model outputs the index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base. This process may include:

[0084] Input each target chapter text block into the pre-established large model, so that the large model outputs the index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base and the prompt word requirements.

[0085] It can be understood that after the user inputs a prompt requirement, when the large model performs reasoning based on the thought chain corresponding to the index to be queried, it needs to follow this instruction of the user to further screen out the index information summary that meets the prompt requirement.

[0086] In some embodiments of the present application, the process of extracting each keyword in the chapter document block mentioned in the foregoing embodiments for each chapter document block is introduced. This process may include:

[0087] S1. For each chapter document block, delete the stop words in the chapter document block to obtain a refined chapter document block.

[0088] Specifically, examples of stop words are "of", "is", "has", etc.

[0089] For example, when the chapter document block is "In the power engineering, the substation is a key hub, relying on the transformer to regulate the voltage, using the circuit breaker to cut off the circuit in case of a fault, using the disconnecting switch to isolate the power supply, obtaining the current and voltage through the current transformer and voltage transformer, cooperating with the relay protection device to ensure safety, and then distributing the processed electric energy through the busbar to ensure stable power supply", after deleting the stop words, the refined chapter document block is " / Power engineering / , substation / key hub, relying on the transformer to regulate the voltage, using the circuit breaker / fault / cut off the circuit, using the disconnecting switch to isolate the power supply, obtaining the current / voltage through the current transformer / voltage transformer, cooperating with the relay protection device to ensure safety, and then distributing the processed / / electric energy through the busbar to ensure stable power supply"

[0090] S2. For each refined chapter document block, match the power engineering professional technical terms in the refined chapter document block through a pre-established power engineering professional technical term library to obtain the keywords of each power engineering specialty.

[0091] For example, after matching the power engineering professional technical terms in the refined chapter document block " / Power engineering / , substation / key hub, relying on the transformer to regulate the voltage, using the circuit breaker / fault / cut off the circuit, using the disconnecting switch to isolate the power supply, obtaining the current / voltage through the current transformer / voltage transformer, cooperating with the relay protection device to ensure safety, and then distributing the processed / / electric energy through the busbar to ensure stable power supply" through the power engineering professional technical term library, the keywords that can be obtained include "power engineering", "substation", "key hub", "transformer", "voltage", "circuit breaker", "circuit", "disconnecting switch", "power supply", "current transformer", "voltage transformer", "current", "relay protection device", "busbar" and "electricity".

[0092] Further, the word bag index construction result of the chapter document block exemplified in this embodiment can be: "Power Engineering": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Substation": [{"Word Frequency": 1, "Chapter": 1, "Weight": 2}], "Key Hub": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Transformer": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Voltage": [{"Word Frequency": 2, "Chapter": 1, "Weight": 1}], "Circuit Breaker": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Circuit": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Isolator Switch": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Power Supply": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Current Transformer": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Voltage Transformer": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Current": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Relay Protection Device": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Busbar": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}], "Electric Power": [{"Word Frequency": 1, "Chapter": 1, "Weight": 1}].

[0093] Next, the device for extracting indicators of power documents provided in the embodiments of the present application will be described. The device for extracting indicators of power documents described below can be correspondingly referred to the method for extracting indicators of power documents described above.

[0094] See Figure 3 , Figure 3 which is a schematic structural diagram of a device for extracting indicators of power documents disclosed in the embodiments of the present application.

[0095] As Figure 3 shown, the device may include:

[0096] A document segmentation unit 11, configured to obtain a power document and an indicator to be queried, and segment the power document according to the chapters of the power document to obtain a plurality of chapter document blocks;

[0097] An index construction unit 12, which is used to extract each keyword in the chapter document block for each chapter document block, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword;

[0098] An index combination unit 13, which is used to combine the bag-of-words indexes of each chapter to obtain a bag-of-words index for the power document;

[0099] A target chapter document determination unit 14, which is used to perform semantic expansion on the to-be-query index to obtain to-be-query index expansion information, and retrieve the bag-of-words index for the power document based on the to-be-query index expansion information to determine several target chapter document blocks;

[0100] An index information summary output unit 15, which is used to input each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, and the large model is deployed with the power engineering professional knowledge base and a pre-constructed thinking chain example library;

[0101] A thinking chain reasoning unit 16, which is used to input the index information summaries of each target chapter text block into the large model, so that the large model performs reasoning according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain example library, and obtains and outputs the index value of the to-be-query index.

[0102] Optionally, the index information summary output unit includes:

[0103] A knowledge rule dynamic loading unit, which is used to input each target chapter text block into a pre-established large model, so that after the large model extracts the target keyword in the to-be-query index expansion information, it determines the knowledge rule corresponding to the target keyword in the pre-constructed power engineering professional knowledge base, and after the large model loads the knowledge rule, it outputs an index information summary of the target chapter text block according to the knowledge rule.

[0104] Optionally, the thinking chain reasoning unit includes:

[0105] A thinking chain dynamic loading unit, which is used to input the index information summaries of each target chapter text block into the large model, so that the large model determines the target thinking chain matching the to-be-query index in the thinking chain example library, and after the large model loads the target thinking chain, it performs reasoning according to the knowledge rule and in accordance with the logic of the target thinking chain, and obtains and outputs the index value of the to-be-query index.

[0106] Optionally, the large model also loads a prompt word requirement input by the user;

[0107] The index information summary output unit includes:

[0108] A prompt word requirement guiding unit, which is used to input each target chapter text block into a pre-established large model, so that the large model outputs the index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base and the prompt word requirements.

[0109] Optionally, the index construction unit includes:

[0110] A stop word deletion unit, which is used to delete the stop words in each chapter document block for each chapter document block to obtain a refined chapter document block;

[0111] A keyword extraction unit, which is used to match the power engineering professional terms in each refined chapter document block through a pre-established power engineering professional term library for each refined chapter document block, obtain the keywords of each power engineering specialty, determine the word frequency, weight and occurrence position of each keyword, and construct the word bag index of the chapter document block based on the word frequency, weight and occurrence position of each keyword.

[0112] Optionally, the target chapter document determination unit includes:

[0113] A BM25 algorithm retrieval unit, which is used to use the BM25 algorithm to retrieve the power document word bag index based on the to-be-query index expansion information and determine several target chapter document blocks.

[0114] The device for extracting indexes of power documents provided by the embodiments of the present application can be applied to devices for extracting indexes of power documents, such as terminals: mobile phones, computers, etc. Optionally, Figure 4 shows the hardware structure block diagram of the device for extracting indexes of power documents. Refer to Figure 4 , the hardware structure of the device for extracting indexes of power documents may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0115] In the embodiments of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4;

[0116] The processor 1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;

[0117] The memory 3 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory;

[0118] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used for:

[0119] Obtain a power document and the metrics to be queried, and segment the power document according to the chapters of the power document to obtain multiple chapter document blocks;

[0120] For each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight and occurrence position of each keyword;

[0121] Combine the bag-of-words indexes of each chapter to obtain a bag-of-words index for the power document;

[0122] Semantically expand the metrics to be queried to obtain extended information on the metrics to be queried, and retrieve the bag-of-words index of the power document based on the extended information on the metrics to be queried to determine several target chapter document blocks;

[0123] Input each target chapter text block into a pre-established large model, so that the large model outputs a summary of the metric information of the target chapter text block according to a pre-constructed power engineering professional knowledge base. The large model is deployed with the power engineering professional knowledge base and a pre-constructed thinking chain example library;

[0124] Input the summary of the metric information of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain example library, and obtains and outputs the metric value of the metric to be queried.

[0125] Optionally, the refined functions and extended functions of the program can be referred to the above description.

[0126] The embodiment of the present application also provides a storage medium, which can store a program suitable for execution by a processor. The program is used for:

[0127] Obtain a power document and the metrics to be queried, and segment the power document according to the chapters of the power document to obtain multiple chapter document blocks;

[0128] For each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight and occurrence position of each keyword;

[0129] Combine the bag-of-words indexes of each chapter to obtain the bag-of-words index of the power document;

[0130] Semantically expand the to-be-query index to obtain the extended information of the to-be-query index, and retrieve the bag-of-words index of the power document based on the extended information of the to-be-query index to determine several target chapter document blocks;

[0131] Input each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base. The large model is deployed with the power engineering professional knowledge base and a pre-constructed chain-of-thought example library;

[0132] Input the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the chain of thought in the chain-of-thought example library, and obtains and outputs the index value of the to-be-query index.

[0133] Optionally, the refinement function and the extension function of the program can be referred to the above description.

[0134] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0135] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0136] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting indicators of power documents, characterized in that, Including: Obtain a power document and an index to be queried, and segment the power document according to the chapters of the power document to obtain multiple chapter document blocks; For each chapter document block, extract each keyword in the chapter document block, determine the word frequency, weight, and occurrence position of each keyword, and construct a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword; Combine the bag-of-words indexes of each chapter to obtain a power document bag-of-words index; Semantically expand the index to be queried to obtain extended information of the index to be queried, and retrieve the power document bag-of-words index based on the extended information of the index to be queried to determine several target chapter document blocks; Input each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, and the large model is deployed with the power engineering professional knowledge base and a pre-constructed thinking chain example library; Input the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain example library, and obtains and outputs the index value of the index to be queried.

2. The method according to claim 1, wherein The step of inputting each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, includes: Input each target chapter text block into a pre-established large model, so that after the large model extracts the target keyword in the extended information of the index to be queried, it determines the knowledge regulations corresponding to the target keyword in the pre-constructed power engineering professional knowledge base, and after the large model loads the knowledge regulations, it outputs an index information summary of the target chapter text block according to the knowledge regulations.

3. The method according to claim 2, wherein The step of inputting the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the thinking chain in the thinking chain example library, and obtains and outputs the index value of the index to be queried, includes: Input the index information summaries of each target chapter text block into the large model, so that the large model determines the target thinking chain matching the index to be queried in the thinking chain example library, and after the large model loads the target thinking chain, it reasons according to the knowledge regulations and in accordance with the logic of the target thinking chain, and obtains and outputs the index value of the index to be queried.

4. The method according to claim 1, wherein The large model is also loaded with prompt requirements input by the user; The step of inputting each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, includes: Input each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to the pre-constructed power engineering professional knowledge base and the prompt requirements.

5. The method according to claim 1, characterized in that For each chapter document block, extract each keyword in the chapter document block, including: For each chapter document block, delete the stop words in the chapter document block to obtain a refined chapter document block; For each refined chapter document block, match the power engineering professional terms in the refined chapter document block through a pre-established power engineering professional term library to obtain keywords for each power engineering specialty.

6. The method according to any one of claims 1-5, characterized in that, The retrieval of the power document bag-of-words index based on the to-be-query index expansion information to determine several target chapter document blocks includes: Use the BM25 algorithm to retrieve the power document bag-of-words index based on the to-be-query index expansion information to determine several target chapter document blocks.

7. An index extraction device for power documents, characterized in that, Including: A document segmentation unit for obtaining a power document and a to-be-query index, and segmenting the power document according to the chapters of the power document to obtain a plurality of chapter document blocks; An index construction unit for, for each chapter document block, extracting each keyword in the chapter document block, determining the word frequency, weight, and occurrence position of each keyword, and constructing a bag-of-words index for the chapter document block based on the word frequency, weight, and occurrence position of each keyword; An index combination unit for combining the bag-of-words indexes of each chapter to obtain a power document bag-of-words index; A target chapter document determination unit for semantically expanding the to-be-query index to obtain to-be-query index expansion information, and retrieving the power document bag-of-words index based on the to-be-query index expansion information to determine several target chapter document blocks; An index information summary output unit for inputting each target chapter text block into a pre-established large model, so that the large model outputs an index information summary of the target chapter text block according to a pre-constructed power engineering professional knowledge base, and the large model is deployed with the power engineering professional knowledge base and a pre-constructed thought chain example library; A thought chain reasoning unit for inputting the index information summaries of each target chapter text block into the large model, so that the large model reasons according to the power engineering professional knowledge base and in accordance with the logic of the thought chain in the thought chain example library to obtain and output the index value of the to-be-query index.

8. The device according to claim 7, characterized in that, The index information summary output unit includes: A knowledge rule dynamic loading unit for inputting each target chapter text block into a pre-established large model, so that after the large model extracts the target keyword in the to-be-query index expansion information, it determines the knowledge rule corresponding to the target keyword in the pre-constructed power engineering professional knowledge base, and after the large model loads the knowledge rule, outputs an index information summary of the target chapter text block according to the knowledge rule.

9. An index extraction device for power documents, characterized in that, Including a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the index extraction method for power documents as described in any one of claims 1-6.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the index extraction method for power documents as described in any one of claims 1-6.