A method, system, device and storage medium for archiving knowledge compilation and research

By constructing a quadruple knowledge graph and using prompt engineering technology to rewrite the editing and research needs, the problem of insufficient knowledge coverage in the proprietary fields of large language models is solved, and more accurate and credible archival knowledge editing and research results are achieved.

CN119623604BActive Publication Date: 2025-06-27STATE GRID TIANJIN ELECTRIC POWER COMPANY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510148562.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-27
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The insufficient knowledge coverage of large language models in proprietary fields leads to insufficient expression of knowledge characteristics and insufficient knowledge reasoning ability, and prone to knowledge hallucinations.

Method used

By constructing a quadruple-based archive knowledge graph, accurately locate archive data that matches the editing and research needs, use prompt engineering technology to rewrite the editing and research requirements information, generate the editing and research description, determine the target quadruple and associate additional features, and finally input the relevant input information into the large language model to generate the editing and research results.

Benefits of technology

It improves the accuracy of input instructions of the large language model, overcomes the illusion of the large language model, improves the credibility and text quality of the target editing content of the output, and enhances the depth and breadth of the editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623604B_ABST
    Figure CN119623604B_ABST
Patent Text Reader

Abstract

The present application proposes a method, system, device and storage medium for archival knowledge compilation and research. The method includes: obtaining compilation and research requirement information and an archival knowledge graph, using prompt engineering technology to rewrite the compilation and research requirement information to obtain a compilation and research description; in the archival knowledge graph, determining and recalling target quadruples that match the compilation and research description; determining additional features according to the requirement type associated with the compilation and research description, and associating the additional features with the target quadruples; using the target quadruples, the source original text associated with the target quadruples, the additional features associated with the target quadruples, and the compilation and research description as input information for a large language model, and outputting an archival knowledge compilation and research result through the large language model. This technical solution improves the accuracy level of the input instructions of the large language model, helps to overcome the hallucination of the large language model, and improves the credibility and text quality of the target compilation and research content output by the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of archival data processing, and particularly to a method, system, device, and storage medium for archival knowledge compilation and research. Background Art

[0002] Archival knowledge is characterized by being unstructured and fragmented. When conducting archival knowledge compilation and research, it is necessary to accurately extract core knowledge elements from a large number of archival texts, images, and other resources, and reveal the semantic associations between the elements. Currently, relevant technical personnel directly input the archival knowledge analysis requirements into a large language model, and the large language model automatically outputs the corresponding archival knowledge analysis results. Although the large language model has advantages in semantic understanding, generation ability, etc., its knowledge coverage in specialized fields is insufficient, facing problems such as insufficient expression of knowledge features and insufficient knowledge reasoning ability, thus prone to the phenomenon of knowledge hallucination, that is, the phenomenon that the generated content does not match the actual situation. Therefore, how to improve the accuracy of archival knowledge compilation and research results has become an urgent technical problem to be solved currently. Summary of the Invention

[0003] Embodiments of this application provide a method, system, device, and storage medium for archival knowledge compilation and research, aiming to solve the technical problems in the related art that the large language model has insufficient knowledge coverage in specialized fields, faces problems such as insufficient expression of knowledge features and insufficient knowledge reasoning ability, and thus is prone to the phenomenon of knowledge hallucination, that is, the phenomenon that the generated content does not match the actual situation.

[0004] In a first aspect, embodiments of this application provide a method for archival knowledge compilation and research, which is applicable to a power system and includes:

[0005] Obtain compilation and research requirement information and an archival knowledge graph, and use prompt engineering technology to rewrite the compilation and research requirement information to obtain a compilation and research description. The archival knowledge graph is constructed based on quadruples extracted from the text to be extracted, and the text to be extracted is archival materials obtained in advance;

[0006] In the archival knowledge graph, determine and recall target quadruples that match the compilation and research description;

[0007] According to the requirement type associated with the compilation and research description, determine additional features, and associate the additional features with the target quadruples. The additional features are used to determine the proportion of the description text of the relevant text content corresponding to the target quadruples in the archival knowledge compilation and research results;

[0008] Use the target quadruples, the source original text associated with the target quadruples, the additional features associated with the target quadruples, and the compilation and research description as the input information of the large language model, and output the archival knowledge compilation and research results through the large language model.

[0009] In a second aspect, an embodiment of the present application provides an archival knowledge compilation and research system, which is applicable to the power system and includes:

[0010] An acquisition module, configured to acquire compilation and research requirement information and an archival knowledge graph, rewrite the compilation and research requirement information by using prompt engineering technology to obtain a compilation and research description. The archival knowledge graph is constructed based on quadruples extracted from the text to be extracted, and the text to be extracted is archival materials obtained in advance;

[0011] A recall module, configured to determine and recall target quadruples matching the compilation and research description in the archival knowledge graph;

[0012] A determination module, configured to determine additional features according to the requirement type associated with the compilation and research description, and associate the additional features with the target quadruples. The additional features are used to determine the proportion of the description text of the relevant text content corresponding to the target quadruples in the archival knowledge compilation and research result;

[0013] A compilation and research module, configured to use the target quadruples, the source original text associated with the target quadruples, the additional features associated with the target quadruples, and the compilation and research description as input information of a large language model, and output the archival knowledge compilation and research result through the large language model.

[0014] In a third aspect, an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the first aspect above.

[0015] In a fourth aspect, an embodiment of the present application provides a storage medium storing computer-executable instructions for executing the method described in the first aspect above.

[0016] Through the above technical solutions, by constructing an archival knowledge graph based on quadruples, accurately locate archival materials matching the compilation and research requirements, rewrite the compilation and research requirement information by using prompt engineering technology to generate a compilation and research description, then determine the target quadruples and associate additional features, and finally input the relevant input information into a large language model to generate a compilation and research result. This process not only improves the accuracy level of the input instructions of the large language model, helps to overcome the hallucination of the large language model, improves the credibility and text quality of the target compilation and research content output by the large language model, but also can flexibly adjust the compilation and research ideas and perspectives according to different requirements, and expand the depth and breadth of compilation and research. Description of the Drawings

[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Shows a schematic flowchart of a file knowledge compilation and research method provided by an embodiment of the present application;

[0019] Figure 2 Shows a schematic flowchart of a file knowledge compilation and research method provided by an embodiment of the present application;

[0020] Figure 3 Shows a flowchart of a file knowledge compilation and research system according to an embodiment of the present application;

[0021] Figure 4 Shows a schematic diagram of the process of generating compilation and research content according to an embodiment of the present application;

[0022] Figure 5 Shows a block diagram of a computer device according to an embodiment of the present application. Detailed implementation manners

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0024] This method is applicable to the power system, especially the file knowledge compilation and research in the power system scenario. The file knowledge compilation and research work in the power system involves in-depth systematic sorting, detailed analysis, and comprehensive research of the file materials related to the power system, aiming to transform these materials into knowledge resources with practical application value. These knowledge resources play an important supporting role in all aspects such as the planning, design, operation, and maintenance of the power system, and can help relevant personnel make more efficient decisions and operations.

[0025] The archival knowledge graph is a way of information organization. By constructing a graphical data structure, it visually displays various information elements and their interrelationships in the power system archives. This graph can not only reveal the internal connections between archival materials but also provide users with intuitive query and analysis tools, thus greatly improving the retrieval efficiency and utilization value of archival materials. Through the archival knowledge graph, users can quickly locate the information they need, deeply explore the underlying knowledge behind the archival materials, and thereby provide strong data support for the optimization and innovation of the power system.

[0026] An embodiment of this application provides a method for archival knowledge compilation and research, as Figure 1 shown. This method includes:

[0027] 101. Obtain the compilation and research requirement information and the archival knowledge graph, and use prompt engineering technology to rewrite the compilation and research requirement information to obtain a compilation and research description. The archival knowledge graph is constructed based on the quadruples extracted from the text to be extracted, and the text to be extracted is the archival materials obtained in advance.

[0028] Among them, during the construction of the archival knowledge graph, a quadruple is a basic data structure used to represent entities, relationships, and time in archival materials. Quadruples can more comprehensively describe the complex information and temporal relationships in the archives, which not only helps to deeply explore the potential value in archival materials but also enhances the ability to understand and analyze archival content.

[0029] The archival knowledge graph organizes and represents archival materials in the form of a graph. Through technologies such as entity recognition and relationship extraction, it extracts time, entities, and the relationships between them from archival texts and constructs a knowledge network in the form of nodes and edges. In this technical solution, the archival knowledge graph is constructed based on the quadruples extracted from the text to be extracted, and the text to be extracted is the archival materials obtained in advance, that is, the available archival materials currently owned. By extracting these materials before specific compilation and research requirements, it can prepare for subsequent various specific compilations and research. The archival knowledge graph can visually display the internal connections and knowledge structure of archival materials and provide strong support for the compilation and research work.

[0030] Prompt engineering technology is to guide large language models (LLMs) to generate accurate and insightful outputs by designing and improving input prompts. In this solution, by using prompt engineering technology to rewrite the compilation and research requirement information to obtain a compilation and research description, in this way, a more clear and targeted task instruction can be given to the large language model. In this way, the model can better understand and focus on the core points of the compilation and research requirements, so that in the subsequent archival knowledge graph retrieval and compilation and research result generation processes, it can more accurately match and extract relevant information, improving the efficiency and quality of the compilation and research work.

[0031] The results of archival knowledge compilation and research refer to the achievements formed through in-depth research, analysis, and collation of archival materials. They are usually presented in the form of texts, charts, etc., and can provide valuable information and insights for specific research topics or practical needs. In this technical solution, through technical means such as archival knowledge graphs, prompt engineering techniques, and large language models, the automated and intelligent generation of archival knowledge compilation and research results has been achieved. Such compilation and research results are not only more accurate, comprehensive, and in-depth in content, but also more readable and practical in form, capable of better meeting the needs of different users for archival materials, and enhancing the service efficiency and social value of archival work.

[0032] 102. In the archival knowledge graph, determine and recall the target quadruples that match the compilation and research description.

[0033] In the embodiments of this application, the compilation and research description clarifies information such as the theme and key points of the compilation and research. By retrieving and matching in the archival knowledge graph, relevant quadruples can be found. The information such as entities and relationships contained in these target quadruples is the core content that needs to be focused on and deeply explored in the compilation and research work. For example, if the compilation and research description is about the research on the technological development status in a certain historical period, then the target quadruples can be quadruples containing relevant entities such as specific time, innovative processes, technical parameters, and technical achievements and their relationships that match the technological development status in that period. They provide an accurate material basis for the subsequent generation of compilation and research content.

[0034] 103. According to the type of requirements associated with the compilation and research description, determine additional features, and associate the additional features with the target quadruples. The additional features are used to determine the proportion of the description text of the relevant text content corresponding to the target quadruples in the archival knowledge compilation and research results.

[0035] In the embodiments of this application, the additional features are determined according to the type of requirements associated with the compilation and research description, and are features used to quantify and guide the proportion of the description text of the relevant text content corresponding to the target quadruples in the archival knowledge compilation and research results. The introduction of additional features makes the compilation and research work more efficient and targeted, and also provides a new perspective and method for the in-depth exploration and utilization of archival knowledge.

[0036] 104. Use the target quadruples, the source original text associated with the target quadruples, the additional features associated with the target quadruples, and the compilation and research description as the input information of the large language model, and output the archival knowledge compilation and research results through the large language model.

[0037] In the embodiments of the present application, the large language model is a model with powerful language understanding and generation capabilities constructed based on deep learning technology through training on a vast amount of text data. It can understand the semantics of the input information and generate coherent, accurate, and logical text outputs based on this information. In this technical solution, the target quadruple, the source original text associated with the target quadruple, the additional features associated with the target quadruple, and the compilation and research description are used as input information. The large language model will comprehensively consider this information, apply its internal complex algorithms and knowledge reserves, and generate the final compilation and research results of archival knowledge.

[0038] The archival knowledge compilation and research method provided in this embodiment constructs an archival knowledge graph based on quadruples to accurately locate archival materials that match the compilation and research requirements. With the help of prompt engineering technology, the compilation and research requirement information is rewritten to generate a compilation and research description, and then the target quadruple is determined and associated with additional features. Finally, the relevant input information is input into the large language model to generate the compilation and research results. This process not only improves the accuracy level of the input instructions of the large language model, helps to overcome the hallucination of the large language model, and improves the credibility and text quality of the target compilation and research content output by the large language model, but also can flexibly adjust the compilation and research ideas and perspectives according to different needs, expanding the depth and breadth of compilation and research.

[0039] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the implementation process of this embodiment, the embodiments of the present application provide an archival knowledge compilation and research method, as Figure 2 shown, the method includes:

[0040] 201. Obtain the compilation and research requirement information, and use prompt engineering technology to rewrite the compilation and research requirement information to obtain a compilation and research description.

[0041] The compilation and research requirement information for the power system reflects the compilation and research requirements in the operation and management work of the power system. For example, the compilation and research requirement information such as "the financial status of Company A" can be input into the compilation and research system. In short, the compilation and research requirement information reflects the theme of the content to be compiled and researched.

[0042] Furthermore, for specific compilation and research requirement information, prompt engineering technology can be used to rewrite its expression into a more targeted compilation and research description. In the embodiments of the present application, the instruction template of "task description - requirement to be rewritten" is used as the input for each rewriting task of the large model. The Prompt template for requirement rewriting is formulated in advance as follows:

[0043] "Rewrite the following compilation and research requirement information into a suitable compilation and research description for database index query. Note that only the rewritten content is output. Compilation and research requirement information: {query_text}." Then use the user's compilation and research requirement information Replace the placeholder {query_text} in the template to obtain the Prompt text , and the model rewrite output is the compilation and research description . Prompt engineering guides the model to generate the desired output text by providing the initial text prompt or question, enabling users to interact with the model in natural language and guiding the model to generate the required text through appropriate prompts.

[0044] 202. Build an archive knowledge graph and knowledge index.

[0045] In the embodiment of the present application, the system first reads the stored archive materials in the preset storage space as the text to be extracted. Then, through a pre-trained information extraction model, entities and the relationships between entities and relevant time information in the text to be extracted are identified to form a quadruple, that is, the quadruple includes entities, relationships, and time. Among them, time, as an important dimension in the knowledge graph, can record the specific time point or time range when an event occurs and has important application value. It should be noted that when inputting the text to be extracted into the information extraction model for quadruple extraction, the metadata associated with the text to be extracted can also be synchronously input into the information extraction model so that the information extraction model can better understand the text content of the text to be extracted. Further, based on the quadruple, a knowledge graph is quickly constructed. The knowledge graph realizes the effective association between different document knowledge and will be used as the basis for generating the content of archive compilation and research to ensure the integrity and reliability of the automatically compiled content.

[0046] It should be noted that to improve the extraction effect of the information extraction model, the information extraction model is obtained by performing supervised fine-tuning and / or low-rank adaptation fine-tuning on the model parameters of the historical text information extraction model based on the training set, and the training set is obtained by performing knowledge distillation on the historical information set. The specific training process is as follows:

[0047] First, an initial historical text information extraction model can be set. This model is fundamental and is used for subsequent processing work. Then, a series of historical information is collected, and these information form a set. This set includes the historical archive materials themselves, as well as the metadata information corresponding to these archive materials. In addition, it also includes the historical quadruples corresponding to the historical archive materials, and these quadruples can provide additional context information to help understand the historical materials more accurately. In this way, a high-quality unlabeled text dataset can be constructed. This dataset not only contains rich historical information, but also due to the diversity of its sources, it can ensure the diversity and coverage breadth of the data, thus providing a solid foundation for subsequent analysis and research.

[0048] Next, perform knowledge distillation on the above-mentioned historical information set to obtain a training set. Specifically, according to the principle of knowledge distillation, transfer the knowledge possessed by the teacher model to the student model. In this process, those excellent-performing teacher models will be selected, such as popular closed-source large models like the GPT series, to deeply extract and annotate the previously mentioned high-quality unlabeled text data set, and finally form a training set. The so-called knowledge distillation means that in the absence of available labeled data, the teacher model is required to provide accurate hard labels for quadruple entities, relationships, and time in the original text. Among them, the hard label refers to the most likely class label, that is, the predicted probability distribution of all classes. The student model will use these hard labels for training to learn how to identify quadruple entities, relationships, and time in the text. The main advantage of using hard labels is to simplify the training process and at the same time ensure that the student model can accurately reproduce the decision-making process of the teacher model. Therefore, by directly using hard labels for training, the student model can adapt faster and focus on the execution of the core task. This method reduces the complexity required for the model to interpret soft labels, thereby improving the training efficiency and model performance.

[0049] Furthermore, based on the training set, fine-tune the historical text information extraction model to obtain an information extraction model. In this process, the ways to fine-tune the historical text information extraction model include Supervised Finetuning (SFT) and Low-Rank Adaptation (LoRA). Among them, supervised fine-tuning is used to make the historical text information extraction model adapt to specific goals, and low-rank adaptation fine-tuning is used to achieve efficient parameter fine-tuning of the historical text information extraction model.

[0050] Finally, use the vLLM framework to deploy the information extraction model.

[0051] Among them, the supervised fine-tuning of the model parameters of the historical text information extraction model includes using the following formula 1 to minimize the loss function of the historical text information extraction model on the historical information set.

[0052] Formula 1:

[0053] Among them, is the minimum value of the loss function, is the model parameter of the historical text information extraction model, is the number of samples in the training set; and are respectively the model input information and model output result of the th sample in the training set, represents for the The predicted output result of the model input information of one sample; Indicates the error level between the model output result and the predicted output result.

[0054] Specifically, for a training sample in the training set, denoted as S, S consists of two parts. One is the input, denoted as query, and the other is the output, denoted as , that is , when training the information extraction model, the predicted result for the input is , the SFT technology, i.e., the sequence-to-sequence fine-tuning technology, will intervene and play a role. It will calculate the loss function based on the predicted result and the target result . Then, through the backpropagation algorithm, the model will adjust its internal parameters according to the value of the loss function with the aim of minimizing this loss function. This process will be repeated continuously until the model parameters are adjusted to a relatively ideal level. In the application scenario of the power system, the above training process will be repeated for all training samples in the training set, so that the information extraction model can gradually learn and master the mapping relationship between the input and the output. Through such learning and adjustment, the model can finally achieve the goal of knowledge distillation, that is, to adapt to the knowledge extraction task of the power system, so as to effectively extract the information crucial for the operation of the power system from the data.

[0055] Among them, performing low-rank adaptation fine-tuning on the model parameters of the historical text information extraction model includes performing low-rank adaptation processing on the model parameters of the historical text information extraction model using the following formula 2,

[0056] Formula 2:

[0057] Among them, represents the adjustment amount of the model parameters, and are low-rank matrices, , , and respectively represent the number of rows and columns of and respectively represent the number of rows and columns of 。The weight matrix of the model is decomposed and adjusted with low rank by LoRA to achieve fine-tuning of the model without changing the overall parameter architecture of the model. Specifically, the LoRA technique replaces the original large matrix with two smaller matrices, which can reduce the number of parameters to be trained while retaining the expressive power of the model. During calculation, only the parameters of these small matrices are adjusted, while most of the original model parameters remain unchanged. When using the LoRA technique, the loss function usually remains the cross-entropy loss, which is used to measure the difference between the model prediction and the actual label. The parameter update focuses on the low-rank matrices introduced by LoRA rather than the parameters of the entire model. The LoRA technique does not directly fine-tune the complete weight matrix W in the large language model, but realizes the update through the product of two low-rank matrices A and B, reducing the number of parameters.

[0058] In addition, in order to accelerate similarity calculation and query processing when dealing with large-scale data, and thus facilitate the rapid retrieval and recall of subsequent quadruples, the embodiment of the present application also constructs a knowledge index for the quadruples based on the BGE-M3 model and automated code, and stores the knowledge graph and the knowledge index in the PostgreSQL database for addressing and recalling the corresponding quadruples based on the knowledge index during the process of determining the target quadruples related to the compilation and research description. PostgreSQL is an open-source object-relational database system. PostgreSQL supports efficient vector storage and retrieval through the pgvector extension. In addition, it has high scalability and flexibility and can handle complex data structures. The parallel query function of PostgreSQL greatly accelerates the processing speed of vectors and complex queries. In addition, PostgreSQL can also provide transaction consistency and ACID support, which ensures the integrity and consistency of data, thus providing a solid guarantee for data processing.

[0059] It should be noted that the knowledge index includes the vector index of the knowledge graph quadruples, the source original text of the knowledge graph quadruples, the source original metadata (source documents, etc.), the text tokenization results of the knowledge graph quadruples, etc., and realizes the storage of vectors, original texts, metadata, and text tokenization results in PostgreSQL through pgvector. Let a knowledge graph quadruple be , correspondingly, there is a corresponding vector index stored in PostgreSQL , the source original text of the knowledge graph quadruples , the source original metadata , and the text tokenization results of the knowledge graph quadruples . pgvector is an extension for PostgreSQL, specifically designed to support the storage and retrieval of vector data. For example, there is a specific knowledge graph quadruple. Corresponding to this knowledge graph quadruple, its vector index, its source text, and the metadata information of these original texts are stored in the PostgreSQL database. At the same time, the word segmentation results of the knowledge graph quadruple text are also stored. In this way, when it is necessary to retrieve or analyze information in the knowledge graph, these indexes can be used to quickly locate specific data, thereby improving the efficiency and accuracy of data processing.

[0060] 203. In the archival knowledge graph, identify and recall the target quadruple that matches the editorial description.

[0061] In an embodiment of the present application, multiple recall strategies are pre-set, namely, the BM25 algorithm recall strategy, the semantic recall strategy, and the hybrid recall strategy. In the actual operation process, a recall strategy can be randomly selected from multiple recall strategies to perform a first recall on all quadruple groups in the archive knowledge graph to obtain a first quadruple group to be recalled. Next, for the quadruple groups that have not been recalled, that is, the other quadruple groups except the first quadruple group to be recalled, the system randomly selects a recall strategy from the remaining recall strategies again, and uses the recall strategy to perform a second recall on the quadruple groups that have not been recalled to obtain a second quadruple group to be recalled. Finally, the first quadruple group to be recalled and the second quadruple group to be recalled are set as target quadruple groups.

[0062] In addition, the variance of the normalized sparse search relevance score, dense search relevance score and comprehensive score associated with each quadruple can be calculated to obtain the participation evaluation parameter associated with each quadruple. Then, the quadruple whose participation evaluation parameter reaches the first preset threshold is determined and recalled from all quadruple groups as the target quadruple.

[0063] Specifically, when the selected recall strategy is the BM25 algorithm recall strategy, for each quadruple in the archive knowledge graph or each quadruple that has not been recalled, the sparse retrieval relevance score associated with each quadruple is determined based on the following formula 3, and among all quadruple groups or quadruple groups that have not been recalled, the quadruple whose sparse retrieval relevance score reaches the second preset threshold is determined as the first quadruple to be recalled or the second quadruple to be recalled.

[0064] Formula 3:

[0065] in, It is a four-tuple, which is composed of several terms; , To compile and research descriptions, multiple terms are used composition; is the relevance of the quadruple to the sparse retrieval of the compilation and research description; is the term in the compilation and research description appears in the target quadruple with a frequency; is the quadruple with a length; is the average length of all quadruples; and are adjustment parameters; is the inverse document frequency, as shown in the following formula 4.

[0066] Formula 4:

[0067] where, is the total number of documents; is the number of texts containing the term.

[0068] When the selected recall strategy is the semantic recall strategy, for each quadruple in the archival knowledge graph or each unrecalled quadruple, based on the following formula 5, determine the dense retrieval relevance score associated with each quadruple, and among all quadruples or unrecalled quadruples, determine the quadruples with the dense retrieval relevance score reaching the third preset threshold as the first quadruples to be recalled or the second quadruples to be recalled.

[0069] Formula 5: ;

[0070] where, is the embedding vector after the compilation and research description is transformed by the BGE-M3 model; , , is the th embedding vector retrieved from the PostgreSQL database through the knowledge index of the quadruple. The knowledge index is constructed by the BGE-M3 model based on the quadruple, including but not limited to the vector index of the quadruple, the source original text of the quadruple, the source original text metadata of the quadruple, and the text tokenization result of the quadruple; and are both L2 norms; is the dense retrieval relevance score of the th quadruple.

[0071] When the selected recall strategy is the hybrid recall strategy, for each quadruple in the archival knowledge graph or each unrecalled quadruple, determine the sparse retrieval relevance score and the dense retrieval relevance score associated with each quadruple. Perform a weighted sum processing on the sparse retrieval relevance score and the dense retrieval relevance score associated with each quadruple to obtain the comprehensive score corresponding to each quadruple. Finally, among all quadruples or unrecalled quadruples, determine the quadruples whose comprehensive scores reach the fourth preset threshold as the first quadruples to be recalled. That is to say, take the sparse retrieval relevance score and the dense retrieval relevance score of the quadruple and the compilation and research theme as the necessary calculation conditions for each quadruple when screening the target group, and calculate the comprehensive score that can comprehensively reflect the relevance between the quadruple and the compilation and research description. Optionally, before calculating the comprehensive score, perform a normalization process on the sparse retrieval relevance score and the dense retrieval relevance score of the quadruple and the compilation and research description to make the two parameters at the same magnitude, further improving the accuracy of the calculation result.

[0072] It should be noted that when the recall strategy selected in the first recall or the recall strategy selected in the second recall is a hybrid recall strategy, the retrieval results may include some knowledge graph quadruples that are partially irrelevant or have an unsatisfactory ranking. Therefore, a re-ranking technique is introduced. The re-ranking technique in the embodiments of this application is a result re-ranking technique that integrates the BGE-Reranker-V2 semantic re-ranking model, which can ensure that the most relevant knowledge graph quadruples are ranked at the front. The re-ranking technique can combine other features to reduce redundant information. Therefore, after hybrid recall, it is necessary to optimize the compilation materials provided to the large model through further evaluation and ranking. The BGE-Reranker-V2 semantic re-ranking is similar to the aforementioned vector retriever based on BGE-M3. The BGE-Reranker-V2-M3 semantic re-ranking model is used to convert the query text and the knowledge graph quadruples into semantic vectors, and the cosine similarity between the semantic vectors is used to determine the vector similarity. Specifically, the BGE-Reranker-V2-M3 semantic re-ranking model is used to vectorize the compilation description through a mapping function to obtain a high-dimensional vector corresponding to the compilation description. The embedding vector corresponding to the target quadruple is retrieved from the PostgreSQL database through the knowledge index of the target quadruple. The high-dimensional vector and the embedding vector are input into a preset similarity calculation function to obtain the calculation result output by the preset similarity calculation function. The Sigmoid function is used to map the calculation result to the probability range of (0,1) to obtain the similarity between the high-dimensional vector and the embedding vector. Finally, all target quadruples are sorted in descending order of the similarity between the high-dimensional vector and the embedding vector. In the actual operation process, in order to optimize the system performance and result quality, a relevance score threshold can also be set to eliminate quadruples below the threshold, and file compilation and research are based on the remaining quadruples after elimination. Thereby, it helps the large language model to more deeply understand the compilation and research requirements and even the compilation and research theme, improves the accuracy level of the input instructions of the large language model, helps to overcome the hallucination of the large language model, and improves the credibility and text quality of the target compilation and research content output by the large language model.

[0073] 204. Determine additional features according to the type of requirement associated with the compilation description, and associate the additional features with the target quadruple.

[0074] In the embodiments of this application, to further improve the quality of the automated compilation and research results, additional features can be determined according to the type of requirement of the compilation description, and the additional features are associated with the target quadruple and used as the input to the large prediction model together to highlight the content that is more relevant to the compilation description.

[0075] Specifically, first determine the additional feature determination method associated with the requirement type. When the additional feature determination method is parameter selection, determine the parameter identifier associated with the requirement type, and select the parameter indicated by the parameter identifier from the sparse retrieval relevance score, dense retrieval relevance score, comprehensive score, and participation evaluation parameter associated with the target quadruple as the additional feature, and associate the additional feature with the target quadruple. For example, the compilation and research description is an overall discussion of xx business, its corresponding requirement type is the evaluation feature requirement type, its corresponding additional feature determination method is parameter selection, the parameter identifier associated with the requirement type is the sparse retrieval relevance score, and the sparse retrieval relevance score associated with the quadruple is used as the additional feature and associated with the corresponding quadruple. When the additional feature determination method is parameter calculation, determine the parameter calculation method and at least one parameter identifier associated with the requirement type, and select at least one parameter indicated by at least one parameter identifier from the sparse retrieval relevance score, dense retrieval relevance score, comprehensive score, and participation evaluation parameter associated with the target quadruple, and calculate the selected at least one parameter according to the parameter calculation method, and use the calculation result as the additional feature, and associate the additional feature with the target quadruple.

[0076] Thus, the association of these quadruples with additional features can more accurately reflect the compilation and research requirements, improve the accuracy level of the input instructions of the large language model, help to overcome the hallucination of the large language model, and improve the credibility and text quality of the target compilation and research content output by the large language model.

[0077] 205. Use the target quadruple, the source text associated with the target quadruple, the additional feature associated with the target quadruple, and the compilation and research description as the input information of the large language model, and output the archive knowledge compilation and research result through the large language model.

[0078] In addition, Figure 3 An archive knowledge compilation and research system according to an embodiment of the present application is shown. The compilation and research system includes an algorithm service module, a resource management module, a graph construction module, a graph index module, and a knowledge compilation and research module.

[0079] Among them, the algorithm service module is used to train the information extraction model and the large language model for compilation and research work, and adopts the vLLM framework for model deployment, providing algorithm services such as OCR services, knowledge extraction, knowledge indexing, and knowledge compilation and research for other modules. The resource management module can support uploading, cataloging, and storing archival files and their metadata to form various archival file libraries. The graph construction module uses OCR algorithms and related text processing algorithms to segment and extract text from archival files, and uses the information extraction model to identify time, entities, and their relationships in the text to be extracted, and stores them in the Nebula graph database. The graph indexing module uses knowledge indexing algorithms to construct vector indexes for quadruples in the knowledge graph and stores them in the PostgreSQL database. The knowledge compilation and research module uses knowledge compilation and research algorithms to recall relevant knowledge content in the knowledge graph according to the user's compilation and research requirements, and generates highly readable compilation and research results through the large language model.

[0080] Next, the present application will be further elaborated with examples in the actual scenario of the power system in combination with the above embodiments.

[0081] Taking the compilation and research requirement of the theme of "the deeds of Expert B in City A" as an example, combined with Figure 4 As shown, first, collect materials. Specifically, collect text materials related to the compilation and research description, such as relevant news documents, official documents, and other documents related to Expert B, and upload them to the compilation and research system in pdf or docx file format.

[0082] Next, construct a graph. Based on the graph construction module of the compilation and research system, after the file upload is completed, a file set F = {f1, f2,..., fn} is formed. Among them, each file is automatically split into several text blocks TB(F(n)) = {fntb1, fntb2,..., fntbn}. The quadruple extraction is performed on each text block through a domain-enhanced information extraction model to form a quadruple set. Further, for each text block, the compilation and research system will use the Prompt template of "Perform knowledge graph fact quadruple extraction on the following text: {tb}", and replace {tb} with the text content of the text block to form the input for the large language model for compilation and research work. The following is an example of the input content and output content when the large language model performs quadruple extraction on the text block extracted from the news document.

[0083] For example, the input of the information extraction model is:

[0084] Perform knowledge graph fact triple extraction on the following text:

[0085] "In 1xxx, B, at the age of xx, started working as a power line inspector at a certain unit in xx City after graduating from xx School. B has been rooted in the front line of power emergency repair and has won various awards, including a, b, and c."

[0086] The output of the information extraction model is the quadruples extracted from the above information:

[0087] [{"time":"1xxx","subject":"B", "predicate":"graduated from", "object": "xx School"},

[0088] {"time":"1xxx","subject":"B", "predicate":"won", "object": "various awards, including a, b, and c"},

[0089] {"time":"1xxx","subject":"B", "predicate":"worked as", "object": "power line inspector"}, ……]

[0090] Of course, the information input into the information extraction model is diverse, and the number of quadruples output by the information extraction model is also large, not limited to the samples given in this example.

[0091] Third, construct a knowledge index. Based on the graph index module of the compilation and research system, after completing the quadruple extraction work for all text blocks, a triple set is formed , and based on the BGE-M3 model, the above quadruples are transformed into 1024-dimensional vector representations and stored in the PostgreSQL database together with the text tokenization results, etc.

[0092] Fourth, conduct knowledge compilation and research. For the aforementioned compilation and research descriptions, recall the quadruple set through the target quadruple screening method described in the previous embodiments.

[0093] In a possible design, referring to the above embodiments, additional features are determined for the quadruples, and the additional features are associated with the corresponding quadruples and input into the large language model together. The form of the additional features can be {"Relevance":0.28}. In other words, the size of Relevance in each group is positively correlated with the proportion of the subsequent description content in the description text of the compilation and research results.

[0094] The input content of the final large language model is the above quadruples, the additional features associated with the quadruples, the source original text associated with the quadruples, and the compilation and research descriptions.

[0095] The embodiments of the present application provide an archive knowledge compilation and research system, including:

[0096] An acquisition module, configured to acquire compilation requirements information and an archive knowledge graph, rewrite the compilation requirements information using prompt engineering technology to obtain a compilation description, where the archive knowledge graph is constructed based on quadruples extracted from text to be extracted, and the text to be extracted is pre-acquired archive materials;

[0097] A recall module, configured to determine and recall target quadruples matching the compilation description in the archive knowledge graph;

[0098] A determination module, configured to determine additional features according to the requirement type associated with the compilation description, and associate the additional features with the target quadruples, where the additional features are used to determine the proportion of the description text of the relevant text content corresponding to the target quadruples in the archive knowledge compilation result;

[0099] A compilation module, configured to use the target quadruples, the source original text associated with the target quadruples, the additional features associated with the target quadruples, and the compilation description as input information of a large language model, and output the archive knowledge compilation result through the large language model

[0100] In a specific application scenario, the recall module is configured to select a recall strategy from a preset multiple recall strategies to perform a first recall on all quadruples in the archive knowledge graph to obtain first quadruples to be recalled, and select another recall strategy from the remaining recall strategies after selection to perform a second recall on other quadruples except the first quadruples to be recalled to obtain second quadruples to be recalled, and set the first quadruples to be recalled and the second quadruples to be recalled as the target quadruples, where the multiple recall strategies include a BM25 algorithm recall strategy, a semantic recall strategy, and a hybrid recall strategy; or, for any quadruple in all the quadruples, calculate the variance of the sparse retrieval relevance score, the dense retrieval relevance score, and the normalized comprehensive score associated with the quadruple, use the variance as a participation evaluation parameter, and determine and recall the quadruples whose participation evaluation parameter reaches a first preset threshold in all the quadruples as the target quadruples.

[0101] In a specific application scenario, when the selected recall strategy is a BM25 algorithm recall strategy, the recall module is configured to, for each quadruple in the archive knowledge graph or for other quadruples except the first quadruples to be recalled, determine the sparse retrieval relevance score associated with each quadruple based on the following formula, and determine the quadruples whose sparse retrieval relevance score reaches a second preset threshold in all the quadruples or other quadruples except the first quadruples to be recalled as the corresponding quadruples to be recalled;

[0102] ;

[0103] Among them, is the quadruple, and the quadruple consists of several terms; , is the compilation and research description, which consists of multiple terms ; is the sparse retrieval relevance between the quadruple and the compilation and research description; is the term in the compilation and research description appears in the quadruple is the length of the quadruple ; is the average length of all quadruples; and are adjustment parameters; is the inverse document frequency; when the selected recall strategy is the semantic recall strategy, for each quadruple in the archival knowledge graph or for other quadruples except the first quadruple to be recalled, based on the following formula, determine the dense retrieval relevance score associated with each quadruple, and among all the quadruples or other quadruples except the first quadruple to be recalled, determine the quadruples with the dense retrieval relevance score reaching the third preset threshold as the corresponding quadruples to be recalled;

[0104] ;

[0105] Among them, is the embedding vector after the compilation and research description is transformed by the BGE-M3 model; , , is the th embedding vector retrieved from the PostgreSQL database through the knowledge index of the quadruple. The knowledge index is constructed by the BGE-M3 model according to the quadruple, including but not limited to the vector index of the quadruple, the source original text of the quadruple, the source original text metadata of the quadruple, and the text tokenization result of the quadruple; and are both L2 norms; is the dense retrieval relevance score of the th quadruple.

[0106] In a specific application scenario, the recall module is used to, when the selected recall strategy is a hybrid recall strategy, determine the sparse retrieval relevance score and the dense retrieval relevance score associated with each quadruple in the archival knowledge graph or for quadruples other than the first quadruple to be recalled; perform a weighted summation process on the sparse retrieval relevance score and the dense retrieval relevance score associated with each quadruple to obtain a comprehensive score corresponding to each quadruple; and determine, among all the quadruples or quadruples other than the first quadruple to be recalled, the quadruples whose comprehensive scores reach a fourth preset threshold as the corresponding quadruples to be recalled.

[0107] In a specific application scenario, the recall module is further used to, when the recall strategy selected in the first recall or the recall strategy selected in the second recall is a hybrid recall strategy, vectorize the compilation and research description to obtain a high-dimensional vector corresponding to the compilation and research description; retrieve the embedding vector corresponding to the target quadruple in the PostgreSQL database through the knowledge index of the target quadruple; input the high-dimensional vector and the embedding vector into a preset similarity calculation function to obtain the calculation result output by the preset similarity calculation function; use the Sigmoid function to map the calculation result to the probability range of (0, 1) to obtain the similarity between the high-dimensional vector and the embedding vector; and sort all the target quadruples in descending order of the similarity between the high-dimensional vector and the embedding vector.

[0108] In a specific application scenario, the determination module is used to determine the additional feature determination method associated with the requirement type; when the additional feature determination method is parameter selection, determine the parameter identifier associated with the requirement type, and select the parameter indicated by the parameter identifier from the sparse retrieval relevance score, the dense retrieval relevance score, the comprehensive score, and the participation evaluation parameter associated with the target quadruple as the additional feature, and associate the additional feature with the target quadruple; when the additional feature determination method is parameter calculation, determine the parameter calculation method and at least one parameter identifier associated with the requirement type, and select at least one parameter indicated by the at least one parameter identifier from the sparse retrieval relevance score, the dense retrieval relevance score, the comprehensive score, and the participation evaluation parameter associated with the target quadruple, and calculate the selected at least one parameter according to the parameter calculation method, and use the calculation result as the additional feature, and associate the additional feature with the target quadruple.

[0109] In a specific application scenario, the device further includes: a training module.

[0110] The training module is used to read the stored archival materials in the preset storage space as the text to be extracted; and use the pre-constructed information extraction model to extract multiple quadruples from the text to be extracted to construct the archival knowledge graph, where the quadruple includes an entity, a relationship, and a time; among them, the information extraction model is obtained by performing supervised fine-tuning and / or low-rank adaptation fine-tuning on the model parameters of the historical text information extraction model based on a training set, and the training set is obtained by performing knowledge distillation on the historical information set, and the historical information set includes historical archival materials, the metadata corresponding to the historical archival materials, and the historical quadruples corresponding to the historical archival materials; among them, performing supervised fine-tuning on the model parameters of the historical text information extraction model includes: using the following expression to minimize the loss function of the historical text information extraction model on the historical information set,

[0111] ;

[0112] is the minimum value of the loss function, are the model parameters of the historical text information extraction model, is the number of samples in the training set; and are respectively the model input information and the model output result of the th sample in the training set, represents the predicted output result of the model input information for the th sample; represents the error level between the model output result and the predicted output result.

[0113] Among them, performing low-rank adaptation fine-tuning on the model parameters of the historical text information extraction model includes: performing low-rank adaptation processing on the model parameters of the historical text information extraction model,

[0114]

[0115] represents the adjustment amount of the model parameters, and are low-rank matrices, , , and respectively represent the number of rows and columns of, and respectively represent the number of rows and columns of, .

[0116] The device uses the solution described in any one of the above embodiments, and thus has all the above technical effects, which will not be elaborated herein.

[0117] In an exemplary embodiment, refer to Figure 5 , a computer device is further provided. The device includes a communication bus, a processor, a memory, and a communication interface, and may further include an input / output interface and a display device. Among them, each functional unit can complete mutual communication through the bus. The memory stores a computer program, and the processor is configured to execute the program stored on the memory to execute the file knowledge compilation and research method in the above embodiments.

[0118] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the file knowledge compilation and research method are implemented.

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented through hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0120] It should be understood that although the terms first, second, etc. may be used in the embodiments of the present application to describe execution units, these execution units should not be limited to these terms. These terms are only used to distinguish the execution units from each other. For example, without departing from the scope of the embodiments of the present application, the first execution unit may also be referred to as the second execution unit, and similarly, the second execution unit may also be referred to as the first execution unit.

[0121] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0122] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0123] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0124] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional unit.

[0125] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0126] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for compiling and researching archival knowledge, characterized in that: The method is applicable to a power system, comprising: Obtaining editing and research demand information and an archival knowledge graph, and rewriting the editing and research demand information using a prompt engineering technique to obtain an editing and research description, wherein the archival knowledge graph is constructed based on quadruples extracted from the text to be extracted, and the text to be extracted is the archival material obtained in advance; In the archive knowledge graph, determine and recall the target quadruple that matches the editing and research description; Determine additional features according to the type of requirements associated with the compilation and research description, and associate the additional features with the target quadruple, wherein the additional features are used to determine the proportion of the description text of the relevant text content corresponding to the target quadruple in the archive knowledge compilation and research results; The target quadruple, the source text associated with the target quadruple, the additional features associated with the target quadruple and the compilation description are used as input information of the large language model, and the archive knowledge compilation result is output through the large language model.

2. The method according to claim 1, characterized in that In the archive knowledge graph, a target quadruple matching the editing and research description is determined and recalled, including: A recall strategy is selected from a plurality of preset recall strategies to perform a first recall on all quadruples in the archive knowledge graph to obtain a first quadruple to be recalled, and a recall strategy is selected from the remaining recall strategies after the selection to perform a second recall on other quadruples except the first quadruple to be recalled to obtain a second quadruple to be recalled, and the first quadruple to be recalled and the second quadruple to be recalled are set as the target quadruples, wherein the plurality of recall strategies include the BM25 algorithm recall strategy, the semantic recall strategy and a mixed recall strategy; or, for any quadruple among all the quadruple groups, calculating the normalized variance of the sparse retrieval relevance score, dense retrieval relevance score and comprehensive score associated with the quadruple group, taking the variance as a participation evaluation parameter, and determining and recalling the quadruple group whose participation evaluation parameter reaches a first preset threshold value among all the quadruple groups as the target quadruple group, wherein the comprehensive score is obtained by weighted summing up the sparse retrieval relevance score and the dense retrieval relevance score associated with the quadruple group.

3. The method according to claim 2, characterized in that The step of selecting a recall strategy from among a plurality of preset recall strategies to perform a first recall on all quadruple groups in the archive knowledge graph to obtain a first quadruple group to be recalled includes: When the selected recall strategy is the BM25 algorithm recall strategy, for each quadruple in the archive knowledge graph or for other quadruple except the first quadruple to be recalled, based on the following formula, determine the sparse retrieval relevance score associated with each quadruple, and among all the quadruple or other quadruple except the first quadruple to be recalled, determine the quadruple whose sparse retrieval relevance score reaches the second preset threshold as the corresponding quadruple to be recalled; ; in, is the four-tuple, and the four-tuple consists of a number of terms; , For the compilation and research description, there are multiple terms composition; is the sparse search relevance of the quadruple to the compilation description; For the terms in the research description In the four The frequency of occurrence in For the four-tuple Length; is the average length of all quadruples; and To adjust the parameters; is the inverse document frequency; When the selected recall strategy is the semantic recall strategy, for each quadruple in the archive knowledge graph or for other quadruple except the first quadruple to be recalled, based on the following formula, a dense search relevance score associated with each quadruple is determined, and among all the quadruple or other quadruple except the first quadruple to be recalled, a quadruple whose dense search relevance score reaches a third preset threshold is determined as the corresponding quadruple to be recalled; ; in, The embedding vector of the compilation description after conversion by the BGE-M3 model; , , For the The embedding vector of the quadruple is retrieved from the PostgreSQL database through the knowledge index of the quadruple, wherein the knowledge index is constructed by the BGE-M3 model according to the quadruple, including but not limited to the vector index of the quadruple, the source text of the quadruple, the source text metadata of the quadruple, and the text segmentation result of the quadruple; and All are L2 norms; For the Dense retrieval relevance score of quadruple.

4. The method according to claim 3, characterized in that The method comprises: When the selected recall strategy is a hybrid recall strategy, for each quadruple in the archive knowledge graph or for other quadruple except the first quadruple to be recalled, determine a sparse search relevance score and a dense search relevance score associated with each quadruple; Performing weighted sum processing on the sparse search relevance score and the dense search relevance score associated with each of the quadruple groups to obtain a comprehensive score corresponding to each of the quadruple groups; A quadruple whose comprehensive score reaches a fourth preset threshold is determined from all the quadruple groups or other quadruple groups except the first quadruple group to be recalled as the corresponding quadruple group to be recalled.

5. The method according to claim 4, characterized in that The method further comprises: When the recall strategy selected in the first recall or the recall strategy selected in the second recall is a mixed recall strategy, vectorizing the compilation and research description to obtain a high-dimensional vector corresponding to the compilation and research description; Retrieving the embedding vector corresponding to the target quadruple in the PostgreSQL database through the knowledge index of the target quadruple; Inputting the high-dimensional vector and the embedded vector into a preset similarity calculation function, and obtaining a calculation result output by the preset similarity calculation function; The calculation result is mapped to a probability range of (0, 1) using a Sigmoid function to obtain the similarity between the high-dimensional vector and the embedded vector; All the target quadruple groups are sorted in descending order according to the similarity between the high-dimensional vector and the embedded vector.

6. The method according to claim 2, characterized in that Determining additional features according to the type of requirements associated with the compilation and research description includes: Determine a method for determining additional features associated with the requirement type; When the additional feature determination method is parameter selection, determine the parameter identifier associated with the requirement type, and select the parameter indicated by the parameter identifier from the sparse search relevance score, dense search relevance score, comprehensive score and participation evaluation parameters associated with the target quadruple as the additional feature, and associate the additional feature with the target quadruple; When the additional feature determination method is parameter calculation, the parameter calculation method and at least one parameter identifier associated with the requirement type are determined, and at least one parameter indicated by the at least one parameter identifier is selected from the sparse retrieval relevance score, dense retrieval relevance score, comprehensive score and participation evaluation parameters associated with the target quadruple, and the selected at least one parameter is calculated according to the parameter calculation method, the calculation result is used as the additional feature, and the additional feature is associated with the target quadruple.

7. The method according to claim 1, characterized in that After obtaining the editing and research demand information, the method further includes: Reading the archive data stored in the preset storage space as the text to be extracted; Using a pre-built information extraction model to extract multiple quadruples from the text to be extracted to construct the archive knowledge graph, wherein the quadruples include entities, relations, and time; Wherein, the information extraction model is obtained by subjecting the model parameters of the historical text information extraction model to supervised fine-tuning and / or low-rank adaptive fine-tuning based on the training set, wherein the training set is obtained by subjecting the historical information set to knowledge distillation, wherein the historical information set includes historical archival materials, metadata corresponding to the historical archival materials, and historical quadruplets corresponding to the historical archival materials; The supervised fine-tuning of the model parameters of the historical text information extraction model includes: using the following expression to minimize the loss function of the historical text information extraction model on the historical information set, ; is the minimum value of the loss function, is the model parameter of the historical text information extraction model, is the number of samples in the training set; and They are respectively The model input information and model output results of samples, express For The predicted output result of the model input information of samples; Indicates the error level between the model output result and the predicted output result; The method of performing low-rank adaptive fine-tuning on the model parameters of the historical text information extraction model comprises: performing low-rank adaptive processing on the model parameters of the historical text information extraction model, represents the adjustment amount of the model parameters, and is a low-rank matrix, , , and Respectively The number of rows and columns, and Respectively The number of rows and columns, .

8. An archive knowledge compilation and research system, characterized in that: The archive knowledge compilation and research system is applicable to the power system, including: An acquisition module is used to acquire editing and research demand information and an archive knowledge graph, and rewrite the editing and research demand information using a prompt engineering technique to obtain an editing and research description. The archive knowledge graph is constructed based on the quadruple extracted from the text to be extracted, and the text to be extracted is the pre-acquired archive material; A recall module, used to determine and recall a target quadruple matching the editing and research description in the archive knowledge graph; A determination module, used to determine additional features according to the type of requirements associated with the compilation and research description, and associate the additional features with the target quadruple, wherein the additional features are used to determine the proportion of the description text of the relevant text content corresponding to the target quadruple in the archive knowledge compilation and research results; The compilation and research module is used to use the target quadruple, the source text associated with the target quadruple, the additional features associated with the target quadruple and the compilation and research description as input information of the large language model, and output the archive knowledge compilation and research results through the large language model.

9. A computer device, characterized in that: include: at least one processor; And, a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method of any one of claims 1 to 7.

10. A storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Search recall method and system based on knowledge graph representation learning

    CN115618113A

  • Search question-answering system and method based on large model and electronic equipment

    CN117708274A