Knowledge base construction method, knowledge processing method and related products

By integrating and inductively reasoning about documents in the knowledge base, higher-level documents are generated, which solves the problem of insufficient information in the existing knowledge base and improves the output accuracy of the large language model, especially in scenarios that require cross-document reasoning and abstract generalization.

CN120930738APending Publication Date: 2025-11-11SHENZHEN WANGYU COMPUTER NETWORK CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410585445.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

The existing knowledge base contains insufficient information, which makes the output of large language models inaccurate when dealing with complex problems, especially when cross-document reasoning and abstract generalization capabilities are required, making it difficult to provide accurate answers.

Method used

By integrating multiple documents in the first knowledge base to form multiple semantically related document sets, and using a large language model for inductive reasoning, a higher-level second document is generated, thus constructing a second knowledge base to improve the information content and cross-document reasoning capabilities of the knowledge base.

Benefits of technology

It improves the accuracy of output results of large language models when dealing with complex problems, especially in scenarios that require cross-document reasoning and abstract generalization capabilities, and can provide more accurate answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930738A_ABST
    Figure CN120930738A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge base construction method, a knowledge processing method and a related product. The construction method comprises the following steps: acquiring a plurality of first documents included in a first knowledge base; integrating the plurality of first documents to obtain a plurality of different document sets, wherein each document set in the plurality of different document sets comprises the first documents with semantic association relationship; carrying out inductive reasoning on the first document included in each document set based on the large language model and the document generation instruction to obtain a second document corresponding to each document set; and constructing a second knowledge base based on the plurality of first documents and the second documents corresponding to each document set. Therefore, the cross-document reasoning ability can be introduced, and the knowledge hierarchy of the second knowledge base is expanded through the induction reasoning ability of the large language model, so that the second knowledge base contains more information to meet the use requirements of the large language model, and the accuracy of the output result of the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method for constructing a knowledge base, a knowledge processing method, and related products. Background Technology

[0002] Large language models refer to deep learning models trained on massive amounts of text data. They can not only generate natural language text, but also deeply understand the meaning of text and handle various natural language tasks, such as summarizing, question answering, and translation.

[0003] In practical applications, large language models can leverage external knowledge bases for knowledge retrieval to enhance the accuracy of their output. However, existing knowledge bases only include the original documents added during construction, and the amount of information contained in these original documents may not meet the needs of large language models, resulting in inaccurate output from these knowledge bases. Summary of the Invention

[0004] This application provides a method for constructing a knowledge base, a knowledge processing method, and related products to meet the usage requirements of large language models and improve the accuracy of the output results of large language models.

[0005] The embodiments of this application disclose the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a method for constructing a knowledge base, including:

[0007] Retrieve multiple first documents included in the first knowledge base;

[0008] The multiple first documents are integrated to obtain multiple different document sets, and each document set includes first documents that have semantic relationships.

[0009] Based on the large language model and document generation instructions, inductive reasoning is performed on the first document included in each document set to obtain the second document corresponding to each document set;

[0010] A second knowledge base is constructed based on the plurality of first documents and the second documents corresponding to each document set.

[0011] Secondly, embodiments of this application provide a knowledge processing method, including:

[0012] In response to user input information, knowledge related to the user input information is obtained from a second knowledge base, which is obtained based on the knowledge base construction method provided in the first aspect above;

[0013] Text input instructions are generated based on knowledge related to the user input information and the user input information itself;

[0014] The text input command is input into a large language model, and the output information matching the user input information is obtained through the large language model.

[0015] Thirdly, embodiments of this application provide a knowledge base construction apparatus, comprising:

[0016] The document acquisition module is used to acquire multiple first documents included in the first knowledge base;

[0017] The document integration module is used to integrate the multiple first documents to obtain multiple different document sets, and each document set includes first documents that have semantic relationships.

[0018] The inductive reasoning module is used to perform inductive reasoning on the first document included in each document set based on the large language model and document generation instructions to obtain the second document corresponding to each document set.

[0019] The knowledge base construction module is used to construct a second knowledge base based on the plurality of first documents and the second documents corresponding to each document set.

[0020] Fourthly, embodiments of this application provide a knowledge processing apparatus, including:

[0021] The knowledge acquisition module is used to retrieve knowledge related to the user input information from a second knowledge base in response to user input information. The second knowledge base is obtained based on the knowledge base construction method provided in the first aspect above.

[0022] The first instruction generation module is used to generate text input instructions based on knowledge related to the user input information and the user input information;

[0023] The output information acquisition module is used to input the text input command into the large language model and obtain output information that matches the user input information through the large language model.

[0024] Fifthly, embodiments of this application provide an electronic device, the device including a processor and a memory:

[0025] The memory is used to store computer programs and to transfer the computer programs to the processor;

[0026] The processor is configured to execute, according to instructions in the computer program, the steps of the knowledge base construction method provided in the first aspect, or the steps of the knowledge processing method provided in the second aspect.

[0027] Sixthly, embodiments of this application provide a computer-readable storage medium for storing program code, the program code being used to execute the steps of the knowledge base construction method provided in the first aspect above, or the steps of the knowledge processing method provided in the second aspect above.

[0028] In a seventh aspect, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the knowledge base construction method provided in the first aspect, or the steps of the knowledge processing method provided in the second aspect.

[0029] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0030] In this embodiment, after obtaining multiple first documents included in the first knowledge base, the multiple first documents can be integrated to obtain multiple different document sets. Each document set includes first documents with semantic relationships. Then, based on the large language model and document generation instructions, inductive reasoning is performed on the first documents included in each document set to obtain the second document corresponding to each document set. Finally, based on the multiple first documents and the second document corresponding to each document set, a second knowledge base is constructed. In this way, by integrating the first documents with semantic relationships in the first knowledge base to form multiple different document sets, cross-document reasoning capabilities can be introduced, improving the ability of the large language model to perform knowledge retrieval with the help of the second knowledge base. Furthermore, by leveraging the text generation and understanding capabilities of the large language model to perform inductive reasoning on the first documents in each document set to generate second documents, the knowledge hierarchy of the second knowledge base can be expanded through the inductive reasoning capabilities of the large language model. This allows the second knowledge base to contain more information to meet the usage needs of the large language model, thereby improving the accuracy of the output results of the large language model. Attached Figure Description

[0031] Figure 1 A flowchart illustrating a method for constructing a knowledge base as provided in an embodiment of this application;

[0032] Figure 2 A schematic diagram of the overall architecture of a knowledge base construction scheme provided in an embodiment of this application;

[0033] Figure 3 A flowchart illustrating a knowledge processing method provided in an embodiment of this application;

[0034] Figure 4 A schematic diagram of the overall architecture of a knowledge processing method provided in an embodiment of this application;

[0035] Figure 5A schematic diagram illustrating an application scenario of the knowledge processing method provided in the embodiments of this application;

[0036] Figure 6 A schematic diagram of a knowledge base construction apparatus provided in an embodiment of this application;

[0037] Figure 7 This is a schematic diagram of the structure of a knowledge processing device provided in an embodiment of this application;

[0038] Figure 8 This application provides a schematic diagram of the structure of a server according to an embodiment of the present application.

[0039] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0040] As mentioned earlier, in practical applications, large language models can incorporate external knowledge through Retrieval Augmented Language Models (RALM). This involves using external knowledge bases for knowledge retrieval and leveraging the retrieval results to improve the accuracy of the large language model's output, thus addressing the language model illusion problem. Specifically, when a user poses a question to the large language model, the RALM's retrieval mechanism responds by retrieving relevant knowledge from an external knowledge base. This knowledge is then combined with the user's input to generate a text input prompt. This prompt serves as input to the large language model, which outputs information related to the user's input. This output incorporates external knowledge, thus mitigating the language model illusion problem to some extent. However, existing knowledge bases often contain only raw documents added during knowledge base construction. The information contained in these raw documents may not meet the needs of the large language model, leading to inaccurate outputs from the large language model using this knowledge base.

[0041] Specifically, as the capabilities of large language models continue to improve, users may ask more complex questions that require reasoning between different documents. For example, a user input question like "Which regions does River A flow through?" is relatively simple; the answer might be directly found in a document in the knowledge base, and the knowledge from a single document could solve the problem. However, a user input question like "What is the largest city in the region through which River A flows?" involves more knowledge points and may not be answered by a single document. This is because the large language model needs to retrieve documents related to the River A basin, as well as question-and-answer methods related to each city within the River A basin. The answer is only obtained through aggregating and reasoning this information. However, due to the low correlation between documents related to the River A basin and documents related to individual cities, the retrieval system struggles to find these related documents, resulting in the large language model failing to output correct responses and lacking accuracy.

[0042] Furthermore, existing knowledge bases only include original documents. While these documents contain a wealth of detailed knowledge, it lacks further summarization and generalization, resulting in a low level of knowledge hierarchy. Moreover, current retrieval systems generally rely on text segmentation matching or text semantics to obtain search results. Therefore, when faced with a relatively abstract question like "What is the significance of River A to the people of Country H?", if the original document does not include fields that match the word segmentation or semantics of the question, even if the original document contains an answer matching the question, such as "River A is known as the father river of the people of Country H, and River A and River B are collectively called the mother rivers," the retrieval system will not return this knowledge. This will affect the accuracy of the final response from the large language model.

[0043] To address the aforementioned issues, this application provides a method for constructing a knowledge base, which may include: after obtaining multiple first documents included in a first knowledge base, integrating the multiple first documents to obtain multiple different document sets, each of the multiple different document sets including first documents with semantic relationships; then performing inductive reasoning on the first documents included in each document set based on a large language model and document generation instructions to obtain second documents corresponding to each document set; and finally constructing a second knowledge base based on the multiple first documents and the second documents corresponding to each document set.

[0044] In this way, by integrating the first documents with semantic relationships in the first knowledge base to form multiple different document sets, cross-document reasoning capabilities can be introduced, enhancing the ability of the large language model to retrieve knowledge using the second knowledge base. Furthermore, by leveraging the text generation and understanding capabilities of the large language model to perform inductive reasoning on the first documents in each document set to generate second documents, the knowledge hierarchy of the second knowledge base can be expanded through the inductive reasoning capabilities of the large language model. This allows the second knowledge base to contain more information to meet the usage needs of the large language model, thereby improving the accuracy of the output results of the large language model.

[0045] It should be noted that the embodiments of this application do not limit the executing entity of the technical solution of this application. For example, the knowledge base construction method and knowledge processing method provided in the embodiments of this application can be applied to a user terminal or a server, or processed collaboratively by a user terminal and a server. As an example, the user terminal includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The server can be a standalone server, a cluster server, or a cloud server.

[0046] For ease of understanding, the terminology that may be involved in the embodiments of this application will be introduced below.

[0047] Large language models refer to deep learning models trained on massive amounts of text data. Their parameter scale exceeds a certain value, typically exceeding five billion. Large language models can receive user prompts as input and generate corresponding outputs. They can handle various natural language tasks, such as chat, summarization, question answering, and translation. In practical applications, large language models can include Tencent's Hunyuan Large Model or Tongyi Qianwen Large Model, among others.

[0048] Language model illusion refers to the phenomenon where existing large language models, when directly applied to dialogue in certain specific domains, produce incorrect guesses and give seemingly plausible but ultimately flawed results due to a lack of relevant knowledge during training.

[0049] A knowledge base is generally an unstructured knowledge base. It can include text documents containing knowledge from several domains. These documents contain directly readable text, facilitating retrieval by search engines and allowing for direct understanding by large language models.

[0050] Knowledge level refers to the level of abstract generalization ability of the knowledge contained in a text. For a given text, the more primitive and detailed the knowledge information it contains, the lower the knowledge level and the weaker its ability to abstract and generalize. Conversely, the broader the scope of knowledge information and the fewer the details, the higher the knowledge level and the stronger its ability to abstract and generalize. For example, consider text 1, "River A, also known as the A River, is the nth longest river in Asia and the mth longest river in the world, with a total length of t kilometers. Its main stream originates from the C Mountains in the eastern part of Plateau B, traverses the southwest (including provinces D, E, F, and G), central (including provinces X, Y, and Z), and eastern (including provinces P, Q, and W) of Country H, and flows into the ocean at City S," and text 2, "River A is the nth longest river in Asia and the mth longest river in the world, originating from Plateau B, traversing the southwest, central, and eastern parts of Country H before flowing into the ocean." Both texts contain essentially the same content, but text 1 contains more detailed information, while text 2 is more like a summary of text 1. Correspondingly, text 1 has a lower knowledge level, while text 2 has a higher knowledge level.

[0051] Artificial Intelligence (AI) is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. In the knowledge base construction and knowledge processing scenarios provided in this application's embodiments, AI technology can utilize machines to perceive the document content in the knowledge base, thereby introducing cross-document reasoning capabilities. Furthermore, by leveraging the inductive reasoning capabilities of AI technology, the knowledge base's knowledge hierarchy can be expanded, allowing it to contain more information to meet the needs of large language models, thus improving the accuracy of the large language model's output.

[0052] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0053] This application provides a method for constructing a knowledge base and related products, primarily involving Natural Language Processing (NLP) technology. NLP is an important field within computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. NLP is a science integrating linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, i.e., the language people use daily, and thus it is closely related to linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs. In this application, relying on NLP technology, the cross-document reasoning and inductive abilities of the knowledge base can be continuously improved, thereby generating new documents that contain more information to meet the needs of large language models, thus improving the accuracy of the output results of large language models.

[0054] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] Figure 1 This is a flowchart illustrating a method for constructing a knowledge base as provided in an embodiment of this application. (In conjunction with...) Figure 1 As shown in the embodiments of this application, the method for constructing a knowledge base may include:

[0056] S101: Obtain multiple first documents included in the first knowledge base.

[0057] In the embodiments of this application, the first knowledge base may refer to the original knowledge base adopted by the large language model; the first document may refer to the original documents included in the original knowledge base.

[0058] The process of obtaining the multiple first documents included in the first knowledge base is not specifically limited in this embodiment. For example, the first knowledge base can be pre-stored in the execution entity of this embodiment. When knowledge base construction is required, the execution entity can obtain the first knowledge base by local reading to obtain the multiple first documents included therein. Alternatively, the first knowledge base can be stored on another data storage server, and the execution entity of this embodiment can obtain the first knowledge base by accessing the data storage server to obtain the multiple first documents included therein.

[0059] S102: Integrate multiple first documents to obtain multiple different document sets.

[0060] Each of the multiple different document sets can include a first document that has a semantic relationship. The first documents included in different document sets can be completely different or partially different, so that the second documents corresponding to the subsequently generated different document sets will not be completely identical, thus avoiding duplication of content in the second documents.

[0061] In this embodiment of the application, the process of integrating multiple different document sets, namely step S102, is not specifically limited. For ease of understanding, it will be described below in conjunction with a variety of possible implementation methods.

[0062] As one possible implementation, step S102 may include: encoding multiple first documents to obtain encoding vectors corresponding to each first document; constructing feature matrices for multiple first documents based on their respective encoding vectors; and clustering the feature matrices using a clustering model to obtain multiple different clusters, which serve as multiple different document sets. In practical applications, sentence vector models, such as the semantic vector model (BAAI General Embedding, BGE), can be used to encode multiple first documents, obtaining encoding vectors corresponding to each first document. These encoding vectors can be used to measure the degree of similarity between multiple first documents. Then, these encoding vectors can be used to construct a feature matrix. This feature matrix can then be clustered using a clustering model, such as a hierarchical clustering model or a K-means clustering model, to obtain multiple different clusters. In this way, by clustering documents, first documents with semantic relationships can be quickly and accurately integrated, facilitating the subsequent construction of a second knowledge base.

[0063] Furthermore, as mentioned earlier, knowledge level is used to represent the abstract generalization ability of the knowledge contained in a text. For a text, the richer the knowledge details, the worse the abstract generalization ability and the lower the knowledge level; the broader the scope of knowledge coverage, the stronger the abstract generalization ability and the higher the knowledge level. Based on this, in the embodiments of this application, document sets with different clustering granularities can be used to generate second documents, so that the second documents include information of different knowledge levels, thereby further expanding the knowledge levels of the second knowledge base and helping to improve the accuracy of the large language model.

[0064] Specifically, the process of clustering the feature matrix can include: based on a clustering model, clustering the feature matrix at multiple clustering granularities to obtain multiple clusters under each clustering granularity; and using the multiple clusters under each clustering granularity as multiple different document sets. In this way, by adjusting the clustering parameters to obtain multiple clustering granularities, it is helpful to generate second documents at different knowledge levels, thereby facilitating the subsequent expansion of the knowledge levels of the second knowledge base and improving the response accuracy of the large language model. In practical applications, multiple clustering granularities can be obtained by adjusting the clustering parameters of the clustering model. Taking a hierarchical clustering model as an example, the clustering granularity can be adjusted by adjusting the clustering threshold of the hierarchical clustering model. The smaller the clustering threshold, the finer the clustering granularity, the fewer first documents included in each cluster, the weaker the summarizing ability of the generated second documents, and the lower the knowledge level; conversely, the larger the clustering threshold, the coarser the clustering granularity, the more first documents included in each cluster, the stronger the summarizing ability of the generated second documents, and the higher the knowledge level. Taking the K-means clustering model as an example, the clustering granularity can be adjusted by changing the K value. The larger the K value, the finer the clustering granularity, the fewer the number of first documents included in each cluster, the weaker the summarizing ability of the generated second documents, and the lower the knowledge level. The smaller the K value, the coarser the clustering granularity, the more the number of first documents included in each cluster, the stronger the summarizing ability of the generated second documents, and the higher the knowledge level.

[0065] As another possible implementation, step S102 may include: encoding multiple first documents to obtain encoding vectors corresponding to each first document; classifying the encoding vectors corresponding to the multiple first documents based on a classification model to obtain document sets corresponding to multiple categories as multiple different document sets. In practical applications, sentence vector models, such as semantic vector models (BAAI General Embedding, BGE), can be used to encode multiple first documents to obtain encoding vectors corresponding to each first document. These encoding vectors can be used to measure the degree of similarity between multiple first documents. Then, classification models, such as neural network classification models or keyword classification models, can be used to classify the documents to obtain document sets corresponding to multiple categories. In this way, by classifying the first documents, first documents with semantic relationships can be quickly and accurately integrated together, facilitating the subsequent construction of a second knowledge base.

[0066] Furthermore, as mentioned earlier, knowledge level is used to represent the abstract generalization ability of the knowledge contained in a text. For a text, the richer the knowledge details, the worse the abstract generalization ability and the lower the knowledge level; the broader the scope of knowledge coverage, the stronger the abstract generalization ability and the higher the knowledge level. Based on this, in the embodiments of this application, a second document can be generated using document sets with different classification granularities, so that the second document includes information of different knowledge levels, thereby further expanding the knowledge levels of the second knowledge base and helping to improve the accuracy of the large language model.

[0067] Specifically, the process of classifying the encoding vectors corresponding to multiple first documents can include: classifying the encoding vectors corresponding to multiple first documents at multiple classification granularities based on a classification model, obtaining document sets corresponding to multiple categories under each classification granularity; and using the document sets corresponding to multiple categories under each classification granularity as multiple different document sets. In this way, by adjusting the classification parameters to obtain multiple classification granularities, it is helpful to generate second documents at different knowledge levels, thereby facilitating the subsequent expansion of the knowledge levels of the second knowledge base and improving the response accuracy of the large language model. In practical applications, multiple classification granularities can be obtained by adjusting the classification parameters of the classification model. Taking a keyword classification model as an example, the classification granularity can be adjusted by changing the granularity of the category keywords in the keyword classification model; for example, the granularity of the category keyword "food" is greater than the granularity of the category keyword "fruit". The finer the granularity of category keywords, the finer the classification granularity, the fewer the number of first documents corresponding to each category, the weaker the summarizing ability of the generated second documents, and the lower the knowledge level; the coarser the granularity of category keywords, the coarser the classification granularity, the more the number of first documents corresponding to each category, the stronger the summarizing ability of the generated second documents, and the higher the knowledge level.

[0068] Additionally, it should be noted that the integration process for multiple first documents specifically involves integrating them based on their semantics, thereby grouping together first documents that have semantic relationships. This application embodiment does not specifically limit the process of integrating based on the semantics of the first documents; any existing or future clustering or classification model capable of semantic-based text integration can be used to implement the above integration steps.

[0069] S103: Based on the large language model and document generation instructions, perform inductive reasoning on the first document included in each document set to obtain the second document corresponding to each document set.

[0070] In this embodiment, the document generation instruction refers to a prompt used to input a large language model to generate a second document. Therefore, the document generation instruction corresponding to each document set may include the first document in each document set, as well as task information for inductive reasoning on the first document.

[0071] Based on this, in this embodiment, document generation instructions can be obtained through the following steps: obtaining task information of the large language model, which can be used to instruct the large language model to perform inductive and / or inference operations; and generating document generation instructions corresponding to each document set based on the first document and task information included in each document set. In this way, suitable document generation instructions can be constructed using the first document and task information in each document set, facilitating the subsequent generation of a second document corresponding to each document set based on the document generation instructions and the inductive reasoning capabilities of the large language model. The generated second document can include cross-document reasoning information and has a higher knowledge level than the first document, thus containing more information to meet the usage requirements of the large language model and improving the accuracy of the large language model's output.

[0072] It should be noted that the specific text content of the document generation instruction is not limited in this embodiment. A document generation instruction only needs to include the first document in the corresponding document set and related task information. The related task information only needs to be used to instruct the large language model to perform inductive and / or inference operations. For ease of understanding, the content of the document generation instruction is illustrated below with reference to Table 1.

[0073] Table 1

[0074]

[0075] As can be seen in Table 1 above, a specific text content of the document generation instruction is shown, in which the relevant task information indicates that the large language model can perform inductive and inference operations.

[0076] Furthermore, the process of obtaining the second document mentioned above, namely step S103, may include: inputting the document generation instruction corresponding to each document set into the large language model, and obtaining the second document matched by the document generation instruction corresponding to each document set through the large language model.

[0077] S104: Construct a second knowledge base based on multiple first documents and the second documents corresponding to each document set.

[0078] In this embodiment, the second knowledge base may refer to a new knowledge base adopted by the large language model. The second knowledge base may include multiple first documents from the first knowledge base described above, as well as second documents corresponding to each newly generated document set.

[0079] To better understand the knowledge base construction method provided in this application embodiment, the overall architecture of the knowledge base construction scheme will be described below in conjunction with the embodiments and accompanying drawings.

[0080] Figure 2 This is a schematic diagram of the overall architecture of a knowledge base construction scheme provided in an embodiment of this application. Combined with... Figure 2 As shown, the original first knowledge base is denoted as K, which includes N first documents D, i.e., K = {D1, D2, ..., D...} N First, regarding the integration process of the first documents, taking a clustering model as an example, we can first encode the N first documents to obtain N encoding vectors {E1, E2, ..., E...}. N},in, Next, an N×d feature matrix is ​​constructed. Then, a clustering model F that accepts the N×d feature matrix is ​​used. C Clustering this feature matrix yields k clusters C, where the j-th cluster Cj is... j Including N j The first document D, which has semantic relationships, j ,Right now F C ({E1, E2, ..., E N})={C1,C2,...,C k This completes the integration process of the first document. For each of the k clusters, we can first construct the document generation instructions corresponding to each cluster, and then input the document generation instructions corresponding to each cluster into the large language model to obtain the second document corresponding to each cluster output by the large language model. Here, the j-th cluster C... j The corresponding second document can be denoted as The second documents corresponding to the k clusters can be denoted as follows: Finally, k second documents and N first documents (i.e., the first knowledge base K) can be jointly stored in the second knowledge base K', that is...

[0081] Furthermore, to expand the knowledge hierarchy of the second knowledge base, different clustering granularities can be set to generate multiple sets of second documents. Taking M clustering granularities as an example, M sets of second documents can be generated, each set containing k second documents, which can be denoted as follows according to the knowledge hierarchy from low to high: The aforementioned M groups of second documents can all be stored in the second knowledge base K', i.e., K' = K ∩ K a1 ∩K a2 ∩...∩K aM .

[0082] Based on the relevant content of steps S101-S104 above, in this embodiment, after obtaining multiple first documents included in the first knowledge base, the multiple first documents can be integrated to obtain multiple different document sets. Each document set in these multiple different document sets includes first documents with semantic relationships. Then, based on the large language model and document generation instructions, inductive reasoning is performed on the first documents included in each document set to obtain the second document corresponding to each document set. Finally, based on the multiple first documents and the second document corresponding to each document set, a second knowledge base is constructed. In this way, by integrating the first documents with semantic relationships in the first knowledge base to form multiple different document sets, cross-document reasoning ability can be introduced, which improves the ability of the large language model to perform knowledge retrieval with the help of the second knowledge base. Furthermore, by using the text generation and understanding capabilities of the large language model to perform inductive reasoning on the first documents in each document set to generate second documents, the knowledge levels of the second knowledge base can be expanded through the inductive reasoning ability of the large language model. This allows the second knowledge base to contain more information to meet the usage needs of the large language model, thereby improving the accuracy of the output results of the large language model.

[0083] Based on the knowledge base construction method provided in the preceding embodiments, this application embodiment can also provide a knowledge processing method that applies the knowledge base. The knowledge processing method is described below with reference to the embodiments and accompanying drawings.

[0084] Figure 3 A flowchart illustrating a knowledge processing method provided in an embodiment of this application. Figure 4 This is a schematic diagram of the overall architecture of a knowledge processing method provided in an embodiment of this application. (Combined with...) Figure 3 and Figure 4 As shown, the knowledge processing method provided in this application embodiment may include:

[0085] S301: In response to user input, retrieve knowledge related to the user input from a second knowledge base.

[0086] Here, user input information refers to the information that the user inputs into the large language model. The second knowledge base is obtained based on the knowledge base construction method provided in the above embodiments. Knowledge related to user input information refers to the answer information obtained by using the user input information as a question. Figure 4 As shown, when a user inputs the information "Which regions of country H does River A flow through?", RALM's search engine can use the user's input as a question and retrieve relevant knowledge from the second knowledge base, such as "River A is the largest river in country H...", "River A flows through the southwest, central and eastern parts of country H and flows into the sea in city S...", and "River A is known as the father river of the people of country H, and together with River B, it is called the mother river...".

[0087] S302: Generate text input instructions based on knowledge related to user input information and user input information.

[0088] In practical applications, once user input information and related knowledge are obtained, a prompt can be constructed. Figure 4 Taking the example of the prompt shown, the user input information can be used as the question, and the knowledge related to the user input information can be used as the known information to construct a prompt that includes both: "Given the following information: 1) River A is the largest river in country H... 2) River A flows through the southwest, central and eastern parts of country H and flows into the sea in city S... 3) River A is known as the father river of the people of country H, and together with River B, it is called the mother river... Please answer the question based on the known information: Which regions of country H does River A flow through?"

[0089] S303: Input the text input command into the large language model, and obtain the output information that matches the user input information through the large language model.

[0090] The output information matched with user input refers to the answer information output by the large language model, which uses the user input as a question and combines it with the knowledge matched to the user input. Figure 4 As shown, the constructed prompt can be input into the large language model, and the answer corresponding to the user input information can be obtained through the large language model, that is, the output information that matches the user input information: "River A flows through the southwest, central and eastern regions of country H".

[0091] Based on the relevant content of steps S301-S303 above, it can be seen that in this embodiment of the application, since the second knowledge base includes a second document with cross-document reasoning ability, and the knowledge level of the second document is higher than that of the first document, the second knowledge base can contain more information to meet the usage requirements of the large language model. The ability of the large language model to perform knowledge retrieval with the help of the second knowledge base can be improved, thereby improving the accuracy of the output results of the large language model.

[0092] In practical applications, the above knowledge processing method can be applied to robot dialogue scenarios. For example, in certain narrow knowledge domains, it can leverage external knowledge to enhance the accuracy of the output results, such as in legal consultation, financial analysis, or game dialogue scenarios. For ease of understanding, the following explanation uses a game dialogue scenario as an example, along with embodiments and accompanying drawings, to illustrate the knowledge processing method.

[0093] Figure 5This diagram illustrates an application scenario of the knowledge processing method provided in this embodiment. In a game dialogue scenario, the game program can provide an intelligent NPC (non-player character) who can chat with the player and answer questions about the game's world view and strategy. Specifically, a first knowledge base can be constructed by pre-collecting game world view introduction text and game strategy-related question-and-answer pairs as a first document. The world view introduction text includes, for example, game quest introductions and faction introductions, while the game strategy-related question-and-answer pairs include, for example, questions and answers related to game equipment parameters and event conditions. Then, a second document can be generated based on the first document using the knowledge base construction method provided in the above embodiment, and a second knowledge base can be constructed based on the first and second documents. In this way, the second knowledge base can serve as the knowledge source for the intelligent NPC's backend algorithm. Figure 5 In the provided game interface, in response to player A's question "Isn't Zhang San a big villain?", the intelligent NPC 001 can use the second knowledge base to retrieve the knowledge that "Uncle Zhang is the master of the Great Medicine Valley. He does not kill women, the elderly, or children. He is highly skilled in martial arts but has a righteous heart. He is a great hero in 001's heart" as the answer to the question and display it in the game interface.

[0094] Furthermore, in the game dialogue scenario provided in this application embodiment, the accuracy of question answering using the first knowledge base and the second knowledge base of the large language model can be tested through simulation experiments. Based on the same large language model, Table 2 provides an illustrative example of the effects achievable by the two knowledge bases:

[0095] Table 2

[0096]

[0097] As shown in Table 2, when the large language model uses the second knowledge base, the overall response accuracy improves, especially for questions requiring external knowledge support regarding game worldviews and strategies. For casual conversation content that doesn't require external knowledge support, the response accuracy is similar to that using the first knowledge base. Because the second knowledge base includes second documents with cross-document reasoning capabilities, and the knowledge level of the second documents is higher than that of the first documents, it can contain more information to meet the needs of the large language model. This enhances the large language model's ability to retrieve knowledge using the second knowledge base, thereby improving the accuracy of its output.

[0098] Based on the knowledge base construction method and knowledge processing method provided in the preceding embodiments, this application embodiment may also provide a knowledge base construction apparatus and a knowledge processing apparatus. The knowledge base construction apparatus and knowledge processing apparatus will be described below with reference to the embodiments and accompanying drawings.

[0099] Figure 6 This is a schematic diagram of a knowledge base construction apparatus provided in an embodiment of this application. (In conjunction with...) Figure 6 As shown, the knowledge base construction apparatus 600 provided in this application embodiment includes:

[0100] The document acquisition module 601 is used to acquire multiple first documents included in the first knowledge base;

[0101] The document integration module 602 is used to integrate the plurality of first documents to obtain a plurality of different document sets, wherein each document set includes first documents that have semantic relationships.

[0102] The inductive reasoning module 603 is used to perform inductive reasoning on the first document included in each document set based on the large language model and document generation instructions to obtain the second document corresponding to each document set.

[0103] The knowledge base construction module 604 is used to construct a second knowledge base based on the plurality of first documents and the second documents corresponding to each document set.

[0104] Optionally, the document integration module 602 includes:

[0105] The first document encoding module is used to encode the plurality of first documents to obtain encoding vectors corresponding to the plurality of first documents respectively;

[0106] A matrix construction module is used to construct a feature matrix for the plurality of first documents based on the encoding vectors corresponding to the plurality of first documents respectively;

[0107] The clustering module is used to cluster the feature matrix based on the clustering model to obtain multiple different clusters as the multiple different document sets.

[0108] Optionally, the clustering module includes:

[0109] The first clustering submodule is used to cluster the feature matrix at multiple clustering granularities based on the clustering model, so as to obtain multiple clusters under each of the multiple clustering granularities.

[0110] The second clustering submodule is used to treat multiple clusters at each clustering granularity as the multiple different document sets.

[0111] Optionally, the document integration module 602 includes:

[0112] The second document encoding module is used to encode the plurality of first documents to obtain encoding vectors corresponding to the plurality of first documents respectively;

[0113] The classification module is used to classify the encoding vectors corresponding to the multiple first documents based on the classification model, and obtain document sets corresponding to multiple categories as the multiple different document sets.

[0114] Optionally, the classification module includes:

[0115] The first classification submodule is used to classify the encoding vectors corresponding to the multiple first documents respectively with multiple classification granularities based on the classification model, so as to obtain the document set corresponding to multiple categories under each of the multiple classification granularities;

[0116] The second classification submodule is used to take the document sets corresponding to the multiple categories under each classification granularity as the multiple different document sets.

[0117] Optionally, the document generation instructions are obtained through the following module:

[0118] The task information acquisition module is used to acquire task information of the large language model, and the task information is used to instruct the large language model to perform inductive and / or inference operations.

[0119] The second instruction generation module is used to generate document generation instructions corresponding to each document set based on the first document included in each document set and the task information.

[0120] Figure 7 This is a schematic diagram of the structure of a knowledge processing device provided in an embodiment of this application. (In conjunction with...) Figure 7 As shown, the knowledge processing apparatus 700 provided in this application embodiment includes:

[0121] The knowledge acquisition module 701 is used to acquire knowledge related to the user input information from a second knowledge base in response to user input information, wherein the second knowledge base is obtained based on the knowledge base construction method according to any one of claims 1 to 6;

[0122] The first instruction generation module 702 is used to generate text input instructions based on knowledge related to the user input information and the user input information.

[0123] The output information acquisition module 703 is used to input the text input instruction into the large language model and obtain output information that matches the user input information through the large language model.

[0124] The following sections describe the structure of the control equipment used to implement the above knowledge base construction or knowledge processing methods, in both server and terminal device formats.

[0125] Figure 8 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 900 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 922 (e.g., one or more processors) and memory 932, and one or more storage media 930 (e.g., one or more mass storage devices) for storing application programs 942 or data 944. The memory 932 and storage media 930 can be temporary or persistent storage. The program stored in the storage media 930 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the server. Furthermore, the CPU 922 may be configured to communicate with the storage media 930 and execute the series of instruction operations in the storage media 930 on the server 900.

[0126] Server 900 may also include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems 941, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0127] The steps performed by the server in the above embodiments can be based on this Figure 8 The server structure shown.

[0128] The CPU 922 is used in the following steps:

[0129] Retrieve multiple first documents included in the first knowledge base;

[0130] The multiple first documents are integrated to obtain multiple different document sets, and each document set includes first documents that have semantic relationships.

[0131] Based on the large language model and document generation instructions, inductive reasoning is performed on the first document included in each document set to obtain the second document corresponding to each document set;

[0132] A second knowledge base is constructed based on the plurality of first documents and the second documents corresponding to each document set;

[0133] or,

[0134] In response to user input information, knowledge related to the user input information is retrieved from a second knowledge base, which is obtained based on the above-described knowledge base construction method;

[0135] Text input instructions are generated based on knowledge related to the user input information and the user input information itself;

[0136] The text input command is input into a large language model, and the output information matching the user input information is obtained through the large language model.

[0137] This application also provides another control device, such as... Figure 9 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal can be any terminal device including mobile phones, tablets, personal digital assistants (PDAs), point-of-sale (POS) terminals, in-vehicle computers, etc. Taking a mobile phone as an example:

[0138] Figure 9 This is a block diagram illustrating a portion of the structure of a mobile phone related to the terminal provided in the embodiments of this application. (Reference) Figure 9 The mobile phone includes: a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090, etc. Those skilled in the art will understand that... Figure 9 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0139] The following is combined with Figure 9 A detailed introduction to each component of a mobile phone:

[0140] The RF circuit 1010 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 1080; additionally, it transmits uplink data to the base station. Typically, the RF circuit 1010 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the RF circuit 1010 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).

[0141] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0142] The input unit 1030 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1031), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 1080, and can also receive and execute commands sent by the processor 1080. In addition, the touch panel 1031 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1031, the input unit 1030 may also include other input devices 1032. Specifically, other input devices 1032 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0143] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1040 may include a display panel 1041, which may optionally be configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel 1041. Further, a touch panel 1031 may cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it transmits the information to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 9 In this embodiment, the touch panel 1031 and the display panel 1041 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.

[0144] The mobile phone may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1041 according to the ambient light level, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0145] The audio circuit 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and the mobile phone. The audio circuit 1060 converts the received audio data into electrical signals and transmits them to the speaker 1061, where the speaker 1061 converts them into sound signals for output. On the other hand, the microphone 1062 converts the collected sound signals into electrical signals, which are then received by the audio circuit 1060, converted into audio data, and then processed by the processor 1080 before being transmitted via the RF circuit 1010 to, for example, another mobile phone, or the audio data can be output to the memory 1020 for further processing.

[0146] WiFi is a short-range wireless transmission technology. Through the WiFi module 1070, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 9 The WiFi module 1070 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.

[0147] The processor 1080 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 1020 and calls data stored in the memory 1020 to perform various functions and process data, thereby collecting overall data and information from the phone. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1080.

[0148] The mobile phone also includes a power supply 1090 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 1080 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0149] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0150] In this embodiment of the application, the processor 1080 included in the terminal also has the following functions:

[0151] Retrieve multiple first documents included in the first knowledge base;

[0152] The multiple first documents are integrated to obtain multiple different document sets, and each document set includes first documents that have semantic relationships.

[0153] Based on the large language model and document generation instructions, inductive reasoning is performed on the first document included in each document set to obtain the second document corresponding to each document set;

[0154] A second knowledge base is constructed based on the plurality of first documents and the second documents corresponding to each document set;

[0155] or,

[0156] In response to user input information, knowledge related to the user input information is retrieved from a second knowledge base, which is obtained based on the above-described knowledge base construction method;

[0157] Text input instructions are generated based on knowledge related to the user input information and the user input information itself;

[0158] The text input command is input into a large language model, and the output information matching the user input information is obtained through the large language model.

[0159] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0160] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0162] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0164] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0165] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for constructing a knowledge base, characterized in that, include: Retrieve multiple first documents included in the first knowledge base; The multiple first documents are integrated to obtain multiple different document sets, and each document set includes first documents that have semantic relationships. Based on the large language model and document generation instructions, inductive reasoning is performed on the first document included in each document set to obtain the second document corresponding to each document set; A second knowledge base is constructed based on the plurality of first documents and the second documents corresponding to each document set.

2. The method for constructing a knowledge base according to claim 1, characterized in that, The process of integrating the multiple first documents to obtain multiple different document sets includes: Encode the plurality of first documents to obtain encoding vectors corresponding to the plurality of first documents respectively; Based on the encoding vectors corresponding to the multiple first documents, a feature matrix of the multiple first documents is constructed; The feature matrix is ​​clustered based on a clustering model to obtain multiple different clusters, which serve as the multiple different document sets.

3. The method for constructing a knowledge base according to claim 2, characterized in that, The clustering of the feature matrix based on the clustering model yields multiple different clusters as the multiple different document sets, including: Based on the clustering model, the feature matrix is ​​clustered at multiple clustering granularities to obtain multiple clusters under each of the multiple clustering granularities; Each cluster at each clustering granularity is used as a different document set.

4. The method for constructing a knowledge base according to claim 1, characterized in that, The process of integrating the multiple first documents to obtain multiple different document sets includes: Encode the plurality of first documents to obtain encoding vectors corresponding to the plurality of first documents respectively; The encoding vectors corresponding to the multiple first documents are classified based on the classification model to obtain multiple document sets corresponding to multiple categories, which are the multiple different document sets.

5. The method for constructing a knowledge base according to claim 4, characterized in that, The step of classifying the encoding vectors corresponding to the multiple first documents based on a classification model to obtain multiple document sets corresponding to multiple categories as the multiple different document sets includes: Based on the classification model, the encoding vectors corresponding to the multiple first documents are classified at multiple classification granularities to obtain document sets corresponding to multiple categories under each of the multiple classification granularities; The document sets corresponding to the multiple categories under each classification granularity are used as the multiple different document sets.

6. The method for constructing a knowledge base according to any one of claims 1 to 5, characterized in that, The document generation instructions are obtained through the following steps: Obtain task information from a large language model, wherein the task information is used to instruct the large language model to perform inductive and / or inferential operations; Based on the first document included in each document set and the task information, a document generation instruction corresponding to each document set is generated.

7. A knowledge processing method, characterized in that, include: In response to user input information, knowledge related to the user input information is obtained from a second knowledge base, wherein the second knowledge base is obtained based on the knowledge base construction method according to any one of claims 1 to 6; Text input instructions are generated based on knowledge related to the user input information and the user input information itself; The text input command is input into a large language model, and the output information matching the user input information is obtained through the large language model.

8. A knowledge base construction apparatus, characterized in that, include: The document acquisition module is used to acquire multiple first documents included in the first knowledge base; The document integration module is used to integrate the multiple first documents to obtain multiple different document sets, and each document set includes first documents that have semantic relationships. The inductive reasoning module is used to perform inductive reasoning on the first document included in each document set based on the large language model and document generation instructions to obtain the second document corresponding to each document set. The knowledge base construction module is used to construct a second knowledge base based on the plurality of first documents and the second documents corresponding to each document set.

9. A knowledge processing device, characterized in that, include: A knowledge acquisition module is used to retrieve knowledge related to the user input information from a second knowledge base in response to user input information, wherein the second knowledge base is obtained based on the knowledge base construction method according to any one of claims 1 to 6; The first instruction generation module is used to generate text input instructions based on knowledge related to the user input information and the user input information; The output information acquisition module is used to input the text input command into the large language model and obtain output information that matches the user input information through the large language model.

10. An electronic device, characterized in that, The device includes a processor and a memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute, according to instructions in the computer program, the steps of the knowledge base construction method of any one of claims 1 to 6, or the steps of the knowledge processing method of claim 7.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed by a terminal device, implements the steps of the knowledge base construction method according to any one of claims 1 to 6, or the steps of the knowledge processing method according to claim 7.

12. A computer program product, characterized in that, It includes a computer program that, when executed by a terminal device, implements the steps of the knowledge base construction method according to any one of claims 1 to 6, or the steps of the knowledge processing method according to claim 7.

Citation Information

Cited By

  • Key point information arrangement method for normative text, electronic equipment and medium

    CN121480485A

  • A method for summarizing key information of a normative text, an electronic device, and a medium

    CN121480485B