Domain knowledge base construction method and device based on knowledge point expansion
Through the method based on knowledge point expansion, circular expansion and cluster query collection, the shortcomings of knowledge point identification and correlation mining in the construction of domain knowledge bases in the existing technology are solved, and efficient update of knowledge bases and meeting diversified needs are achieved.
Patent Information
- Application Number
- CN202411860394.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-13
AI Technical Summary
When building a domain knowledge base, the existing technology lacks the accuracy of knowledge point recognition, the depth of correlation mining, and the dynamic update ability of the knowledge base, which is difficult to meet the diversified needs in different application scenarios.
Using a method based on knowledge point expansion, the initial query collection is generated through a custom seed query, the relevant knowledge point data is searched and obtained, and the unrelated data is filtered using a binary classification model, loop expansion and cluster query collections are expanded until the construction convergence condition is reached.
It improves the accuracy of knowledge point recognition and the depth of correlation mining, enhances the coverage and update efficiency of the knowledge base, and ensures the timeliness and authority of the knowledge base content.
Smart Images

Figure CN119988539A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge base construction, and more specifically, to a method and device for constructing a domain knowledge base based on knowledge point expansion. Background Art
[0002] At present, the methods for building domain knowledge bases mainly rely on two approaches: one is the traditional manual sorting method, that is, domain experts sort out and enter knowledge points and their relationships one by one based on their professional knowledge and experience; the other is the automated construction method based on natural language processing (NLP) technology, which automatically extracts knowledge points from massive text data through text mining, information extraction and other technologies, and tries to establish relationships between them.
[0003] Specifically, although the traditional manual sorting method can ensure the accuracy and authority of the knowledge base, it is inefficient, costly, and difficult to adapt to the rapidly changing knowledge environment. Although the automated construction method based on NLP technology has greatly improved the construction efficiency, it still has shortcomings in the accuracy of knowledge point identification, the depth of association mining, and the dynamic update ability of the knowledge base. For example, the automated construction methods in the prior art often find it difficult to accurately identify knowledge points in complex contexts, and the mining of implicit relationships between knowledge points is not deep enough. At the same time, there is a lack of effective mechanisms to continuously track and introduce new knowledge, resulting in a lag in the update of knowledge base content. In addition, the existing domain knowledge base construction methods also have the problem of a single knowledge representation method, which is difficult to meet the diverse needs in different application scenarios. Summary of the invention
[0004] In order to solve the shortcomings of the prior art methods for constructing a domain knowledge base in terms of the accuracy of knowledge point recognition, the depth of association mining and the dynamic update capability of the knowledge base, as well as the technical problem of being difficult to meet the diverse needs in different application scenarios, the present invention provides a method and device for constructing a domain knowledge base based on knowledge point expansion.
[0005] According to one aspect of the present invention, the present invention provides a method for constructing a domain knowledge base based on knowledge point expansion, comprising:
[0006] Step 1: Generate an initial query set based on several customized seed queries;
[0007] Step 2: Search each seed query in the initial query set, obtain basic web page information and web page document content related to each seed query, and generate initial knowledge point data of the initial query set;
[0008] Step 3: Based on a custom binary classification model, the initial knowledge point data is classified, knowledge point data irrelevant to the professional field in which the domain knowledge base to be constructed is located is filtered out, and the relevant knowledge point data is stored as valid knowledge point data in the domain knowledge base to be constructed;
[0009] Step 4: Determine the knowledge base construction result according to the valid knowledge point data and the customized knowledge base construction convergence condition, wherein the construction result includes the end of construction and the continuation of construction. When the construction result is the continuation of construction, go to step 5; when the construction result is the end of construction, go to step 6;
[0010] Step 5: Expand and cluster the initial query set based on the valid knowledge point data to generate an updated query set, and set the updated query set as the initial query set, and go to step 2;
[0011] Step 6: Generate a domain knowledge base based on all the knowledge point data stored in the domain knowledge base to be constructed.
[0012] According to another aspect of the present invention, the present invention provides a device for constructing a domain knowledge base based on knowledge point expansion, the device comprising:
[0013] The initial collection module is used to generate an initial query set based on several customized seed queries before the update collection module is started;
[0014] A knowledge point acquisition module, used to search each seed query in the initial query set, obtain basic web page information and web page document content related to each seed query, and generate initial knowledge point data of the initial query set;
[0015] A knowledge point pruning module is used to classify the initial knowledge point data based on a custom binary classification model, filter out knowledge point data that is not relevant to the professional field in which the domain knowledge base to be constructed is located, and store the relevant knowledge point data as valid knowledge point data in the domain knowledge base to be constructed;
[0016] A convergence judgment module is used to determine the knowledge base construction result according to the valid knowledge point data and the self-defined knowledge base construction convergence condition, wherein the construction result includes the end of construction and the continuation of construction. When the construction result is the continuation of construction, it is transferred to the update set module, and when the construction result is the end of construction, it is transferred to the result output module;
[0017] An update set module, used for expanding and clustering the initial query set based on the valid knowledge point data, generating an updated query set, and setting the updated query set as the initial query set;
[0018] The result output module is used to generate a domain knowledge base based on all the knowledge point data stored in the domain knowledge base to be constructed.
[0019] According to another aspect of the present invention, the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above aspects of the present invention.
[0020] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of the above aspects of the present invention.
[0021] The method and device for constructing a domain knowledge base based on knowledge point expansion described in the present invention include: generating an initial query set based on a number of customized seed queries; searching each seed query in the initial query set to obtain basic web page information and web page document content related to each seed query, and generating initial knowledge point data of the initial query set; classifying the initial knowledge point data based on a customized binary classification model, filtering out knowledge point data that are not related to the professional field in which the domain knowledge base to be constructed is located, and storing the related knowledge point data as valid knowledge point data in the domain knowledge base to be constructed; determining the knowledge base construction result based on the valid knowledge point data and the customized knowledge base construction convergence condition, and when the construction result is to continue construction, expanding and clustering the initial query set based on the valid knowledge point data to generate an updated query set, and making the updated query set the initial query set and then repeating the iteration; when the construction result is the end of construction, generating a domain knowledge base based on all the knowledge point data stored in the domain knowledge base to be constructed. The method and device not only improve the accuracy of knowledge point recognition and the depth of association mining, but also improve the coverage and updating efficiency of the knowledge base by adopting query search and expanded clustering, knowledge point data classification model and combining with search engine resources, thereby ensuring the timeliness and authority of the knowledge base content. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:
[0023] Figure 1 A flowchart of a method for constructing a domain knowledge base based on knowledge point expansion according to a preferred embodiment of the present invention;
[0024] Figure 2 A schematic diagram of the structure of a device for constructing a domain knowledge base based on knowledge point expansion according to a preferred embodiment of the present invention;
[0025] Figure 3 Schematic diagram of the structure of an electronic device according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0026] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.
[0027] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.
[0028] Exemplary Methods
[0029] Figure 1 FIG. 1 is a flowchart of a method for constructing a domain knowledge base based on knowledge point expansion according to a preferred embodiment of the present invention. Figure 1 As shown, the method for constructing a domain knowledge base based on knowledge point expansion described in this preferred embodiment starts from step 101.
[0030] In step 101, an initial query set is generated according to a number of customized seed queries.
[0031] Preferably, the generating of the initial query set according to the plurality of customized seed queries includes:
[0032] Determine the professional field in which the domain knowledge base is to be built;
[0033] At least one of the following methods is used to generate multiple seed queries, where:
[0034] Extract frequently asked questions from authoritative websites in the professional field;
[0035] Extract the titles of laws and regulations in the professional field;
[0036] Using a large model, questions are generated based on documents in the professional field;
[0037] Generate an initial query set based on multiple generated queries.
[0038] In this preferred implementation, taking the finance and taxation field as an example, the seed query can be obtained through common FAQs or titles of key finance and taxation regulations on authoritative finance and taxation sites (such as the official website of the State Administration of Taxation), or by generating questions from documents such as accounting standards through a large model.
[0039] In step 102, each seed query in the initial query set is searched to obtain basic web page information and web page document content related to each seed query, and generate initial knowledge point data of the initial query set.
[0040] Preferably, searching each seed query in the initial query set, acquiring basic web page information and web page document content related to each seed query, and generating initial knowledge point data of the initial query set includes:
[0041] Using a search engine to search each seed query in the initial query set, and obtaining the top N pages most relevant to each seed query from the search results;
[0042] Parsing the N pages to obtain basic webpage information of the corresponding webpages, wherein the basic webpage information includes webpage URL, title and summary;
[0043] Using crawlers and / or web page parsing technology to obtain web page document contents corresponding to the URLs of N web pages;
[0044] The webpage basic information and webpage document contents of the N pages are matched one by one to generate the initial knowledge point data of the initial query set.
[0045] In this preferred embodiment, there is no restriction on the search engine, and Baidu, Google or Bing can be used. Through these search engines, each seed query is accurately searched to obtain multiple most relevant pages, such as top 10. Then, the URL, title, summary and other information of the web page are parsed from the page, and the web document content corresponding to all URLs is obtained through crawlers and web page parsing technology, so that 10 times the initial knowledge point data of the initial query set of the search is obtained.
[0046] In step 103, based on a custom binary classification model, the initial knowledge point data is classified, knowledge point data irrelevant to the professional field where the domain knowledge base to be constructed is located is filtered out, and relevant knowledge point data is stored as valid knowledge point data in the domain knowledge base to be constructed.
[0047] Preferably, before classifying the knowledge point data of the initial query set based on the custom binary classification model, the method further includes generating a binary classification model, specifically:
[0048] Obtain and annotate multiple web page samples that are related to and unrelated to the professional field in which the domain knowledge base to be constructed is located, and generate a web page sample set;
[0049] Establish an initial domain classification model based on the RoBerta base model;
[0050] The web page sample set is used to train the initial domain classification model to generate the binary classification model.
[0051] In this preferred embodiment, in order to distinguish whether the acquired initial knowledge point data is related to the professional field in which the domain knowledge base to be constructed is located, the initial knowledge point data needs to be pruned. Therefore, a domain classification model needs to be trained in advance to ensure that the web document content stored in the domain knowledge base is highly relevant to the construction goal of the domain knowledge base.
[0052] In step 104, the knowledge base construction result is determined based on the valid knowledge point data and the customized knowledge base construction convergence condition, wherein the construction result includes the end of construction and continued construction. When the construction result is continued construction, go to step 105; when the construction result is the end of construction, go to step 106.
[0053] Preferably, the knowledge base construction result is determined according to the valid knowledge point data and the self-defined knowledge base construction convergence condition, wherein the self-defined knowledge base construction convergence condition is:
[0054] When the repetition rate between the valid knowledge point data and the knowledge point data already stored in the domain knowledge base to be constructed reaches P%, the knowledge base construction result is determined to be completed, otherwise, the construction continues.
[0055] In this preferred embodiment, P takes a value of 60, indicating that when more than 60% of the acquired valid knowledge point data are already in the constructed domain knowledge base, it indicates that the growth of the knowledge point data in the knowledge base has stabilized and the repeated iteration can be terminated.
[0056] In step 105 , the initial query set is expanded and clustered based on the valid knowledge point data to generate an updated query set, and the updated query set is set as the initial query set, and the process goes to step 102 .
[0057] Preferably, the step of expanding and clustering the initial query set based on the valid knowledge point data to generate an updated query set includes:
[0058] According to the webpage document content in the valid knowledge point data, a number of new seed queries are generated by calling the large model interface;
[0059] Using the web page title in the valid knowledge point data as a new seed query;
[0060] Generate a candidate query set according to the new seed query and the seed query in the initial query set;
[0061] Obtaining a 1024-dimensional semantic vector of each seed query in the candidate query set based on the open source text-to-semantic vector model bge-m3;
[0062] Clustering the seed queries in the candidate query set according to the semantic vector by block matrix operation;
[0063] An updated query set is generated according to the result of clustering the seed queries in the candidate query set.
[0064] Preferably, clustering the seed queries in the candidate query set by block matrix operation according to the semantic vector comprises:
[0065] Calculate the inner product between the semantic vectors of any two seed queries;
[0066] When the inner product of two seed queries is greater than the user-defined inner product threshold, the two seed queries are determined to be a synonymous question pair;
[0067] Aggregating the synonymous question pairs in the candidate query set by co-occurrence terms to generate a number of seed query clusters, wherein the co-occurrence terms refer to seed queries that appear in both synonymous question pairs;
[0068] When the same seed query appears in any two seed query clusters, the same seed query is removed from both seed query clusters.
[0069] Preferably, the calculating the inner product between the semantic vectors of any two seed queries includes:
[0070] Divide each seed query semantic vector in the candidate query set into I matrix blocks;
[0071] Compute the matrix inner product between I matrix blocks of any two seed queries.
[0072] In this preferred implementation, there may be some queries in the candidate query set generated after expanding the seed query that are synonymous with each other although they have different literals. Therefore, it is necessary to cluster the seed queries of the candidate query set and group the synonymous queries together to avoid repeated mining of synonymous question-answer pairs.
[0073] When a seed query has too many semantic vectors, for a large set of seed queries, it will take a very long time to determine synonymous question pairs by simply calculating the inner product of the semantic vectors between the two pairs. Although the calculation efficiency can be improved by using numpy matrix operations, if the full amount of data is loaded into one matrix, the operation will fail due to the explosion of memory overhead. At the same time, calculating two N*1024 matrices will also waste half of the computing power (the inner product remains unchanged when the order of the question pairs is exchanged). In order to solve the above problems, the present invention proposes a block calculation method. After dividing the semantic vectors of all seed queries into N matrix blocks, the matrix inner product between the blocks is calculated, and the calculation of the inner product between all the questions can be completed. This method can get rid of the size limit of the set of questions to be clustered, and the calculation efficiency is more than 100 times higher than that of direct pairwise calculation.
[0074] This preferred implementation achieves deduplication by performing seed query expansion on the acquired valid knowledge point data and then adopting a clustering algorithm to identify and merge semantically identical queries based on the semantic similarity between the seed queries in the expanded candidate query set. This helps to reduce redundant seed query information in the candidate query set and improve the quality of searching for relevant knowledge point data, thereby accelerating the efficiency of constructing the domain knowledge base.
[0075] In step 106, a domain knowledge base is generated based on all the knowledge point data stored in the domain knowledge base to be constructed.
[0076] In this preferred embodiment, when the knowledge base construction convergence conditions are met, the domain knowledge base is constructed. However, as new domain-related knowledge point data emerges, the binary classification model can be retrained, and then knowledge point data can be collected based on this method to ensure changes and updates in the content of the domain knowledge base.
[0077] The method for constructing a domain knowledge base based on knowledge point expansion described in this preferred implementation ensures comprehensive coverage of the professional domain knowledge base by integrating multi-source authoritative data in the professional domain. At the same time, by introducing real-time monitoring and automatic update mechanisms, it greatly reduces human intervention, improves the efficiency and intelligence level of knowledge base construction, and enables the knowledge base to quickly respond to the latest developments in the professional field, ensuring the timeliness of the data and providing users with the latest and most accurate professional field information.
[0078] Exemplary Devices
[0079] Figure 2 FIG. 1 is a schematic diagram of a structure of a device for constructing a domain knowledge base based on knowledge point expansion according to a preferred embodiment of the present invention. Figure 2 As shown, the domain knowledge base construction device 200 based on knowledge point expansion described in this preferred embodiment includes:
[0080] The initial set module 201 is used to generate an initial query set according to a number of self-defined seed queries before the update set module 205 is started;
[0081] The knowledge point acquisition module 202 is used to search each seed query in the initial query set, obtain the basic web page information and web page document content related to each seed query, and generate initial knowledge point data of the initial query set;
[0082] The knowledge point pruning module 203 is used to classify the initial knowledge point data based on a custom binary classification model, filter out knowledge point data that is irrelevant to the professional field in which the domain knowledge base to be constructed is located, and store the relevant knowledge point data as valid knowledge point data in the domain knowledge base to be constructed;
[0083] The convergence judgment module 204 is used to determine the knowledge base construction result according to the valid knowledge point data and the self-defined knowledge base construction convergence condition, wherein the construction result includes the construction end and the construction continuation. When the construction result is the construction continuation, the module goes to the update set module 205; when the construction result is the construction end, the module goes to the result output module 206;
[0084] An update set module 205, configured to expand and cluster the initial query set based on the valid knowledge point data, generate an updated query set, and set the updated query set as the initial query set;
[0085] The result output module 206 is used to generate a domain knowledge base according to all the knowledge point data stored in the domain knowledge base to be constructed.
[0086] Preferably, the initial set module 201 generates an initial query set according to a plurality of self-defined seed queries, including:
[0087] Determine the professional field in which the domain knowledge base is to be built;
[0088] At least one of the following methods is used to generate multiple seed queries, where:
[0089] Extract frequently asked questions from authoritative websites in the professional field;
[0090] Extract the titles of laws and regulations in the professional field;
[0091] Using a large model, questions are generated based on documents in the professional field;
[0092] Generate an initial query set based on multiple generated queries.
[0093] Preferably, the knowledge point acquisition module searches each seed query in the initial query set, obtains web page basic information and web page document content related to each seed query, and generates initial knowledge point data of the initial query set, including:
[0094] Using a search engine to search each seed query in the initial query set, and obtaining the top N pages most relevant to each seed query from the search results;
[0095] Parsing the N pages to obtain basic webpage information of the corresponding webpages, wherein the basic webpage information includes webpage URL, title and summary;
[0096] Using crawlers and / or web page parsing technology to obtain web page document contents corresponding to the URLs of N web pages;
[0097] The webpage basic information and webpage document contents of the N pages are matched one by one to generate the initial knowledge point data of the initial query set.
[0098] Preferably, the knowledge point pruning module 203 is based on a custom binary classification model, and before classifying the initial knowledge point data, it also includes generating a binary classification model, specifically:
[0099] Obtain and annotate multiple web page samples that are related to and unrelated to the professional field in which the domain knowledge base to be constructed is located, and generate a web page sample set;
[0100] Establish an initial domain classification model based on the RoBerta base model;
[0101] The web page sample set is used to train the initial domain classification model to generate the binary classification model.
[0102] Preferably, the convergence judgment module 204 determines the knowledge base construction result according to the valid knowledge point data and the self-defined knowledge base construction convergence condition as follows:
[0103] When the repetition rate between the valid knowledge point data and the knowledge point data already stored in the domain knowledge base to be constructed reaches P%, the knowledge base construction result is determined to be completed, otherwise, the construction continues.
[0104] Preferably, the update set module 205 expands and clusters the initial query set based on the valid knowledge point data to generate an updated query set, including:
[0105] According to the webpage document content in the valid knowledge point data, a number of new seed queries are generated by calling the large model interface;
[0106] Using the web page title in the valid knowledge point data as a new seed query;
[0107] Generate a candidate query set according to the new seed query and the seed query in the initial query set;
[0108] Obtaining a 1024-dimensional semantic vector of each seed query in the candidate query set based on the open source text-to-semantic vector model bge-m3;
[0109] Clustering the seed queries in the candidate query set according to the semantic vector by block matrix operation;
[0110] An updated query set is generated according to the result of clustering the seed queries in the candidate query set.
[0111] Preferably, the update set module 205 clusters the seed queries in the candidate query set according to the semantic vector through block matrix operation, including:
[0112] Calculate the inner product between the semantic vectors of any two seed queries;
[0113] When the inner product of two seed queries is greater than the user-defined inner product threshold, the two seed queries are determined to be a synonymous question pair;
[0114] Aggregating the synonymous question pairs in the candidate query set by co-occurrence terms to generate a number of seed query clusters, wherein the co-occurrence terms refer to seed queries that appear in both synonymous question pairs;
[0115] When the same seed query appears in any two seed query clusters, the same seed query is removed from both seed query clusters.
[0116] Preferably, the update set module 205 calculates the inner product between the semantic vectors of any two seed queries, including:
[0117] Divide each seed query semantic vector in the candidate query set into I matrix blocks;
[0118] Compute the matrix inner product between I matrix blocks of any two seed queries.
[0119] The domain knowledge base construction device based on knowledge point expansion described in this preferred embodiment has the same steps of constructing and expanding the domain knowledge base as the domain knowledge base construction method based on knowledge point expansion, and the technical effects achieved are also the same, which will not be repeated here.
[0120] Exemplary Electronic Devices
[0121] Figure 3 The electronic device according to the preferred embodiment of the present invention is a schematic diagram of the structure of the electronic device. The electronic device can be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device can communicate with the first device and the second device to receive the collected input signal from them. Figure 3 FIG. 1 is a block diagram of an electronic device according to an embodiment of the present disclosure. Figure 3 As shown, the electronic device includes one or more processors 301 and a memory 302 .
[0122] The processor 301 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0123] The memory 302 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 301 may run the program instructions to implement the energy consumption anomaly diagnosis method based on the enterprise energy consumption space of each disclosed embodiment described above and / or other desired functions. In one example, the electronic device may also include: an input device 303 and an output device 304, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0124] In addition, the input device 303 may also include, for example, a keyboard, a mouse, and the like.
[0125] The output device 304 can output various information to the outside, and can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto.
[0126] Of course, to simplify, Figure 3 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device may further include any other appropriate components.
[0127] Exemplary computer program products and computer-readable storage media
[0128] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method for constructing a domain knowledge base based on knowledge point expansion according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0129] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0130] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, causes the processor to execute the steps of the method for constructing a domain knowledge base based on knowledge point expansion according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0131] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0132] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0133] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0134] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0135] The apparatus and method of the present disclosure may be implemented in many ways. For example, the apparatus and method of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0136] It should also be noted that in the apparatus, equipment and method of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present disclosure. The above description of the disclosed aspects is provided to enable any technician in the field to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown here, but to the widest scope consistent with the principles and novel features disclosed herein.
[0137] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A method for constructing a domain knowledge base based on knowledge point expansion, characterized in that: The method comprises: Step 1: Generate an initial query set based on several customized seed queries; Step 2: Search each seed query in the initial query set, obtain basic web page information and web page document content related to each seed query, and generate initial knowledge point data of the initial query set; Step 3: Based on a custom binary classification model, the initial knowledge point data is classified, knowledge point data irrelevant to the professional field in which the domain knowledge base to be constructed is located is filtered out, and the relevant knowledge point data is stored as valid knowledge point data in the domain knowledge base to be constructed; Step 4: Determine the knowledge base construction result according to the valid knowledge point data and the customized knowledge base construction convergence condition, wherein the construction result includes the end of construction and the continuation of construction. When the construction result is the continuation of construction, go to step 5; when the construction result is the end of construction, go to step 6; Step 5: Expand and cluster the initial query set based on the valid knowledge point data to generate an updated query set, and set the updated query set as the initial query set, and go to step 2; Step 6: Generate a domain knowledge base based on all the knowledge point data stored in the domain knowledge base to be constructed.
2. The method according to claim 1, characterized in that The initial query set is generated based on several customized seed queries, including: Determine the professional field in which the domain knowledge base is to be built; At least one of the following methods is used to generate multiple seed queries, where: Extract frequently asked questions from authoritative websites in the professional field; Extract the titles of laws and regulations in the professional field; Using a large model, questions are generated based on documents in the professional field; Generate an initial query set based on multiple generated queries.
3. The method according to claim 1, characterized in that The step of searching each seed query in the initial query set, acquiring basic web page information and web page document content related to each seed query, and generating initial knowledge point data of the initial query set includes: Using a search engine to search each seed query in the initial query set, and obtaining the top N pages most relevant to each seed query from the search results; Parsing the N pages to obtain basic webpage information of the corresponding webpages, wherein the basic webpage information includes webpage URL, title and summary; Using crawlers and / or web page parsing technology to obtain web page document contents corresponding to the URLs of N web pages; The webpage basic information and webpage document contents of the N pages are matched one by one to generate the initial knowledge point data of the initial query set.
4. The method according to claim 1, characterized in that: The method based on the customized binary classification model further includes generating a binary classification model before classifying the initial knowledge point data. Specifically: Obtain and annotate multiple web page samples that are related to and unrelated to the professional field in which the domain knowledge base to be constructed is located, and generate a web page sample set; Establish an initial domain classification model based on the RoBerta base model; The web page sample set is used to train the initial domain classification model to generate the binary classification model.
5. The method according to claim 3, characterized in that: The knowledge base construction result is determined based on the valid knowledge point data and the self-defined knowledge base construction convergence condition as follows: When the repetition rate between the valid knowledge point data and the knowledge point data already stored in the domain knowledge base to be constructed reaches P%, the knowledge base construction result is determined to be completed, otherwise, the construction continues.
6. The method according to claim 3, characterized in that The expanding and clustering the initial query set based on the valid knowledge point data to generate an updated query set includes: According to the webpage document content in the valid knowledge point data, a number of new seed queries are generated by calling the large model interface; Using the web page title in the valid knowledge point data as a new seed query; Generate a candidate query set according to the new seed query and the seed query in the initial query set; Obtaining a 1024-dimensional semantic vector of each seed query in the candidate query set based on the open source text-to-semantic vector model bge-m3; Clustering the seed queries in the candidate query set according to the semantic vector by block matrix operation; An updated query set is generated according to the result of clustering the seed queries in the candidate query set.
7. The method according to claim 6, characterized in that The clustering of the seed queries in the candidate query set by block matrix operation according to the semantic vector includes: Calculate the inner product between the semantic vectors of any two seed queries; When the inner product of two seed queries is greater than the user-defined inner product threshold, the two seed queries are determined to be a synonymous question pair; Aggregating the synonymous question pairs in the candidate query set by co-occurrence terms to generate a number of seed query clusters, wherein the co-occurrence terms refer to seed queries that appear in both synonymous question pairs; When the same seed query appears in any two seed query clusters, the same seed query is removed from both seed query clusters.
8. The method according to claim 7, characterized in that The step of calculating the inner product between the semantic vectors of any two seed queries includes: Divide each seed query semantic vector in the candidate query set into I matrix blocks; Compute the matrix inner product between I matrix blocks of any two seed queries.
9. A device for constructing a domain knowledge base based on knowledge point expansion, characterized in that: The device comprises: The initial collection module is used to generate an initial query set based on several customized seed queries before the update collection module is started; A knowledge point acquisition module, used to search each seed query in the initial query set, obtain basic web page information and web page document content related to each seed query, and generate initial knowledge point data of the initial query set; A knowledge point pruning module is used to classify the initial knowledge point data based on a custom binary classification model, filter out knowledge point data that is not relevant to the professional field in which the domain knowledge base to be constructed is located, and store the relevant knowledge point data as valid knowledge point data in the domain knowledge base to be constructed; A convergence judgment module is used to determine the knowledge base construction result according to the valid knowledge point data and the self-defined knowledge base construction convergence condition, wherein the construction result includes the end of construction and the continuation of construction. When the construction result is the continuation of construction, it is transferred to the update set module, and when the construction result is the end of construction, it is transferred to the result output module; An update set module, used for expanding and clustering the initial query set based on the valid knowledge point data, generating an updated query set, and setting the updated query set as the initial query set; The result output module is used to generate a domain knowledge base based on all the knowledge point data stored in the domain knowledge base to be constructed.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 8.
11. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Knowledge base construction method and device
CN109800879A
Electric power scientific research achievement knowledge base construction method based on large language model
CN118093768A
Knowledge-enriched item set expansion system and method
US11734365B1