Knowledge base construction method and device, equipment and medium

By building a target tree structure, using the association between document sub-information and knowledge document nodes, the problem of inaccurate search results in the rule matching method is solved, and efficient multi-dimensional information retrieval of the knowledge base is realized.

CN120297399AActive Publication Date: 2025-07-11ZHONGDIAN DATA IND CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510783018.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The prior art knowledge base search method based on rule matching is difficult to cover all possible problem variants and complex situations, resulting in inaccurate search results.

Method used

By obtaining the document type of the knowledge document, analyzing the document sub-information and generating the target tree structure, building a knowledge base, using the document sub-information as the first node and the knowledge document as the second node, combining the position identification information and document description information for node association, and building a target tree structure to support multi-dimensional knowledge information retrieval.

Benefits of technology

It improves the accuracy and efficiency of knowledge information retrieval, avoids data loss, and supports multi-dimensional information retrieval needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297399A_ABST
    Figure CN120297399A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a knowledge base construction method and device, equipment and a medium, and the method comprises the steps: obtaining a plurality of knowledge documents, and determining a document type corresponding to each knowledge document; analyzing each knowledge document according to a predetermined document analysis rule corresponding to the document type, obtaining at least one piece of document sub-information of each knowledge document, and determining position identification information of the document sub-information in the knowledge document; obtaining document description information of each knowledge document and a knowledge type to which the document description information belongs; and taking each piece of document sub-information as a first node, taking each knowledge document as a second node, and performing node association on each first node and the corresponding second node to generate a target tree structure of the knowledge base so as to construct the knowledge base. According to the technical scheme, the information retrieval accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a method, apparatus, device, and medium for constructing a knowledge base. Background Art

[0002] With the development of computer technology, it has become a common requirement to provide corresponding answers based on the input retrieval text. For example, in the knowledge system within an enterprise, it is common to retrieve the enterprise's internal knowledge base by inputting retrieval text to obtain retrieval answers.

[0003] In related technologies, a rule matching-based method is used to meet the retrieval requirements. Specifically, rules for obtaining answers are set in advance, and answers corresponding to the retrieval text are obtained in the knowledge base based on the rules for obtaining answers. For example, if the matching rule for the answer is keyword retrieval, then answers corresponding to the retrieval text are matched in the knowledge base based on the keyword matching rule.

[0004] However, in the above method of using rule matching to meet the retrieval requirements, since it is difficult to enumerate all rules and it is difficult to set rules that cover all possible problem variants and complex situations, the retrieval results may be inaccurate. Summary of the Invention

[0005] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for constructing a knowledge base.

[0006] An embodiment of the present disclosure provides a method for constructing a knowledge base. The method includes: obtaining a plurality of knowledge documents and determining the document type corresponding to each knowledge document; parsing each knowledge document according to a document parsing rule corresponding to the document type determined in advance, obtaining at least one document sub-information of each knowledge document, and determining the position identification information of the document sub-information in the knowledge document; obtaining the document description information and the knowledge type to which each knowledge document belongs; associating each document sub-information as a first node with the corresponding second node, which is each knowledge document, to generate a target tree structure of the knowledge base, and further constructing the knowledge base; wherein, the node attribute of each first node includes the corresponding position identification information, and the node attribute of each second node includes the corresponding document description information and the knowledge type, and the target tree structure is used for knowledge information retrieval of the knowledge base.

[0007] The embodiments of the present disclosure also provide a device for constructing a knowledge base. The device includes: a first acquisition module, configured to acquire a plurality of knowledge documents; a determination module, configured to determine the document type corresponding to each of the knowledge documents; a second acquisition module, configured to parse each of the knowledge documents according to a pre-determined document parsing rule corresponding to the document type, acquire at least one document sub-information of each of the knowledge documents, and determine the position identification information of the document sub-information in the knowledge document; a third acquisition module, configured to acquire the document description information and the knowledge type to which each of the knowledge documents belongs; a construction module, configured to use each of the document sub-informations as a first node, each of the knowledge documents as a second node, and associate each of the first nodes with the corresponding second node to generate a target tree structure of the knowledge base, and further construct the knowledge base; wherein, the node attribute of each of the first nodes includes the corresponding position identification information, and the node attribute of each of the second nodes includes the corresponding document description information and the knowledge type, and wherein, the target tree structure is used for knowledge information retrieval of the knowledge base.

[0008] The embodiments of the present disclosure also provide an electronic device. The electronic device includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for constructing a knowledge base provided by the embodiments of the present disclosure.

[0009] The embodiments of the present disclosure also provide a computer-readable storage medium. The storage medium stores a computer program, and the computer program is configured to execute the method for constructing a knowledge base provided by the embodiments of the present disclosure.

[0010] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art: The construction solution of the knowledge base provided by the embodiments of the present disclosure obtains multiple knowledge documents, determines the document type corresponding to each knowledge document, parses each knowledge document according to the pre-determined document parsing rules corresponding to the document type, obtains at least one document sub-information of each knowledge document, and determines the position identification information of the document sub-information in the knowledge document. Furthermore, the document description information and the knowledge type to which each knowledge document belongs are obtained. Each document sub-information is used as a first node, each knowledge document is used as a second node, and each first node is associated with the second node to which it belongs to generate the target tree structure of the knowledge base, and then the knowledge base is constructed. Among them, the node attribute of each first node includes the corresponding position identification information, and the node attribute of each second node includes the corresponding document description information and knowledge type, where the target tree structure is used for knowledge information retrieval in the knowledge base. In this technical solution, based on the knowledge points to which the knowledge documents belong, the included working sub-documents, etc., the target tree structure is constructed. On the basis of avoiding the loss of knowledge document data, the knowledge documents are described from multiple dimensions, and the retrieval requirements of work information can be met from multiple dimensions, improving the retrieval accuracy and efficiency of information. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flowchart of a method for constructing a knowledge base provided by an embodiment of the present disclosure; Figure 2 is a schematic structural diagram of a target tree structure provided by an embodiment of the present disclosure; Figure 3 is a flowchart of another method for constructing a knowledge base provided by an embodiment of the present disclosure; Figure 4 is a flowchart of yet another method for constructing a knowledge base provided by an embodiment of the present disclosure; Figure 5 is a flowchart of still another method for constructing a knowledge base provided by an embodiment of the present disclosure; Figure 6 is a flowchart of still another method for constructing a knowledge base provided by an embodiment of the present disclosure; Figure 7 is a flowchart of still another method for constructing a knowledge base provided by an embodiment of the present disclosure; Figure 8 is a schematic structural diagram of a knowledge base construction device provided by an embodiment of the present disclosure; Figure 9A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0014] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0015] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0016] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0017] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0019] To solve the above problems, an embodiment of the present disclosure provides a method for constructing a knowledge base. The method will be introduced below in combination with specific embodiments.

[0020] Figure 1 A flowchart of a method for constructing a knowledge base provided by an embodiment of the present disclosure. This method can be executed by a knowledge base construction device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, the method includes: Step 101: Obtain multiple knowledge documents and determine the document type corresponding to each knowledge document.

[0021] Among them, the knowledge documents include various documents produced by an enterprise in the work process, including but not limited to technical documents, contract documents, etc. In the embodiments of the present disclosure, the knowledge documents can be obtained periodically at a preset time interval to ensure the comprehensiveness of the document information covered in the subsequent constructed target tree structure.

[0022] In the embodiments of the present disclosure, determine the document type corresponding to each knowledge document. Among them, the document type includes but is not limited to at least one of the following: technical documents, institutional documents, management documents, etc. Among them, there may be a hierarchical relationship between different document types. For example, under technical documents, there may be sub-document types such as technical specification documents, technical R & D inclusion documents, etc.

[0023] Step 102: Parse each knowledge document according to the pre-determined document parsing rules corresponding to the document type, obtain at least one document sub-information corresponding to each knowledge document, and determine the position identification information of each document sub-information in the knowledge document.

[0024] Among them, the document parsing rules are used to parse the knowledge documents to break them down into the granularity of document sub-information. Among them, the document sub-information can be part of the information in the knowledge document. For example, the document sub-information is a document fragment included in the knowledge document. The document sub-information can also be the description information of part of the information included in the knowledge document. The information format of this description information includes but is not limited to text, video, etc. For example, when there are pictures in the knowledge document, the document sub-information can be the descriptive text information of the pictures. Another example is that when there is a PPT in the knowledge document, the document sub-information can be the video information describing the PPT, etc.

[0025] It should be understood that the information types included in the knowledge documents corresponding to different document types are different. Therefore, in order to more comprehensively and better sort out the document sub-information included in each knowledge document, each knowledge document is parsed according to the pre-determined document parsing rules corresponding to the document type to obtain at least one document sub-information corresponding to each knowledge document.

[0026] It should be noted that in different application scenarios, the method of parsing each knowledge document according to the pre-determined document parsing rules corresponding to the document type is different. The examples are as follows: In some possible embodiments, document parsing codes corresponding to each document parsing rule can be pre-constructed, and by calling the document parsing codes corresponding to each document parsing rule, the knowledge documents are parsed to obtain at least one document sub-information corresponding to each knowledge document.

[0027] In some possible embodiments, each type of work matches a corresponding document parsing rule, including the specific processing flow for knowledge documents. The knowledge documents are processed according to the corresponding processing flow to obtain at least one document sub - information corresponding to each knowledge document.

[0028] For example, when the document type includes structured document types, where the structured document type can be a CSV document type, an excel document type, etc. In this example, each structured data included in the knowledge document of the structured document type is parsed. For example, when the knowledge document is an excel document, the tabular data included in the excel is parsed to determine the data association relationship between each structured data. The data association relationship is used to indicate the relationship between each structured data. For example, when the knowledge document is an excel document, the tabular data included in the excel, data belonging to the same category or column can be regarded as having a data association relationship, or a group of data belonging to the same type can be regarded as having a data association relationship. The data association relationship is related to the specific table structure. Furthermore, the structured data with a data association relationship is used as a document sub - information. For example, when the knowledge document is an excel document and the excel table contains m groups of data for m objects, each group of data for each object in the m groups of data can be used as a document sub - information.

[0029] Another example is that when the document type includes unstructured document types, where the unstructured document type can include a word document type, a TXT document type, etc. In this embodiment, each information element included in the knowledge document corresponding to the document type is identified. The information element includes at least one of the following: pictures, videos, text fragments. A natural - language description text of each information element is generated. For example, when the information element includes a text fragment, the text fragment is input into a large - language model to obtain the natural - language description text of the text fragment. In this embodiment, the natural - language description text is used as a document sub - information.

[0030] In the embodiments of the present disclosure, the position identification information of each document sub - information in the knowledge document to which it belongs is also determined. The position identification information may include at least one of the following: the document sub - title of the document fragment where the document sub - information is located, the document sub - abstract of the document fragment where the document sub - information is located, etc. The document position identification information is used to assist in positioning the source position of the document sub - information in the knowledge document. The position identification information corresponding to the document sub - information can be extracted from the document sub - information or obtained by summarizing the content of the document sub - information and the knowledge document to which it belongs through a pre - trained model.

[0031] Step 103, obtain the document description information and the knowledge type to which each knowledge document belongs.

[0032] Among them, the document description information may include any unique identifier information for locating the document, including at least one of the following: document title, document abstract. The position identifier information may include at least one of the following: document subtitle of the document fragment where the document sub-information is located, document sub-abstract of the document fragment where the document sub-information is located.

[0033] In an embodiment of the present disclosure, the document description information of each knowledge document and the knowledge type to which each knowledge document belongs are obtained, where the knowledge type corresponds to a work link, and the knowledge type is used to describe the knowledge type included in the work information generated in the work link. For example, the work link may include: raw material procurement, production manufacturing, quality inspection, product sales, etc. Each work link is related to the knowledge type of the enterprise. For example, the procurement link involves knowledge types such as supplier information and procurement contract terms.

[0034] Step 104, taking each document sub-information as the first node, taking each knowledge document as the second node, and associating each first node with the second node to which it belongs to generate the target tree structure of the knowledge base, and then constructing the knowledge base; among them, the node attribute of each first node includes the corresponding position identifier information, and the node attribute of each second node includes the corresponding document description information and knowledge type, where the target tree structure is used for knowledge information retrieval in the knowledge base.

[0035] In an embodiment of the present disclosure, the target tree structure is generated by comprehensively considering the knowledge type, document type, and document sub-information of the knowledge document. On the basis of ensuring the content integrity of the knowledge document, the entire processing flow is automated, reducing the labor cost. In an embodiment of the present disclosure, each document sub-information is used as the first node, and each knowledge document is used as the second node. A node association is constructed between the first node and the second node to which it belongs, that is, a connection edge pointing from the first node to the second node to which it belongs is constructed to generate the target tree structure. The node attribute of each first node in the target tree structure includes the position identifier information of the first node. For example, it includes the subtitle information of the first node, etc. The node attribute of the second node includes the document description information and knowledge type of the second node. In some other possible embodiments, the node attribute of the second node may also include the document type, etc. In an embodiment of the present disclosure, the target tree structure constructs the knowledge base in the form of a tree structure.

[0036] For example, among the knowledge documents, there are knowledge document A and knowledge document B. The document sub-information included in knowledge document A includes a1, a2, a3... a7, and the document sub-information included in knowledge document B includes b1, b2... b9. Then the corresponding target tree structure is as Figure 2 shown, and the target tree structure can be intuitively displayed on the relevant interface. When the user triggers the corresponding node, the attribute information of the node is displayed. For example, in Figure 2When triggering Knowledge Document A, the node attribute information of Knowledge Document A is displayed. Of course, in other possible ways, the node attribute information can also be displayed in other ways, which will not be enumerated one by one here.

[0037] Among them, the target tree structure constructed in this disclosure can be used for the retrieval of knowledge information. Among them, the retrieval method can be set according to specific scenarios. In some possible embodiments, such as Figure 3 As shown, the steps for retrieving knowledge information based on the target tree structure including the first node and the second node may include: Step 301, in response to obtaining a retrieval request carrying a retrieval text, semantically match the retrieval text with the document sub-information of each first node to obtain a semantic matching degree.

[0038] Among them, the retrieval text can be used to indicate the description text of the obtained work information, etc. For example, the retrieval text can be "Obtain the sales promotion information of Product A".

[0039] In the embodiments of this disclosure, in response to obtaining a retrieval request carrying a retrieval text, semantically match the retrieval text with the document sub-information of each first node to obtain a semantic matching degree, that is, perform information retrieval at the granularity of document sub-information.

[0040] Step 302, determine at least one first node that matches successfully in the target tree structure according to the semantic matching degree.

[0041] In the embodiments of this disclosure, determine at least one first node that matches successfully in the target tree structure according to the semantic matching degree. For example, directly use the first nodes with a semantic matching degree greater than a certain threshold as the first nodes that match successfully.

[0042] Step 303, generate an answer corresponding to the retrieval request according to the document sub-information of at least one first node that matches successfully.

[0043] In the embodiments of this disclosure, an answer corresponding to the retrieval request can be generated according to the document sub-information of at least one first node that matches successfully. For example, the document sub-information of at least one first node can be input into a pre-trained large language model to obtain the answer output by the large language model.

[0044] In summary, for the method for constructing a knowledge base according to the embodiments of the present disclosure, multiple knowledge documents are obtained, and the document type corresponding to each knowledge document is determined. Each knowledge document is parsed according to the pre-determined document parsing rule corresponding to the document type to obtain at least one document sub-information corresponding to each knowledge document. Furthermore, the document description information of each knowledge document and the knowledge type to which each knowledge document belongs are obtained, and the position identification information of each document sub-information in the knowledge document to which it belongs is determined. Each document sub-information is used as a first node, and each knowledge document is used as a second node, and a connection edge pointing from each first node to the second node to which it belongs is constructed to generate a target tree structure. Among them, the node attribute of each first node in the target tree structure includes the position identification information of the first node, and the node attribute of each second node includes the document description information and knowledge type of the second node. The target tree structure is used for work information retrieval. In this technical solution, based on the knowledge points to which the knowledge documents belong, work sub-documents included, etc., a target tree structure is constructed. On the basis of avoiding the loss of knowledge document data, the knowledge documents are described from multiple dimensions, and the retrieval requirements of work information can be met from multiple dimensions, improving the retrieval accuracy and efficiency of information.

[0045] Based on the above embodiments, the dimension of the target tree structure can be further extended to meet the retrieval requirements of more dimensions.

[0046] In an embodiment of the present disclosure, as Figure 4 shown, the method further includes: Step 401, determine each work department corresponding to the business process, and use each work department as a third node.

[0047] Among them, the work department corresponds to the business process. For example, if the business process includes raw material procurement, production manufacturing, quality inspection, product sales, etc., then each business process corresponds to a specific work department. For example, the quality inspection process corresponds to the quality inspection department, etc.

[0048] Step 402, in the target tree structure, associate each second node with the third node to which it belongs, and associate the third nodes corresponding to each work department according to the business process to update the target tree structure.

[0049] In an embodiment of the present disclosure, each working department is used as a third node, and each second node is associated with the belonging third node in the target tree structure. For example, in this embodiment, connection edges pointing from each second node to the belonging third node are constructed to update the target tree structure, that is, the working department to which the knowledge document corresponding to each second node belongs is determined. With the updated target tree structure, various knowledge documents can be associated through department relationships, further meeting the retrieval requirements in more scenarios. In this embodiment, the third nodes corresponding to each working department are also associated according to the business process. For example, the third nodes are connected in the business flow direction of the business process. For example, when the third nodes include the sales department and the finance department, since the sales department needs financial approval from the finance department, there is a connection edge pointing from the sales department to the finance department between the sales department and the finance department.

[0050] Among them, when storing the target tree structure, in order to avoid data loss, a suitable database can be selected according to the data volume corresponding to the target tree structure. Among them, the database can include MySQL, Oracle, etc., and the corresponding nodes and node attributes can be stored in the database according to the subordination relationship corresponding to the target tree structure. Among them, in order to facilitate the complete storage of data, different fields can be used in the database to save the node attributes of different nodes.

[0051] In an embodiment of the present disclosure, continuing with the Figure 2 scenario shown as an example, when knowledge document A belongs to the third node S1 and knowledge document B belongs to the third node S2, there is a connection edge where A points to S1, and a connection edge where B points to S2. When there is a connection relationship where S1 points to S2 between the S1 department and the S2 department, the knowledge documents between different working departments are also associated in the target tree structure, facilitating the query of knowledge documents in a certain dimension of different working departments, etc.

[0052] It should be emphasized that in constructing the target tree structure, the multi-dimensional information retrieval requirements can be met based on different dimensions of information included in the target tree structure. For example, as Figure 5 shown, in the scenario of document traceability based on the target tree structure including the first node, the second node, and the third node, the specific retrieval steps may include: Step 501, in response to obtaining a document traceability request carrying document content, semantically match the document content with the document sub-information of each first node.

[0053] Among them, the document content is used to indicate the content to be retrieved. For example, when retrieving a contract for a certain product, the document content can be the contract content describing a certain product.

[0054] In this embodiment, in response to obtaining a document traceability request carrying document content, the document content is semantically matched with the document sub-information of each first node.

[0055] Step 502, determine at least one second node connected to the first node that matches successfully in the target tree structure according to the semantic matching degree.

[0056] In an embodiment of the present disclosure, determine the second node to which at least one first node that matches successfully belongs in the target tree structure according to the semantic matching degree. Among them, the second node can be one or multiple.

[0057] Step 503, determine a third node connected to the connected second node in the target tree structure.

[0058] Step 504, generate an answer corresponding to the document traceability request according to the working department of the connected third node and the knowledge type of the connected second node.

[0059] After determining at least one second node connected to the first node that matches successfully, determine a third node connected to the connected second node in the target tree structure, where the third node can be one or multiple.

[0060] After determining the third node, generate an answer corresponding to the document traceability request according to the working department of the connected third node and the knowledge type of the connected second node. This answer indicates the working department, knowledge type, etc. associated with the retrieved document content. Thus, when there are multiple second nodes, different knowledge documents, different working departments, etc. associated with the document content can be queried based on the document content.

[0061] In summary, the method for constructing the knowledge base in the embodiment of the present disclosure associates and integrates working sub-documents, knowledge documents, working departments, etc. to construct a target tree structure, ensuring that multi-dimensional information retrieval requirements can be met based on the target tree structure, etc.

[0062] In an embodiment of the present disclosure, as Figure 6 shown, the following steps can be adopted to determine the document type corresponding to each knowledge document, including: Step 601, extract the document features of each knowledge document.

[0063] Among them, the document features can include any features that can distinguish different document types, including but not limited to: document format (such as Word, Excel, PDF, etc.), document function (such as report, contract, manual, etc.), document content property type (such as technical document, business document, management document, etc.).

[0064] Step 602: Match the document features of each knowledge document with the node attributes of the document type nodes in the pre-constructed document tag system. The document tag system includes multiple document type nodes, each document type node corresponds to a document type, and node associations are made between document type nodes with a subordination relationship. The node attributes of each document type node include the document features of the corresponding document type.

[0065] It can be understood that a document tag system is pre-constructed. The document tag system includes multiple document type nodes, each document type node corresponds to a document type, and node associations are made between document type nodes with a subordination relationship. For example, document type nodes with a subordination relationship are connected by edges, where the direction of the edge points to the document type node it belongs to. For example, under the technical document type, it may include document types such as product technical specification documents and technology R & D report documents. Then, in the document tag system, there is a connection edge from the product technical specification document type node to the technical document type direction, and a connection edge from the technology R & D report document type to the technical document type direction. The node attributes of each document type node in the embodiments of the present disclosure include the document features of the corresponding document type.

[0066] In the embodiments of the present disclosure, the document features of each knowledge document are matched with the node attributes of the document type nodes in the pre-constructed document tag system, so as to automatically classify the document types of the knowledge documents according to the matching results.

[0067] Step 603: Determine the document type nodes with successful matches according to the matching results, and determine the document type corresponding to the document type nodes with successful matches as the document type corresponding to each knowledge document.

[0068] In this embodiment, the document type nodes with successful matches are determined according to the matching results. Furthermore, the document type corresponding to the document type nodes with successful matches is determined as the document type corresponding to each knowledge document.

[0069] In some possible embodiments, the document features of each knowledge document can be matched with the node attributes of the document type nodes of the first level in the document tag system, where the first level is the highest level. After determining the document type nodes of the first level with successful matches, in the order from the highest node level to the lowest, the document features of each knowledge document are matched with the node attributes of the document types belonging to the document type nodes of the first level with successful matches, that is, hierarchical matching, to improve the efficiency of determining the document type.

[0070] In this embodiment, the document type node with the lowest level among all the document type nodes with successful matches is determined as the document type node with a successful match.

[0071] In one embodiment of the present disclosure, the node attributes of each document type node further include the document parsing rules for the corresponding document type. Before parsing each knowledge document according to the document parsing rules corresponding to the document type determined in advance, the document parsing rules corresponding to the document type node with successful matching are determined, which are the document parsing rules corresponding to the document type. Thus, the document parsing rules are automatically determined, further improving the parsing efficiency and accuracy of parsing the knowledge documents.

[0072] In one embodiment of the present disclosure, as Figure 7 shown, the steps of determining the knowledge type to which each knowledge document belongs include: Step 701: Match the document content of each knowledge document with the node attributes of each work link node in the pre-constructed knowledge label system. The knowledge label system includes multiple work link nodes, and node associations are made between work link nodes with a subordination relationship. The node attributes of each work link node include the knowledge types associated with the corresponding work link.

[0073] In one embodiment of the present disclosure, a knowledge label system is pre-constructed. The knowledge label system includes multiple work link nodes, and node associations are made between work link nodes with a subordination relationship. For example, work link nodes with a subordination relationship are connected by edges, and the direction of the connection edge points from the work link node to the work link node to which it belongs. That is, the knowledge label system reflects the hierarchical relationship between work link nodes. Work links are related to work processes. In the embodiments of the present disclosure, the node attributes of each work link node include the knowledge types associated with the corresponding work link. Based on the knowledge label system, knowledge can be utilized more comprehensively.

[0074] Step 702: Determine the work link node with successful matching according to the matching result, and determine the knowledge type corresponding to the work link node with successful matching, which is the knowledge type of each document sub-information.

[0075] In this embodiment, the document content of each knowledge document is matched with the node attributes of each work link node in the pre-constructed knowledge label system. Furthermore, according to the matching result, the work link node with successful matching is determined, and the knowledge type corresponding to the work link node with successful matching is determined, which is the knowledge type of each knowledge document.

[0076] In some possible embodiments, the document content of each knowledge document can be matched with the node attributes of the work process nodes at the first level in the knowledge label system, where the first level is the highest-level node. After determining the work process nodes at the first level with successful matching, in the order from the highest node level to the lowest, the document content of each work document is matched with the node attributes of the work process nodes belonging to the work process nodes at the first level with successful matching, and the lowest-level work process node among all the work process nodes with successful matching is determined as the work process node with successful matching.

[0077] In summary, the method for constructing a knowledge base according to the embodiments of the present disclosure determines the document work type based on the constructed document label system, determines the knowledge type of the knowledge document based on the constructed knowledge label system, realizes the automation of determining the document work type and the knowledge type, and ensures the consistency of determining the document work type and the knowledge type.

[0078] To implement the above embodiments, the present disclosure also proposes a device for constructing a knowledge base.

[0079] Figure 8 As shown in the structural schematic diagram of a device for constructing a knowledge base provided by the embodiments of the present disclosure, the device can be implemented by software and / or hardware and is generally integrated in an electronic device. Figure 8 As shown, the device includes: a first acquisition module 810, a determination module 820, a second acquisition module 830, a third acquisition module 840, and a construction module 850, where The first acquisition module 810 is configured to acquire a plurality of knowledge documents; The determination module 820 is configured to determine the document type corresponding to each knowledge document; The second acquisition module 830 is configured to parse each knowledge document according to a document parsing rule corresponding to the document type determined in advance, acquire at least one document sub-information corresponding to each knowledge document, and determine the position identification information of the document sub-information in the knowledge document; The third acquisition module 840 is configured to acquire the document description information of each of the knowledge documents and the knowledge type to which it belongs; The construction module 850 is configured to use each document sub-information as a first node, each knowledge document as a second node, and associate each first node with the second node to which it belongs to generate a target tree structure of the knowledge base, and further construct the knowledge base; where the node attribute of each first node includes the corresponding position identification information, and the node attribute of each second node includes the corresponding document description information and knowledge type, where the target tree structure is used for knowledge information retrieval in the knowledge base.

[0080] The knowledge base construction apparatus provided by the embodiments of the present disclosure can execute the knowledge base construction method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0081] To implement the above embodiments, the present disclosure also proposes a computer program product, including computer programs / instructions, which, when executed by a processor, implement the knowledge base construction method in the above embodiments.

[0082] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.

[0083] Specifically, referring to Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The electronic device 900 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0084] As Figure 9 shown, the electronic device 900 may include a processor (such as a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a memory 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0085] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 shows the electronic device 900 having various devices, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.

[0086] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the memory 908, or installed from the ROM 902. When the computer program is executed by the processor 901, the above functions defined in the method for constructing the knowledge base of the embodiment of the present disclosure are performed.

[0087] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0088] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0089] The above computer-readable medium can be included in the above electronic device; it can also exist separately and not be assembled into the electronic device.

[0090] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to execute the above method for constructing the knowledge base.

[0091] The electronic device can write computer program code for performing the operations of the present disclosure in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0093] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0094] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0095] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0096] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0097] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0098] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for constructing a knowledge base, characterized in that, Including: Obtain a plurality of knowledge documents and determine the document type corresponding to each of the knowledge documents; Parse each of the knowledge documents according to a pre-determined document parsing rule corresponding to the document type, obtain at least one document sub-information of each of the knowledge documents, and determine the position identification information of the document sub-information in the knowledge document; Obtain the document description information and the knowledge type to which each of the knowledge documents belongs; Use each of the document sub-informations as a first node, use each of the knowledge documents as a second node, and associate each of the first nodes with the second node to which it belongs to generate a target tree structure of the knowledge base, and then construct the knowledge base; wherein, the node attribute of each of the first nodes includes the corresponding position identification information, and the node attribute of each of the second nodes includes the corresponding document description information and the knowledge type, wherein the target tree structure is used for knowledge information retrieval of the knowledge base.

2. The method according to claim 1, characterized in that, The method further includes: Determine each work department corresponding to the business process, and use each of the work departments as a third node; In the target tree structure, associate each of the second nodes with the third node to which it belongs, and associate the third nodes corresponding to each of the work departments according to the business process to update the target tree structure.

3. The method according to claim 1, characterized in that, The determining the document type corresponding to each of the knowledge documents includes: Extract the document features of each of the knowledge documents; Match the document features of each of the knowledge documents with the node attributes of the document type nodes in a pre-constructed document tag system, wherein the document tag system includes a plurality of document type nodes, each of the document type nodes corresponds to a document type, the document type nodes with a subordination relationship are associated with each other, and the node attribute of each of the document type nodes includes the document features of the corresponding document type; Determine the document type node with a successful match according to the matching result, and determine the document type corresponding to the document type node with a successful match as the document type corresponding to each of the knowledge documents.

4. The method according to claim 3, wherein The matching the document features of each of the knowledge documents with the node attributes of the document type nodes in a pre-constructed document tag system includes: Match the document features of each of the knowledge documents with the node attributes of the document type nodes of the first level in the document tag system; After determining the document type node of the first level with a successful match, match the document features of each of the knowledge documents with the node attributes of the document type nodes belonging to the document type node of the first level with a successful match in order from the highest node level to the lowest node level.

5. The method according to claim 4, wherein The determining the document type node with a successful match according to the matching result includes: Determine the document type node of the lowest level among all the document type nodes with a successful match as the document type node with a successful match.

6. The method according to claim 3, characterized in that, The node attribute of each of the document type nodes further includes the document parsing rule of the corresponding document type. Before parsing each of the knowledge documents according to the pre-determined document parsing rule corresponding to the document type, it further includes: Determine the document parsing rule corresponding to the document type node with successful matching, which is the document parsing rule corresponding to the document type.

7. The method according to claim 1, characterized in that, The steps for determining the knowledge type to which each of the knowledge documents belongs include: Match the document content of each of the knowledge documents with the node attributes of each work link node in the pre-constructed knowledge label system. The knowledge label system includes multiple work link nodes, and node associations are made between work link nodes with a subordination relationship. The node attributes of each work link node include the knowledge type associated with the corresponding work link. Determine the work link node with successful matching according to the matching result, and determine the knowledge type corresponding to the work link node with successful matching as the knowledge type of each of the knowledge documents.

8. The method according to claim 7, wherein The matching of the document content of each of the knowledge documents with the node attributes of each work link node in the pre-constructed knowledge label system includes: Match the document content of each of the knowledge documents with the node attributes of the work link nodes at the first level in the knowledge label system. After determining the work link nodes at the first level with successful matching, in the order from the highest node level to the lowest, match the document content of each of the knowledge documents with the node attributes of the work link nodes belonging to the work link nodes at the first level with successful matching.

9. The method according to claim 8, characterized in that The determining of the work link node with successful matching according to the matching result includes: Determine the work link node at the lowest level among all the work link nodes with successful matching as the work link node with successful matching.

10. The method according to claim 1, characterized in that, When the document type includes a structured document type, parsing the corresponding knowledge document according to the document parsing rule corresponding to the document type to obtain at least one document sub-information corresponding to each of the knowledge documents includes: Parse each structured data included in the knowledge document corresponding to the structured document type. Determine the data association relationship between the structured data. Take the structured data with the data association relationship as one of the document sub-informations.

11. The method according to claim 1, characterized in that When the document type includes an unstructured document type, parsing the corresponding knowledge document according to the document parsing rule corresponding to the document type to obtain at least one document sub-information corresponding to each of the knowledge documents includes: Identify each information element included in the document corresponding to the document type, where the information element includes at least one of the following: picture, video, text segment. Generate a natural language description text for each information element, and take the natural language description text as one of the document sub-informations.

12. The method according to any one of claims 1-11, characterized in that, It also includes: In response to obtaining a retrieval request carrying a retrieval text, perform semantic matching between the retrieval text and the document sub-informations of each of the first nodes to obtain a semantic matching degree. Determine at least one first node with successful matching in the target tree structure according to the semantic matching degree. Generate an answer corresponding to the retrieval request according to the document sub-informations of the at least one first node with successful matching.

13. The method according to any one of claims 1-11, characterized in that, The method also includes: In response to obtaining a document traceability request carrying document content, semantically match the document content with the document sub-information of each of the first nodes; Determine at least one second node connected to the first node with a successful match in the target tree structure according to the semantic matching degree; Determine a third node connected to the connected second node in the target tree structure; Generate an answer corresponding to the document traceability request according to the working department of the connected third node and the knowledge type of the connected second node.

14. The method according to claim 1, wherein The document description information includes at least one of the following: Document title, document abstract; The location identification information includes at least one of the following: Document sub-title of the document fragment where the document sub-information is located, document sub-abstract of the document fragment where the document sub-information is located.

15. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method for constructing the knowledge base according to any one of claims 1-14 above.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to execute the method for constructing the knowledge base according to any one of claims 1-14 above.

Citation Information

Patent Citations

  • Knowledge graph determination method and device, computer equipment and storage medium

    CN113190687A

  • Metadata-based knowledge graph processing method

    CN115391571A

  • Document knowledge system automatic construction method and device, electronic equipment and storage medium

    CN117076686A

  • Knowledge base visual configuration method and device, equipment, storage medium and product

    CN118861272A

  • Knowledge question and answer method and device, readable medium, electronic equipment and program product

    CN118964694A