Automatic updating method of knowledge base, electronic equipment and storage medium

By automatically obtaining text information from preset information sources and storing it into the knowledge base, the problem of inefficient manual updates in the existing technology is solved, and efficient and accurate automatic update of the knowledge base is achieved.

CN120069034APending Publication Date: 2025-05-30UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510156068.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing knowledge base system relies on manual updates, resulting in inefficient efficiency, insufficient accuracy, high cost and poor real-time performance.

Method used

By obtaining keywords from the keyword database corresponding to the knowledge base, text information including the keyword is automatically extracted from the preset information source and stored in the knowledge base.

Benefits of technology

Automatic update of the knowledge base is realized, reducing labor costs and improving the accuracy and efficiency of updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069034A_ABST
    Figure CN120069034A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent scheduling, in particular to an automatic updating system and method of a knowledge base, electronic equipment and a storage medium. The method comprises the following steps: acquiring a first keyword from a keyword library corresponding to the knowledge base, wherein a plurality of keywords related to the knowledge base are pre-stored in the keyword library; extracting first text information comprising the at least one keyword from a preset information source; and storing the first text information in a knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent scheduling, and more particularly, to an automatic update method for a knowledge base, an electronic device, and a storage medium. Background Art

[0002] In the field of computer science, research hotspots emerge continuously, and new results are introduced very quickly, especially in interdisciplinary research. This has led to a sharp increase in the publication volume of academic papers, and many papers are published on public online platforms. Researchers need to quickly obtain the latest knowledge, especially cross-domain knowledge, to stay at the forefront of their respective fields. Large Language Models (LLMs) have become one of the important tools for researchers. However, these models have some limitations, such as poor timeliness and lack of knowledge in specific professional fields. To overcome these problems, Retrieval Augmented Generation (RAG) has become a popular solution. By using existing data to build a knowledge base, RAG can provide support for large language models, thereby helping them generate more accurate and relevant outputs.

[0003] Existing knowledge base systems rely on manual updates, resulting in low efficiency, insufficient accuracy, high costs, and poor real-time performance. This manual update method not only consumes a large amount of manpower and time, but also is prone to errors and cannot quickly respond to information changes. There is an urgent need for an automated solution to improve the update efficiency and accuracy. Summary of the Invention

[0004] An object of the present disclosure is to provide an automatic update method for a knowledge base to solve the problems of low efficiency and insufficient accuracy in manual updates.

[0005] According to a first aspect of the present disclosure, there is provided a method for automatically updating a knowledge base, including: obtaining a first keyword from a keyword library corresponding to the knowledge base, where a plurality of keywords related to the knowledge base are pre-stored in the keyword library;

[0006] extracting first text information including the at least one keyword from a preset information source;

[0007] storing the first text information into the knowledge base.

[0008] Optionally, after extracting the first text information including the at least one keyword from the preset information source, the method further includes:

[0009] determining the relevance of each keyword in the keyword library to the first text information;

[0010] And adjust the weight of each keyword according to the relevance, where the weight is used to determine the first keyword.

[0011] Optionally, determining the relevance of each keyword in the keyword library to the first text information document includes:

[0012] Calculating the relevance score of each keyword to the first text information, where the relevance score score is:

[0013]

[0014] Where Q is any one of the keywords, d is the first text information, and q i is a word in keyword Q, where IDF(q i ) represents the inverse document frequency of word q i , and R(q i , d) represents the relevance score of word q i to the first text information d; where:

[0015]

[0016] N is the total number of text information in the knowledge base, n(q i ) represents the number of text information in the knowledge base that contains word q i , k 1 and b are adjustable parameters, tf(q i , d) is the frequency of word q i appearing in the first text information d, is the adjusted word frequency, L d is the length of the first text information d, and L avg is the average length of the text information in the knowledge base.

[0017] Optionally, extracting the first text information including the at least one keyword from the preset information source includes:

[0018] Based on a preset time interval, obtaining a document file containing the at least one keyword from the preset information source;

[0019] Parsing the document file to obtain the first text information including the at least one keyword.

[0020] Optionally, before storing the first text information into the knowledge base, the method further includes:

[0021] Determining whether the first text information is repeated with the text information in the knowledge base and / or determining whether the first text information is relevant to the knowledge base;

[0022] Filter the first text information when the first text information is repeated with the text information in the knowledge base or the first text information is not relevant to the knowledge base.

[0023] Optionally, determining whether the first text information is repeated with the text information in the knowledge base includes:

[0024] Determine the first similarity hash value of the first text information;

[0025] Respectively determine the Hamming distance between the first similarity hash value and the similarity hash value of each text information in the knowledge base;

[0026] When any of the Hamming distances is less than a first preset threshold, determine that the first text information is repeated with the text information in the knowledge base.

[0027] Optionally, determining whether the first text information is relevant to the knowledge base includes:

[0028] Extract the core information in the first text information, where the core information includes at least one of a title, an abstract, and a second keyword;

[0029] Count the number of third keywords included in each of the core information, where the third keyword is a keyword included in the keyword library;

[0030] Determine the relevance score of the first text information and the knowledge base according to the number of the third keywords and the weight coefficient of the core information;

[0031] When the relevance score of the first text information and the knowledge base is less than a preset threshold, determine that the first text information is not relevant to the knowledge base.

[0032] Optionally, a corresponding knowledge graph is further stored in the knowledge base; after storing the first text information into the knowledge base, it includes:

[0033] Calculate the cosine similarity between the first text information and each knowledge in the knowledge graph;

[0034] When any of the cosine similarity speeds is greater than a second preset threshold, update the knowledge graph in the knowledge base.

[0035] According to a second aspect of the present disclosure, there is provided an electronic device, including a processor and a memory, where computer instructions are stored in the memory, and when the computer instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0036] According to a third aspect of the present disclosure, there is provided a storage medium having computer instructions stored thereon, and when the computer instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0037] One technical effect of the present disclosure is to provide a method for automatically updating a knowledge base. By obtaining keywords in a corresponding keyword library, text information including the keyword is automatically obtained from a preset information source and stored in the knowledge base. In this way, text information can be automatically obtained from the information source according to the keyword, thereby realizing the automatic update of the knowledge base, reducing labor costs, and at the same time improving the accuracy of the knowledge base update.

[0038] Through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings, other features and advantages of the embodiments of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the description, are used to explain the principles of the embodiments of the present disclosure.

[0040] Figure 1 is a flowchart of a method for automatically updating a knowledge base according to an embodiment;

[0041] Figure 2 is a schematic structural diagram of an electronic device according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0043] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended as a limitation on the present invention or its application or use.

[0044] Techniques and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques and devices should be considered as part of the specification.

[0045] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0046] It should be noted that: like reference numerals and letters denote like items in the following drawings; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0047] It should be noted that all actions related to data collection, storage, use, processing, transmission, provision, disclosure, deletion, etc. in this disclosure are carried out on the premise of complying with relevant data protection regulations and policies of the country or region where the data is located and with the full authorization of the corresponding data owners.

[0048] The embodiments of this application disclose an automatic update method for a knowledge base, as Figure 1 shown, including steps S11 - S13.

[0049] Step S11: Obtain a first keyword from the keyword library corresponding to the knowledge base. Multiple keywords related to the knowledge base are pre - stored in the keyword library.

[0050] In this embodiment, a knowledge base usually corresponds to a specific field or specific classification. On this basis, a corresponding keyword library can be configured for the knowledge base. Multiple keywords related to the corresponding knowledge of the knowledge base can be pre - stored in the keyword library. For example, in the field of medical and health, the corresponding keyword library may include various disease names, symptoms, drugs, etc. Or in the field of artificial intelligence, the corresponding keyword library may include model names, development tools, frameworks, etc. The keyword can be in the form of a single word, such as "cancer", "framework", etc., or it can be a keyword composed of multiple words, such as "development tools, application platforms, medical devices, etc.".

[0051] In this embodiment, the keywords in the keyword library can be added manually by users or automatically updated through a preset algorithm.

[0052] In one example, the number of the first keywords can be one or more, and each keyword in the keyword library can also be set with a corresponding weight. In one example, the number of the first keywords can only include one. The weight of the keyword is used to determine the frequency of using this keyword as the first keyword, that is, the frequency of updating the knowledge base through this keyword. In another example, the number of the first keywords can be multiple, and the weight of the keyword is used to determine which keywords are included in the first keywords. For example, keywords with weights greater than a threshold can be used as the first keywords.

[0053] Step S12: Extract first text information including at least one keyword from a preset information source;

[0054] In this embodiment, the preset information source can be correspondingly set based on the knowledge base. For example, for the knowledge bases in certain academic fields, the preset information source can be various academic websites, forums, or platforms, etc. Users can set the preset information source corresponding to the knowledge base according to their needs. In this embodiment, the HTML structure can be identified and parsed through web crawler technology and corresponding parsing tools, which can effectively locate the links and files storing text information, so as to obtain...

[0055] In an example of this embodiment, extracting the first text information including at least one keyword from the preset information source includes: obtaining the document files containing at least one keyword from the preset information source based on a preset time interval; parsing the document files to obtain the first text information including at least one keyword.

[0056] In an example, after determining the first keyword, text information including the first keyword can be regularly obtained from the information source based on a preset time interval. In this example, the preset time interval can be set according to the user's needs, or can be set based on the weight parameter of the first keyword. For the first keyword with a higher weight parameter, the preset time interval can be shorter; for the first keyword with a lower weight parameter, the preset time interval can be set longer.

[0057] After obtaining the corresponding file, it can be identified whether the format of the file can be processed subsequently. If it is a file in an unprocessable format such as pdf, etc., text information can be extracted from the file through a parsing tool.

[0058] Step S13, storing the first text information into the knowledge base.

[0059] In this example, an automatic update method for the knowledge base is provided. By obtaining the keywords in the corresponding keyword library, automatically obtaining the text information including the keyword from the preset information source, and storing the text information into the knowledge base. In this way, text information can be automatically obtained from the information source according to the keyword, so as to realize the automatic update of the knowledge base, reduce the labor cost, and at the same time improve the accuracy of the knowledge base update.

[0060] In an example of this embodiment, after extracting the first text information including at least one keyword from the preset information source, the method further includes: determining the relevance between each keyword in the keyword library and the first text information; and adjusting the weight of each keyword according to the relevance, where the weight is used to determine the first keyword.

[0061] In the embodiments of the present application, after obtaining the first text information, the weights of each keyword in the keyword library can also be updated based on the text information. In this example, the weights of keywords with higher relevance can be increased, and the weights of keywords with lower relevance can be decreased. In this way, it can adapt to the dynamic changes in the corresponding field, reflect the current hotspots in the field, and facilitate subsequent adaptive updates of the knowledge base.

[0062] Determining the relevance between each keyword in the keyword library and the first text information document includes:

[0063] Calculating the relevance score between each keyword and the first text information, where the relevance score score is:

[0064]

[0065] where Q is any keyword, d is the first text information, and q i is the word in keyword Q, where IDF(q i ) represents the inverse document frequency of word q i , and R(q i , d) represents the relevance score between word q i and the first text information d; where:

[0066]

[0067] N is the total number of text information in the knowledge base, n(q i ) represents the number of text information containing word q i in the knowledge base, k 1 and b are adjustable parameters, tf(q i , d) is the frequency of word q i appearing in the first text information d, is the adjusted term frequency, L d is the length of the first text information d, and L avg is the average length of the text information in the knowledge base.

[0068] In this example, if a word appears in more than half of the documents, this usually indicates that the word is not important at all because the discrimination is low. At this time, IDF(q i ) is negative, and the contribution of this word to the relevance score is negative. In addition, 0.5 is added to the numerator and denominator of the formula respectively to prevent the situation where the denominator of IDF(q i ) is 0 or out of the log domain.

[0069] It is to avoid the penalty of the algorithm on overly long texts. The role of b is to adjust the magnitude of the influence of the document length on relevance. The hyperparameter k1 plays a role in adjusting the scale of the text frequency of feature words. In one example, b = 0.75, k 1 ∈ [1.2, 2.0].

[0070] In an example of this embodiment, before storing the first text information in the knowledge base, the method further includes: determining whether the first text information is repeated with the text information in the knowledge base and / or determining whether the first text information is relevant to the knowledge base; filtering the first text information in the case where the first text information is repeated with the text information in the knowledge base or the first text information is not relevant to the knowledge base.

[0071] In this embodiment, after obtaining the text information in the preset information source, the text information can also be filtered. For example, filtering methods such as duplicate checking or excluding irrelevant content are used to ensure the relevance and storage efficiency of the content in the knowledge base. The extracted information can be refined using natural language processing techniques. Through semantic analysis and text classification algorithms, duplicate or irrelevant documents are identified and removed, eliminating redundant and irrelevant data, thereby effectively reducing text noise.

[0072] In an example of this embodiment, determining whether the first text information is repeated with the text information in the knowledge base includes: determining the first similarity hash value of the first text information; respectively determining the Hamming distance between the first similarity hash value and the similarity hash values of each text information in the knowledge base; determining that the first text information is repeated with the text information in the knowledge base in the case where any Hamming distance is less than the first preset threshold.

[0073] In this embodiment, the similarity hash value (SimHash) of each text information in the knowledge base can be determined in advance and stored for duplicate checking. In this example, the first text information or the text information in the knowledge base can be all the text information in the document or partial document information in the document, such as a certain paragraph or several paragraphs. The granularity of the text information can be specifically set according to needs. After obtaining the first text information, the hash value of the first text information can be calculated through the similarity hash algorithm. Then, this value is compared with the similarity hash values of each text information in the knowledge base, and the Hamming distance between them is calculated. If the Hamming distance between the first text information and a certain text information in the knowledge base is less than the threshold, it indicates that the two text information are repeated. If the Hamming distance is greater than or equal to the threshold, it indicates that the two text information are not repeated.

[0074] In an example of this embodiment, determining whether the first text information is relevant to the knowledge base includes: extracting the core information in the first text information, where the core information includes at least one of a title, an abstract, and a second keyword; counting the number of third keywords included in each core information, where the third keyword is a keyword included in the keyword library; determining a relevance score of the first text information and the knowledge base according to the number of third keywords and the weight coefficient of the core information; and determining that the first text information is not relevant to the knowledge base when the relevance score of the first text information and the knowledge base is less than a preset threshold.

[0075] In this embodiment, the core information of the first text information can be extracted through natural language processing technology. In this example, the core information can include at least one of a title, an abstract, and a second keyword. In this example, the second keyword in the core information can be different from or the same as the first keyword, which represents the core content in the first text information. Then, each core information can be respectively matched with all the keywords in the keyword library to determine the number of keywords included in the keyword library for each core information. After that, the relevance score of the first text information and the knowledge base is determined. Taking the core information including the title, abstract, and second keyword as an example, the relevance score sco is:

[0076] sco = ω 1 *a + ω 2 *b + ω 3 *c

[0077] where ω 1 、ω 2 、ω 3 are the weight coefficients of the title, abstract, and second keyword respectively, and a, b, and c are the numbers of third keywords included in the title, abstract, and second keyword respectively.

[0078] In an example of this embodiment, a corresponding knowledge graph is also stored in the knowledge base; after storing the first text information in the knowledge base, it includes: calculating the cosine similarity between the first text information and each knowledge in the knowledge graph; and updating the knowledge graph in the knowledge base when any cosine similarity is greater than a second preset threshold.

[0079] In this embodiment, a corresponding knowledge graph is also stored in the knowledge base, and the knowledge graph can be formed based on the text information in the knowledge base. First, knowledge extraction can be performed on the text information in the knowledge base to extract entities, relationships, attributes, etc. from the text information. Taking the knowledge base in the field of papers as an example, the extracted entities may include authors, papers, institutions, years, journals, etc. Then, the relationships between various entities are determined, such as author - publishes paper, etc. Extracting entities, relationships, and attributes from the text information can be done through a deep learning model, such as the BERT - BiLSTM - CRF model for recognition, or can be extracted through regular expressions or template matching. Then, the extracted entities and relationships are fused, including alignment and deduplication, to ensure that the same entity appears only once in the knowledge graph. The unification of entities is achieved through string similarity and semantic similarity calculations. The problems of duplication and ambiguity in the extraction results are solved. The fused knowledge is stored in a structured manner to support efficient querying and operations. The processed entities and relationships are stored in the knowledge base to formally form the knowledge graph.

[0080] In this example, the first text information and knowledge can be converted into vector form. For example, through the BERT (Bidirectional Encoder Representations from Transformers) model, the vector representation of the hidden layer is extracted, and then the sentence is encoded to obtain a fixed - length vector. Then, the cosine similarity between the two vectors is calculated. When the cosine similarity between the first text information and a certain knowledge is less than the threshold, it indicates that the first text information is relevant to the corresponding knowledge. The knowledge graph can be updated based on the first text information. In this example, the way to update the knowledge graph can be to generate a new knowledge graph based on the first text information, or to update and expand the knowledge graph in the knowledge base through an automated pipeline and inference tools.

[0081] As Figure 2 As shown in the figure, an electronic device 200 is further provided in an embodiment of the present application, including a processor 201 and a memory 202. Computer instructions are stored in the memory 202, and when the computer instructions are executed by the processor 201, the steps of any one of the methods in the embodiment of the automatic update method of the knowledge base are implemented.

[0082] An embodiment of the present application further provides a storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, any one of the above - mentioned embodiments of the automatic update of the knowledge base is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0083] The various embodiments in the present disclosure are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the description of the method embodiments.

[0084] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0085] Embodiments of the present disclosure may be systems, methods, and / or computer program products. A computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the embodiments of the present disclosure.

[0086] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0087] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0088] The computer program instructions for performing the operations of the embodiments of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the embodiments of the present disclosure.

[0089] Aspects of the embodiments of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0090] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0091] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0092] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As will be apparent to those skilled in the art, implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0093] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for automatically updating a knowledge base, characterized in that: include: Acquire a first keyword from a keyword library corresponding to the knowledge base, wherein the keyword library pre-stores a plurality of keywords related to the knowledge base; extracting first text information including the at least one keyword from a preset information source; The first text information is stored in a knowledge base.

2. The method according to claim 1, characterized in that: After extracting the first text information including the at least one keyword from the preset information source, the method further includes: Determining the relevance between each keyword in the keyword library and the first text information; And the weight of each keyword is adjusted according to the relevance, wherein the weight is used to determine the first keyword.

3. The method according to claim 2, characterized in that The determining the relevance between each keyword in the keyword library and the first text information document includes: Calculate the relevance score between each keyword and the first text information, where the relevance score is: Wherein, Q is any of the keywords, d is the first text information, q i is a word in keyword Q, where IDF(q i ) represents the word q i The inverse document frequency, R(q i ,d) represents word q i The relevance score with the first text information d; wherein: N is the total number of text information in the knowledge base, n(q i ) indicates that the knowledge base contains word q i The number of text information, k1 and b are adjustable parameters, tf(q i ,d) is word q i The frequency of occurrence in the first text information d, is the adjusted word frequency, L d is the length of the first text message d, L avg is the average length of text information in the knowledge base.

4. The method according to claim 1, characterized in that: The step of extracting the first text information including the at least one keyword from the preset information source includes: Based on a preset time interval, obtaining a document file containing the at least one keyword from the preset information source; The document file is parsed to obtain first text information including the at least one keyword.

5. The method according to claim 1, characterized in that Before storing the first text information in the knowledge base, the method further includes: Determining whether the first text information is repeated with text information in the knowledge base and / or determining whether the first text information is related to the knowledge base; In the case where the first text information is repeated with text information in the knowledge base or the first text information is irrelevant to the knowledge base, the first text information is filtered.

6. The method according to claim 5, characterized in that Determining whether the first text information is repeated with text information in the knowledge base includes: Determining a first similarity hash value of the first text information; respectively determining the Hamming distance between the first similarity hash value and the similarity hash value of each text information in the knowledge base; In the case where any of the Hamming distances is less than a first preset threshold, it is determined that the first text information is repeated with text information in the knowledge base.

7. The method according to claim 5, characterized in that The determining whether the first text information is related to the knowledge base includes: Extracting core information from the first text information, the core information including at least one of a title, an abstract and a second keyword; Counting the number of third keywords included in each of the core information, where the third keywords are keywords included in the keyword library; Determining a relevance score between the first text information and the knowledge base according to the number of the third keywords and the weight coefficient of the core information; When the relevance score between the first text information and the knowledge base is less than a preset threshold, it is determined that the first text information is not relevant to the knowledge base.

8. The method according to claim 1, characterized in that The knowledge base also stores a corresponding knowledge graph; after storing the first text information in the knowledge base, the method includes: Calculate the cosine similarity between the first text information and each piece of knowledge in the knowledge graph; In the case where any of the cosine phase velocities is greater than a second preset threshold, the knowledge graph in the knowledge base is updated.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the steps of the method described in any one of claims 1 to 8 are implemented.

10. A storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.