Knowledge retrieval method and apparatus
Patent Information
- Application Number
- CN202311125335.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-09-01
AI Technical Summary
普通检索只能够通过网页内容返回结果,使用算法根据相关性和受欢迎程度对结果进行排名和显示,因此信息来源不清晰,结果不准确或不可靠
[0010] The knowledge retrieval method and apparatus provided in the embodiments of this disclosure, by using technologies such as knowledge graphs and semantic matching, can more quickly locate and retrieve relevant information, improving retrieval efficiency. Fusing data from different modalities can increase the richness of retrieval results, providing users with more useful information. By fusing data from different modalities and technologies such as machine learning, information extraction, and natural language processing, four major retrieval capabilities are achieved: comprehensive retrieval, content retrieval, rich summary retrieval, and general resource retrieval. This allows for a more comprehensive understanding of the user's query intent, thereby providing more accurate retrieval results. Based on vertical industry sectors, it provides multimodal fine-grained document fragmentation resource retrieval and location functions, such as PDF document location and table location, effectively improving the user retrieval experience and meeting fine-grained retrieval needs.
Smart Images

Figure CN117171319B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, particularly to the field of intelligent search, and can be specifically applied to scenarios such as human-computer dialogue and electronic libraries. Specifically, it is a knowledge retrieval method and apparatus. Background Technology
[0002] Currently, the amount of online information is becoming increasingly vast and complex. Users often obtain knowledge from traditional search engines that is disorganized and lacks structure, resulting in low efficiency and poor quality, requiring significant time to filter for satisfactory results. For example, traditional enterprise document retrieval methods offer full-text search services at the whole document level. However, given the diverse search scenarios within an enterprise, they cannot provide more granular content than general domain searches to meet user needs.
[0003] Traditional search engines can provide general knowledge across broad domains, but the sheer volume and fragmented nature of knowledge in specific industry sectors makes searching difficult and time-consuming. General search, achieved through web search engines, only analyzes web pages for keywords to find relevant content. It returns results solely based on webpage content, ranking and displaying results using algorithms based on relevance and popularity; therefore, the information source is unclear, and the results may be inaccurate or unreliable. Furthermore, it fails to meet the needs of fine-grained searches for specialized documents within specific industry sectors, especially given users' limited understanding of such content. Summary of the Invention
[0004] This disclosure provides a knowledge retrieval method, apparatus, device, storage medium, and computer program product.
[0005] According to a first aspect of this disclosure, a knowledge retrieval method is provided, comprising: in response to receiving a search request, extracting entity information and faceted information from the search request; searching for target nodes that match the entity information and faceted information in a pre-constructed knowledge graph; obtaining source data for constructing the target nodes; parsing the semantics in the search request; and retrieving retrieval results that match the semantics from the source data.
[0006] According to a second aspect of this disclosure, a knowledge retrieval apparatus is provided, comprising: an extraction unit configured to extract entity information and faceted information from the search request in response to receiving a search request; a search unit configured to search for target nodes matching the entity information and faceted information in a pre-constructed knowledge graph; an acquisition unit configured to acquire source data for constructing the target nodes; a parsing unit configured to parse the semantics in the search request; and a matching unit configured to retrieve retrieval results matching the semantics from the source data.
[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in any one of the first aspects.
[0008] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method according to any one of the first aspects.
[0009] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of the first aspects.
[0010] The knowledge retrieval method and apparatus provided in the embodiments of this disclosure, by using technologies such as knowledge graphs and semantic matching, can more quickly locate and retrieve relevant information, improving retrieval efficiency. Fusing data from different modalities can increase the richness of retrieval results, providing users with more useful information. By fusing data from different modalities and technologies such as machine learning, information extraction, and natural language processing, four major retrieval capabilities are achieved: comprehensive retrieval, content retrieval, rich summary retrieval, and general resource retrieval. This allows for a more comprehensive understanding of the user's query intent, thereby providing more accurate retrieval results. Based on vertical industry sectors, it provides multimodal fine-grained document fragmentation resource retrieval and location functions, such as PDF document location and table location, effectively improving the user retrieval experience and meeting fine-grained retrieval needs.
[0011] In summary, knowledge retrieval design based on knowledge engines can provide more accurate, richer, and more personalized search results, helping to improve user experience and retrieval efficiency.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied;
[0015] Figure 2 This is a flowchart of an embodiment of the knowledge retrieval method according to the present disclosure;
[0016] Figures 3a-3f This is a schematic diagram of an application scenario of the knowledge retrieval method disclosed herein;
[0017] Figure 4 This is a flowchart of yet another embodiment of the knowledge retrieval method according to the present disclosure;
[0018] Figure 5 This is a schematic diagram of the structure of an embodiment of the knowledge retrieval device according to the present disclosure;
[0019] Figure 6 This is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the knowledge retrieval methods or devices disclosed herein can be applied.
[0022] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0023] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0024] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.
[0025] Server 105 can be a server that provides various services, such as a search engine server that supports the search results displayed on terminal devices 101, 102, and 103. The search engine server can analyze and process data such as received search requests and feed back the processing results (such as search results) to the terminal devices.
[0026] It's important to note that a server can be either hardware or software. When a server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When a server is software, it can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here. A server can also be a server for a distributed system, or a server integrated with blockchain technology. A server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0027] It should be noted that the knowledge retrieval method provided in the embodiments of this disclosure is generally executed by server 105, and correspondingly, the knowledge retrieval device is generally set in server 105.
[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0029] Continue to refer to Figure 2 The diagram illustrates a flow 200 of an embodiment of a knowledge retrieval method according to the present disclosure. The knowledge retrieval method includes the following steps:
[0030] Step 201: In response to receiving a search request, extract entity information and faceted information from the search request.
[0031] In this embodiment, the execution entity of the knowledge retrieval method (e.g.) Figure 1 The server shown can receive search requests from the user's web browsing terminal via a wired or wireless connection. The search request includes keywords. Optionally, the search request may also include user information, search function options, etc.
[0032] Entity information and faceted information can be extracted from search requests. Faceted information refers to the attribute information of an entity. This extraction can be achieved using pre-trained natural language processing models. For example, named entity recognition models or reading comprehension models can be used to extract entity and faceted information from keywords in a search request. If the keywords do not contain faceted information, all faceted information related to that entity can be used as default faceted information for the user to choose from. If the user does not select specific faceted information, the search will proceed based on all faceted information related to that entity.
[0033] Step 202: Search for target nodes that match entity information and faceted information in the pre-built knowledge graph.
[0034] In this embodiment, the knowledge graph is also known as a scientific knowledge graph, or in the library and information science community as a knowledge domain visualization or knowledge domain mapping map. It is a series of different graphics that show the development process and structural relationships of knowledge. It uses visualization technology to describe knowledge resources and their carriers, and to mine, analyze, construct, draw and display knowledge and the interrelationships between them.
[0035] Specifically, knowledge graphs are a modern theory that combines theories and methods from applied mathematics, computer graphics, information visualization, and information science with bibliometric methods such as citation analysis and co-occurrence analysis. It uses visualized graphs to vividly display the core structure, development history, cutting-edge fields, and overall knowledge architecture of a discipline, achieving multidisciplinary integration. It reveals complex knowledge domains through data mining, information processing, knowledge measurement, and graphical representation, showcasing the dynamic development patterns of these domains and providing practical and valuable references for disciplinary research.
[0036] Using knowledge graphs to organize search results links different result nodes together to form a structured knowledge network. This helps users better understand and apply the search results.
[0037] A knowledge graph can be constructed based on existing knowledge (source data). The knowledge graph consists of nodes and edges. Each node represents an entity, and each entity corresponds to faceted information. The edges between nodes represent the relationships between entities.
[0038] Nodes whose entity and facet information in the knowledge graph match the entity and facet information in the search request more than a predetermined threshold can be selected as target nodes through graph search.
[0039] Step 203: Obtain the source data for constructing the target node.
[0040] In this embodiment, each node in the knowledge graph is an integration of entity information, faceted information, and relational information extracted from a large amount of source data. It can not only obtain source data that directly contains entity information and faceted information, but also entity information and faceted information associated with entity information and faceted information, such as the superordinate and subordinate concepts of the entity corresponding to the target node.
[0041] Step 204: Parse the semantics in the search request.
[0042] In this embodiment, the semantics in the search request can be parsed using a pre-trained semantic recognition model. This can be achieved using natural language processing techniques, such as decomposing the query into words and extracting semantic roles.
[0043] Step 205: Retrieve the semantically matched search results from the source data.
[0044] In this embodiment, semantic information can be parsed from source data using a pre-trained semantic recognition model, and this semantic information can be annotated to the source data during the construction of the knowledge graph. During the search, the semantic information corresponding to the search request is matched with the semantic information of the source data. The matching degree can be calculated through string matching or through a pre-trained matching model. The input to the matching model is two semantic terms, and the output is the semantic matching degree.
[0045] The methods provided in the above embodiments of this disclosure, by using a combination of knowledge graph and semantic matching techniques, can more quickly locate and retrieve relevant information, thereby improving retrieval efficiency.
[0046] In some optional implementations of this embodiment, the search request includes keywords and at least one of the following search function options: comprehensive search, content search, rich summary search, and resource search; and the method further includes: selecting targeted search results from the search results based on the search function and displaying them. Based on the pain points of traditional search engines, the search function is divided into four major functional modules: comprehensive search, content search, rich summary search, and resource search (e.g., ...). Figure 3a(As shown). Comprehensive search provides a full range of results. Content search allows for mixed-format searches based on chapters, paragraphs, tables, etc. Rich summary search displays summaries of the search results. Resource search displays resources in different modalities. Users can select functions based on their specific search scenarios, providing a more efficient, accurate, and comprehensive knowledge retrieval service. The search interface offers four search function tabs. After selecting a tab and entering keywords, clicking "confirm" will filter and display targeted search results based on the selected search function.
[0047] In some optional implementations of this embodiment, the resource retrieval includes data formats of different modalities; and the step of filtering the search results from the search results to obtain accurate search results based on the search function includes: filtering search results from the search results for data formats selected by the user. Data of different modalities (such as text, images, videos, etc.) is merged, and the search results include other formats such as compressed files, Word, Excel, PPT, and PDF. This provides more comprehensive and accurate search results.
[0048] In some optional implementations of this embodiment, the method further includes: outputting the theoretical basis for the relevance of the search results to the search request. This provides interpretability in the search results, i.e., explaining why a particular result is relevant to the user's query. This can be achieved using an interpretability model, such as a decision tree or rule set. The nodes involved in the decision-making process or the rules used in the decision can be displayed as the theoretical basis.
[0049] The decision tree or rule set in this application is a mechanism for building and managing a knowledge base that can store and manage knowledge in a structured way to support intelligent decision-making and automated processing.
[0050] A decision tree (semantic tree) is a tree-like structure where each node represents an attribute or condition, and each branch represents a decision or prediction. By recursively making judgments based on conditions, a dataset can be divided into different subsets, and classification and prediction can be performed based on the features of each subset. Decision trees have the advantages of being intuitive, easy to construct, and easy to interpret.
[0051] A rule set represents decision rules as a set of conditions and conclusions. Each rule contains a premise and a conclusion, where the premise is the rule's precondition and the conclusion is the rule's prediction. Rule sets can be applied and executed through logical reasoning and matching, thereby achieving automated decision-making and processing. Rule sets are characterized by being explicit, concise, and easy to understand.
[0052] In some optional implementations of this embodiment, the method further includes: updating the knowledge graph in real time; and retrieving the search request based on the changed information if the knowledge graph changes. Integrating the retrieval system with the knowledge graph or other data sources allows for real-time updates and improvements to the knowledge in the knowledge graph. This helps provide more accurate search results and reduces the risk of knowledge becoming outdated.
[0053] In some optional implementations of this embodiment, the method further includes: obtaining the user's historical behavior information; and filtering recommended information for the user from the search results based on the historical behavior information. During the search process, personalized recommendation technology is used to recommend the most relevant results based on the user's interests, historical behavior, and other information. This can be achieved using a recommendation system, such as content-based recommendation or collaborative filtering recommendation.
[0054] In some optional implementations of this embodiment, the method further includes: filtering categorized search results from the search results based on the entity information and the faceted information. Before searching, such as... Figure 3b As shown, users can filter and search by entity and facet on the page to improve search accuracy. After the search card results are displayed, as shown... Figure 3c As shown, the results will also be categorized by entity and facet to help users better understand the knowledge.
[0055] In some optional implementations of this embodiment, the method further includes: highlighting the most relevant search results using box selection and color coding. For example... Figure 3d As shown, the search results card highlights the most relevant results to the user through box selection and color coding.
[0056] In some optional implementations of this embodiment, the method further includes: outputting the directory structure of the search results for users to preview and locate detailed content. For example... Figure 3d As shown, users can preview all detailed directory locations related to the document structure through the document directory structure on the left, and quickly switch between them. For example... Figure 3e As shown, clicking the card will take you to the document previewer, allowing you to navigate to the corresponding page for preview, thus improving the efficiency of finding knowledge.
[0057] In some optional implementations of this embodiment, the method further includes: extracting at least one of the following table information from the searched table file: title, header, and content; and matching the search request with the table information to obtain multi-granularity recall results. Tables are a data format with a large amount of information in enterprises. For table content in rich documents, table recall technology is applied, combining keywords in the search request with information from the table's title, header, and content to perform multi-granularity recall matching. String matching methods can be used to calculate the matching degree between keywords and the title, header, and content respectively, and titles, headers, and content with matching degrees greater than a predetermined threshold are recalled.
[0058] In some optional implementations of this embodiment, the method further includes: sorting the recall results using a pre-built tabular data sorting model. The tabular data sorting model can be equivalent to a fully connected layer, representing the weights of the matching scores of the title, header, and content. The tabular data sorting model can calculate a weighted sum of the matching scores of the title, header, and content. The weights can be obtained through supervised training using a large number of tables labeled with the overall matching score and the matching scores of the title, header, and content as samples.
[0059] In some optional implementations of this embodiment, the method further includes: locating block coordinates of the searched table file based on the semantics and labeling the recall results. End-to-end block location technology, combined with keyword intent requirements, can be used to directly locate block coordinates based on the entire table, satisfying retrieval needs with finer granularity. For example... Figure 3f As shown, the search results card displays the table titles and screenshots with high relevance, and marks the hit results with boxes, highlighting the hit table rows and columns, allowing users to quickly and clearly focus on the relevant results.
[0060] Further reference Figure 4 This illustrates a flow 400 of another embodiment of the knowledge retrieval method. Flow 400 of the knowledge retrieval method includes the following steps:
[0061] Step 401: In response to receiving a search request, extract entity information and faceted information from the search request.
[0062] Step 402: Search for target nodes that match entity information and faceted information in the pre-built knowledge graph.
[0063] Step 403: Obtain the source data for constructing the target node.
[0064] Step 404: Parse the semantics in the search request.
[0065] Step 405: Retrieve the semantically matched search results from the source data.
[0066] Steps 401-405 are basically the same as steps 201-205, so they will not be repeated here.
[0067] Step 406: Obtain a semantic ranking model trained based on samples from vertical domains.
[0068] In this embodiment, the initial semantic ranking model is a pre-trained model that can rank search results and keywords based on their semantic matching degree in certain domains. This application can retrain the initial semantic ranking model using samples from vertical domains, fine-tuning some parameters to enable the semantic ranking model to accurately identify the semantics of search results in vertical domains, thus obtaining an industry-migrated semantic ranking model. The input to the semantic ranking model is the keywords of the search request and multiple search results; the output is the ranking result based on semantic matching degree. The semantic ranking model can be an Aurora large model, utilizing its powerful semantic recognition and similarity calculation capabilities.
[0069] Based on specific industry sectors, it provides accurate retrieval of multimodal, fine-grained document fragmentation resources. During training, it integrates the Aurora large model and knowledge graph-driven sample augmentation semantic ranking technology to achieve accurate mixed ranking of multimodal, fine-grained fragmentation resources such as chapters, paragraphs, and tables. Unlike traditional keyword understanding, knowledge graph keyword understanding, in addition to basic recognition such as word segmentation and normalization, also applies content understanding technology to identify entity and faceted information in keywords, providing necessary feature inputs for subsequent retrieval and aggregation.
[0070] The filtering results are refined by using an industry semantic ranking model. Relying on the aforementioned "industry transfer technology", the semantic ranking model not only matches on a literal level, but also focuses on the semantic matching of keywords and content, thus more efficiently meeting the final requirements.
[0071] Step 407: Semantically rank the search results using a semantic ranking model.
[0072] In this embodiment, the original search results are from the general domain, which has poor comprehension of specialized documents in vertical industry sectors. A semantic ranking model can obtain fine-grained ranking results, ensuring effective search migration across different industries.
[0073] from Figure 4 It can be seen from this that, with Figure 2 Compared to the corresponding embodiments, the knowledge retrieval method process 400 in this embodiment reflects a multimodal, fine-grained document fragmentation resource retrieval and positioning function based on vertical industry fields, such as PDF document positioning and table positioning, which effectively improves the user retrieval experience and meets fine-grained retrieval needs.
[0074] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a knowledge retrieval device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0075] like Figure 5 As shown, the knowledge retrieval device 500 of this embodiment includes: an extraction unit 501, a search unit 502, an acquisition unit 503, a parsing unit 504, and a matching unit 505. The extraction unit 501 is configured to extract entity information and faceted information from the search request in response to receiving the search request; the search unit 502 is configured to search for target nodes matching the entity information and faceted information in a pre-constructed knowledge graph; the acquisition unit 503 is configured to acquire the source data for constructing the target nodes; the parsing unit 504 is configured to parse the semantics in the search request; and the matching unit 505 is configured to retrieve search results matching the semantics from the source data.
[0076] In this embodiment, the specific processing of the extraction unit 501, search unit 502, acquisition unit 503, parsing unit 504, and matching unit 505 of the knowledge retrieval device 600 can be referred to Figure 2 The corresponding steps are 201, 202, 203, 204 and 205 in the embodiment.
[0077] In some optional implementations of this embodiment, the search request includes keywords and at least one of the following search function options: comprehensive search, content search, rich summary search, resource search; and the device further includes a display unit configured to: select targeted search results from the search results according to the search function and display them.
[0078] In some optional implementations of this embodiment, the resource retrieval includes data formats of different modalities; and the display unit is further configured to: filter the retrieval results of the data format selected by the user from the retrieval results.
[0079] In some optional implementations of this embodiment, the device further includes an interpretation unit configured to output the theoretical basis for the relevance of the search results to the search request.
[0080] In some optional implementations of this embodiment, the device further includes an update unit configured to: update the knowledge graph in real time; and if the knowledge graph changes, retrieve the search request based on the changed information.
[0081] In some optional implementations of this embodiment, the device further includes a recommendation unit configured to: acquire the user's historical behavior information; and filter recommendation information for the user from the search results based on the historical behavior information.
[0082] In some optional implementations of this embodiment, the device further includes a classification unit configured to: filter and obtain classified search results from the search results based on the entity information and the faceted information.
[0083] In some optional implementations of this embodiment, the apparatus further includes a sorting unit configured to: acquire a semantic sorting model trained on samples from a vertical domain; and perform semantic sorting on the retrieval results using the semantic sorting model.
[0084] In some optional implementations of this embodiment, the device further includes a labeling unit configured to: label the most relevant search results by using box selection and color.
[0085] In some optional implementations of this embodiment, the device further includes a positioning unit configured to output a directory structure of the search results for users to preview and locate detailed content.
[0086] In some optional implementations of this embodiment, the apparatus further includes a table recall unit, configured to: extract at least one of the following table information from the searched table file: title, header, and content; and match the search request with the table information to obtain multi-granularity recall results.
[0087] In some optional implementations of this embodiment, the table recall unit is further configured to sort the recall results using a pre-built table data sorting model.
[0088] In some optional implementations of this embodiment, the table recall unit is further configured to: locate the block coordinates of the searched table file according to the semantics, and label the recall results.
[0089] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0090] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0091] An electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method described in process 200 or 400.
[0092] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described in process 200 or 400.
[0093] A computer program product includes a computer program that, when executed by a processor, implements the method described in process 200 or 400.
[0094] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0095] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0096] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0097] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as knowledge retrieval methods. For example, in some embodiments, the knowledge retrieval method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the knowledge retrieval method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the knowledge retrieval method by any other suitable means (e.g., by means of firmware).
[0098] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0099] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0100] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0101] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0102] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0103] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0104] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0105] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A knowledge retrieval method, comprising: In response to receiving a search request, entity information and faceted information are extracted from the keywords in the search request using a pre-trained natural language processing model. If the keywords do not contain faceted information, all faceted information involved in the entity information is used as the default faceted information for the user to select. If the user does not select specific faceted information, the search is performed based on all faceted information involved in the entity information. Search for target nodes that match the entity information and the faceted information in a pre-constructed knowledge graph, wherein each node in the knowledge graph represents an entity and each entity corresponds to faceted information; Obtaining the source data for constructing the target node includes: obtaining source data containing entity information and facet information; Parse the semantics in the search request; Retrieve search results that match the semantics from the source data, wherein the source data of the knowledge graph is labeled with semantic information.
2. The method according to claim 1, wherein, The search request includes keywords and at least one of the following search function options: comprehensive search, content search, rich summary search, and resource search; as well as The method further includes: The search function is used to filter and display targeted search results from the search results.
3. The method according to claim 2, wherein, The resource retrieval includes data formats of different modalities; as well as The step of filtering the search results from the search results using the search function to obtain accurate search results includes: The search results are filtered to select the data format chosen by the user from the search results.
4. The method according to claim 1, wherein, The method further includes: The theoretical basis for outputting the relevance between the search results and the search request.
5. The method according to claim 1, wherein, The method further includes: The knowledge graph is updated in real time; If the knowledge graph changes, the search request is retrieved based on the changed information.
6. The method according to claim 1, wherein, The method further includes: Obtain user's historical behavior information; Recommended information for the user is filtered from the search results based on the historical behavior information.
7. The method according to claim 1, wherein, The method further includes: Based on the entity information and the faceted information, categorized search results are obtained by filtering from the search results.
8. The method according to claim 1, wherein, The method further includes: Obtain a semantic ranking model trained based on samples from vertical domains; The search results are semantically ranked using the semantic ranking model.
9. The method according to claim 1, wherein, The method further includes: Highlight the most relevant search results by using selection boxes and color coding.
10. The method according to claim 1, wherein, The method further includes: Output the directory structure of the search results for users to preview and locate detailed content.
11. The method according to claim 1, wherein, The method further includes: Extract at least one of the following table information from the searched table files: title, header, and content; The search request is matched with the table information to obtain multi-granularity recall results.
12. The method according to claim 11, wherein, The method further includes: The recall results are sorted using a pre-built tabular data sorting model.
13. The method according to claim 11, wherein, The method further includes: Based on the semantics parsed from the search request, the searched table files are located by block coordinates, and the recall results are labeled.
14. A knowledge retrieval device, comprising: The extraction unit is configured to, in response to receiving a search request, extract entity information and faceted information from the keywords in the search request using a pre-trained natural language processing model. If the keywords do not contain faceted information, all faceted information involved in the entity information is provided as the default faceted information for the user to select. If the user does not select specific faceted information, the search is performed based on all faceted information involved in the entity information. The search unit is configured to search for target nodes that match the entity information and the faceted information in a pre-built knowledge graph, wherein each node in the knowledge graph represents an entity and each entity corresponds to faceted information; The acquisition unit is configured to acquire source data for constructing the target node, including: acquiring source data containing entity information and facet information; The parsing unit is configured to parse the semantics in the search request; The matching unit is configured to retrieve a search result that matches the semantics from the source data, wherein the source data of the knowledge graph is annotated with semantic information.
15. The apparatus according to claim 14, wherein, The search request includes keywords and at least one of the following search function options: comprehensive search, content search, rich summary search, resource search; and The device also includes a display unit configured to: The search function is used to filter and display targeted search results from the search results.
16. The apparatus according to claim 15, wherein, The resource retrieval includes data formats of different modalities; and The display unit is further configured to: The search results are filtered to select the data format chosen by the user from the search results.
17. The apparatus according to claim 14, wherein, The device further includes an interpretation unit configured to: The theoretical basis for outputting the relevance between the search results and the search request.
18. The apparatus according to claim 14, wherein, The device further includes an update unit configured to: The knowledge graph is updated in real time; If the knowledge graph changes, the search request is retrieved based on the changed information.
19. The apparatus according to claim 14, wherein, The device also includes a recommendation unit configured to: Obtain user's historical behavior information; Recommended information for the user is filtered from the search results based on the historical behavior information.
20. The apparatus according to claim 14, wherein, The device further includes a classification unit, configured to: Based on the entity information and the faceted information, categorized search results are obtained by filtering from the search results.
21. The apparatus according to claim 14, wherein, The device further includes a sorting unit, configured to: Obtain a semantic ranking model trained based on samples from vertical domains; The search results are semantically ranked using the semantic ranking model.
22. The apparatus according to claim 14, wherein, The device further includes a labeling unit, configured to: Highlight the most relevant search results by using selection boxes and color coding.
23. The apparatus according to claim 14, wherein, The device further includes a positioning unit configured to: Output the directory structure of the search results for users to preview and locate detailed content.
24. The apparatus according to claim 14, wherein, The device also includes a form recall unit, configured to: Extract at least one of the following table information from the searched table files: title, header, and content; The search request is matched with the table information to obtain multi-granularity recall results.
25. The apparatus according to claim 24, wherein, The table recall unit is further configured to: The recall results are sorted using a pre-built tabular data sorting model.
26. The apparatus according to claim 24, wherein, The table recall unit is further configured to: Based on the semantics parsed from the search request, the searched table files are located by block coordinates, and the recall results are labeled.
27. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
28. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.
29. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-13.
Citation Information
Patent Citations
Question answering interaction method and system based on intelligent robot
CN108959627A
Searching method and device, electronic equipment and storage medium
CN114036373A