Semantic retrieval method, apparatus, device, and storage medium

By identifying query intent and generating query statements, and combining entity dictionaries and knowledge graphs, the accuracy and efficiency of semantic retrieval are improved, solving the problem of insufficient accuracy of query statements in existing technologies.

CN116775823BActive Publication Date: 2026-02-10CHINA MERCHANTS BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310611049.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-26
Publication Date
2026-02-10
Estimated Expiration
2043-05-26

AI Technical Summary

Technical Problem

In existing semantic retrieval methods, the accuracy of query statements is low, resulting in insufficient retrieval accuracy.

Method used

By identifying the query intent of the statement to be retrieved, obtaining the target query attributes, determining the query entities that match the statement to be retrieved from the target entity dictionary, generating the query statement, and performing a retrieval in the target knowledge graph, the query statement generated by combining the target query attributes and query entities can directly reflect the query intent of the statement to be retrieved.

Benefits of technology

It improves the accuracy of query statements, enhances the accuracy and efficiency of semantic retrieval, and solves the problem of low accuracy in existing semantic retrieval methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116775823B_ABST
    Figure CN116775823B_ABST
Patent Text Reader

Abstract

The application discloses a semantic retrieval method and device, equipment and storage medium, relate to semantic retrieval technical field, method includes: obtaining the sentence to be retrieved; identifying the query intention of the sentence to be retrieved, obtaining the target query attribute; determining the first query entity matched with the sentence to be retrieved from the target entity dictionary; wherein, all entities in the target entity dictionary are extracted from a plurality of natural language sentences in a target field; generating the first query sentence based on the target query attribute and the first query entity; based on the first query sentence, retrieval is carried out in the target knowledge graph, and the first query result of the sentence to be retrieved is obtained; wherein, the target knowledge graph is constructed based on a plurality of natural language sentences. The application makes the accuracy of the query sentence higher, thereby improving the accuracy of semantic retrieval, and solves the technical problem of low accuracy of the existing semantic retrieval method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic retrieval technology, and in particular to a semantic retrieval method, apparatus, device, and storage medium. Background Technology

[0002] In related technologies, semantic retrieval methods typically use single language features such as word segmentation and recognition, similarity scoring models, and pre-trained language models like BERT (Bidirectional Encoder Representations from Transformers) to process the retrieval statement. The resulting query statements have low accuracy, leading to low semantic retrieval accuracy. Summary of the Invention

[0003] The main objective of this invention is to provide a semantic retrieval method, apparatus, device, and storage medium, aiming to solve the technical problem of low accuracy in existing semantic retrieval methods.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] In a first aspect, the present invention provides a semantic retrieval method, the method comprising:

[0006] Retrieve the search query;

[0007] Identify the query intent of the statement to be retrieved and obtain the target query attributes;

[0008] The first query entity that matches the query statement is determined from the target entity dictionary; wherein all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain;

[0009] Based on the target query attribute and the first query entity, generate the first query statement;

[0010] Based on the first query statement, a search is performed in the target knowledge graph to obtain the first query result of the query statement; wherein, the target knowledge graph is constructed based on multiple natural language statements.

[0011] Optionally, after identifying the query intent of the statement to be retrieved and obtaining the target query attributes, the method further includes:

[0012] Entity extraction is performed on the search statement to extract the second query entity;

[0013] Generate a second query statement based on the target query attribute and the second query entity;

[0014] Based on the second query statement, a search is performed in the target knowledge graph to obtain the second query result of the query statement to be searched.

[0015] If the second query result is empty, then the step of determining the first query entity that matches the query statement from the target entity dictionary is executed.

[0016] Optionally, the first query entity that matches the query statement is determined from the target entity dictionary, including:

[0017] Identify multiple ambiguous entities from the target entity dictionary that match the query statement;

[0018] Identify the first query entity from multiple ambiguous entities that has the same category as the second query entity.

[0019] Optionally, a first query entity of the same category as the second query entity is determined from a plurality of ambiguous entities, including:

[0020] If multiple ambiguous entities include multiple ambiguous entities of the same category as the second query entity, then the first query entity with the highest similarity to the query statement is determined from the preset ambiguous entity dictionary.

[0021] Optionally, the first query entity that matches the query statement is determined from the target entity dictionary, including:

[0022] Identify multiple nested entities from the target entity dictionary that match the query statement;

[0023] Identify the first query entity with the longest entity length from multiple nested entities.

[0024] Optionally, after extracting the second query entity from the query statement, the method further includes:

[0025] Determine if the second query entity exists in the target entity dictionary;

[0026] If the second query entity does not exist in the target entity dictionary, then at least one similar natural language statement that meets the preset conditions for similarity with the statement to be retrieved is determined from multiple natural language statements.

[0027] Based on at least one similar natural language statement and the statement to be retrieved, obtain the target extended triplet;

[0028] Based on the target extended triples, generate an extended query statement;

[0029] The extended query statement is added to the target knowledge graph, and the extended entities in the target extended triples are added to the target entity dictionary, resulting in the extended knowledge graph and the extended entity dictionary.

[0030] Optionally, based on at least one similar natural language statement and the statement to be retrieved, a target extended triple is obtained, including:

[0031] Based on a preset splicing template, at least one similar natural language statement and the statement to be retrieved are spliced ​​together to obtain at least one extended language statement;

[0032] Extract at least one extended triple from at least one extended language statement;

[0033] Identify the target extended triplet that is closest to the query statement from at least one extended triplet.

[0034] Secondly, the present invention also provides a semantic retrieval device, the device comprising:

[0035] The retrieval module is used to retrieve the search query.

[0036] The intent recognition module is used to identify the query intent of the query statement to be retrieved and obtain the target query attributes;

[0037] The entity matching module is used to determine the first query entity that matches the query statement from the target entity dictionary; wherein, all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain.

[0038] The statement generation module is used to generate the first query statement based on the target query attribute and the first query entity;

[0039] The retrieval module is used to perform a retrieval in the target knowledge graph based on the first query statement to obtain the first query result of the query statement; wherein, the target knowledge graph is constructed based on multiple natural language statements.

[0040] Thirdly, the present invention also provides a semantic retrieval device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of any of the semantic retrieval methods described above.

[0041] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the semantic retrieval method as described above.

[0042] This invention provides a semantic retrieval method, apparatus, device, and storage medium, comprising: acquiring a statement to be retrieved; identifying the query intent of the statement to be retrieved to obtain target query attributes; determining a first query entity that matches the statement to be retrieved from a target entity dictionary; wherein all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain; generating a first query statement based on the target query attributes and the first query entity; and performing a retrieval in a target knowledge graph based on the first query statement to obtain a first query result for the statement to be retrieved; wherein the target knowledge graph is constructed based on multiple natural language statements.

[0043] Therefore, this invention identifies the query intent of the query statement to be retrieved, obtains the target query attributes, and generates a query statement by combining the query entities that match the query statement from the target entity dictionary. The query results of the query statement are obtained by searching in the target knowledge graph. The query statement generated by combining the target query attributes and query entities can directly reflect the query intent of the query statement, making the query statement more accurate and thus improving the accuracy of semantic retrieval. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of the semantic retrieval device of the present invention;

[0046] Figure 2 This is a flowchart illustrating the first embodiment of the semantic retrieval method of the present invention;

[0047] Figure 3 This is a flowchart illustrating the second embodiment of the semantic retrieval method of the present invention;

[0048] Figure 4 This is a flowchart illustrating the third embodiment of the semantic retrieval method of the present invention;

[0049] Figure 5 This is a schematic diagram of the modules of the first embodiment of the semantic retrieval device of the present invention.

[0050] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0052] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0053] In this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that an apparatus or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an apparatus or system. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the apparatus or system that includes that element.

[0054] A knowledge graph is a graph structure for storing information, composed of numerous triples, each consisting of a head entity, a tail entity, and relationships between entities. Through the structured representation of data using knowledge graphs, computers can better "understand" human knowledge and perform knowledge reasoning. Typically, knowledge graphs are built upon natural language text within a specific professional domain. The construction process mainly includes data preprocessing and data storage. First, unstructured data in the natural language text undergoes entity recognition, entity disambiguation, and relation extraction to transform the natural language text of that professional domain into a structured knowledge graph, which is then stored in a database such as Neo4j, OrientDB, and JanusGraph. To implement functions such as machine question answering and intelligent information retrieval using knowledge graphs, Natural Language Processing (NLP) technology is needed to help machines recognize the semantics of human questions, transforming the questions into query languages ​​(QL) with predefined syntax and structure that can be "understood" by machines, such as SPARQL, CypherQL, and FunQL.

[0055] However, in existing knowledge graph-based retrieval methods, most query statements are generated by semantic processing of user questions using single word segmentation recognition, similarity scoring models, BERT models, etc., resulting in low accuracy of query statements and consequently low accuracy of retrieval methods.

[0056] Furthermore, the coverage of entities and relationships in the knowledge graph also affects retrieval accuracy. If the graph does not store relevant entities, then it will be impossible to retrieve information related to that entity. This is because, when extracting entities from unstructured natural language text, the lack of labeled samples to train the extraction model makes it difficult for the model to identify all entities of interest to the user, resulting in an incomplete knowledge graph and inaccurate retrieval results.

[0057] In view of the technical problem of low accuracy in existing semantic retrieval methods, this invention provides a semantic retrieval method, the overall idea of ​​which is as follows:

[0058] The method includes: obtaining a query statement to be retrieved; identifying the query intent of the query statement to obtain target query attributes; determining a first query entity that matches the query statement from a target entity dictionary; wherein all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain; generating a first query statement based on the target query attributes and the first query entity; and performing a search in a target knowledge graph based on the first query statement to obtain the first query result of the query statement; wherein the target knowledge graph is constructed based on multiple natural language statements.

[0059] This invention provides a semantic retrieval method that identifies the query intent of the statement to be retrieved, obtains the target query attributes, and generates a query statement by combining the query entities that match the statement to be retrieved from the target entity dictionary. The query results are then obtained by searching the target knowledge graph. The query statement generated by combining the target query attributes and query entities can directly reflect the query intent of the statement to be retrieved, resulting in higher accuracy of the query statement. This improves the accuracy of semantic retrieval and solves the technical problem of low accuracy in existing semantic retrieval methods.

[0060] The semantic retrieval method, apparatus, device, and storage medium used in the technical implementation of this invention will be described in detail below:

[0061] Reference Figure 1 , Figure 1 This is a schematic diagram of the semantic retrieval device of the present invention;

[0062] like Figure 1 As shown, the device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include user devices such as mobile phones, smartphones, and computers; optionally, the user interface 1003 may also include standard wired interfaces and wireless interfaces. The network interface 1004 may optionally include standard wired interfaces and wireless interfaces (such as Wireless-Fidelity (Wi-Fi) interfaces). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0063] Those skilled in the art will understand that Figure 1The structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0064] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a semantic retrieval program.

[0065] exist Figure 1 In the device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with user devices; the processor 1001 and the memory 1005 in the semantic retrieval method of the present invention can be set in the device. The semantic retrieval method calls the semantic retrieval program stored in the memory 1005 through the processor 1001 and executes the semantic retrieval method provided in the embodiment of the present invention.

[0066] The semantic retrieval method, apparatus, device, and storage medium of the present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0067] Based on, but not limited to, the above hardware structure, refer to Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the semantic retrieval method of the present invention. This embodiment provides a semantic retrieval method, which may include:

[0068] Step S100: Obtain the query statement to be searched.

[0069] In this embodiment, the executing entity is as follows: Figure 1 The semantic retrieval device shown can be a physical server consisting of a standalone host, or a virtual server hosted in a host cluster. The query to be retrieved can be a question sent by a user device.

[0070] Step S200: Identify the query intent of the statement to be retrieved and obtain the target query attributes.

[0071] In this embodiment, the query intent can be the utterance intent of the statement to be retrieved. The query intent of the statement to be retrieved can be identified using a fine-tuning classification model of the pre-trained language model BERT or a BiLSTM (Bi-directional Long Short-Term Memory)-CRF (Conditional Random Field) model. For example, if the statement to be retrieved is "What is the interest rate of project XX?", then the query intent is "interest rate".

[0072] Step S600: Determine the first query entity that matches the statement to be retrieved from the target entity dictionary; wherein, all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain.

[0073] In this embodiment, the target domain is determined based on the specific usage scenario. Multiple natural language statements can be all natural language statements in the natural language database of that target domain. The target entity dictionary can perform information extraction (IE) on the multiple natural language statements, extracting named entities (entities identified by names, such as person names, organization names, and place names) from them, and storing all named entities in a trie format according to the category of entity names. The first query entity can be an identical entity in the target entity dictionary that is completely identical to the entity to be retrieved in the query statement. This can typically be obtained by hard matching the query statement against all entities in the target entity dictionary, or it can be the most similar entity in the target entity dictionary to the entity to be retrieved. This can typically be obtained by similarity matching the query statement against all entities in the target entity dictionary.

[0074] Step S700: Generate the first query statement based on the target query attribute and the first query entity.

[0075] In this embodiment, taking the query statement as "What is the interest rate of XX project?" as an example, if the first query entity is "XX project" or "YY project" which has the highest similarity to "XX project", then the first query statement can be generated based on "XX project" and "interest rate", or it can be generated based on "YY project" and "interest rate".

[0076] Step S800: Based on the first query statement, perform a search in the target knowledge graph to obtain the first query result of the query statement; wherein, the target knowledge graph is constructed based on multiple natural language statements.

[0077] In this embodiment, the target knowledge graph can be constructed as follows. First, information extraction (IE) is performed on multiple natural language statements to extract named entities (entities identified by names, such as people, organizations, and places), attributes, and relationships, which are then presented as triples. The structure of a triple is (s, p, o), where s represents the subject entity (which can be one of the named entities mentioned above), o represents the object entity (which can be one of the attributes mentioned above), and p represents the predicate relationship between the two entities (which can be one of the relationships mentioned above). Extracting named entities, attributes, and relationships from multiple natural language statements can be achieved using methods such as CNN (Convolutional Neural Network) or fine-tuning the pre-trained language model BERT. Based on the extracted triples, corresponding QL (Query Language) statements are generated and stored in a database in the form of a graph, thus obtaining the knowledge graph.

[0078] It's understandable that when a knowledge graph involves many knowledge categories, multiple sub-knowledge graphs can be constructed based on different scenarios. For example, a knowledge graph about financial products can be divided into multiple sub-knowledge graphs based on categories, such as stocks, bonds, futures, options, and funds, and further categorized as short-term products, long-term products, high-risk products, and low-risk products. The entities stored in each sub-graph can overlap. When performing a query, the search can also be conducted within the corresponding sub-knowledge graph based on the category involved in the query, improving search efficiency.

[0079] Specifically, taking the query statement "What is the interest rate of Project XX?" as an example, if the first query entity is "Project XX" or "Project YY" which has the highest similarity to "Project XX", then the first query statement can be generated based on "Project XX" and "interest rate", or it can be generated based on "Project YY" and "interest rate". Based on the first query statement, a search is performed in the target knowledge graph, and the first query result can be "The interest rate of Project XX is a%" or "The interest rate of Project YY is b%".

[0080] This embodiment provides a semantic retrieval method. By identifying the query intent of the statement to be retrieved, target query attributes are obtained. These attributes are then combined with query entities from the target entity dictionary that match the statement to be retrieved to generate a query statement. The query results are obtained by searching the target knowledge graph. The query statement generated by combining the target query attributes and query entities directly reflects the query intent of the statement to be retrieved, resulting in higher accuracy and thus improving the accuracy of semantic retrieval. This solves the technical problem of low accuracy in existing semantic retrieval methods. Furthermore, in this embodiment, entities in the target entity dictionary are stored in a trie structure, which improves the efficiency of matching the first query entity, making the semantic retrieval method more efficient.

[0081] Furthermore, referring to Figure 3 , Figure 3 This is a flowchart illustrating a second embodiment of the semantic retrieval method of the present invention. As one implementation, after step S200, the method may further include:

[0082] Step S300: Extract entities from the query statement to extract the second query entity.

[0083] Step S400: Generate a second query statement based on the target query attribute and the second query entity.

[0084] Step S500: Based on the second query statement, perform a search in the target knowledge graph to obtain the second query result of the query statement to be searched.

[0085] Step S600 may include:

[0086] Step S600a: If the second query result is empty, determine the first query entity that matches the query statement from the target entity dictionary.

[0087] In this embodiment, the second query entity of the query statement can be directly extracted from the query statement to be retrieved. This second query statement is then generated by combining the target query attributes and retrieved from the target knowledge graph to obtain the second query result. If the second query result is not found in the target knowledge graph when combining the second query entity and the target query intent, the first query entity is then matched from the target entity dictionary to improve the coverage and accuracy of semantic retrieval. Entity extraction can be implemented using the pre-trained language model BERT.

[0088] As a specific implementation, step S600 may include: determining multiple ambiguous entities from the target entity dictionary that match the query statement; and determining a first query entity from the multiple ambiguous entities that is of the same category as the second query entity.

[0089] In this embodiment, when the second query result is empty and the target entity dictionary includes multiple ambiguous entities that match the query statement, the entity categories of the multiple ambiguous entities can be compared with the entity category of the first query entity. Based on the entity categories of the ambiguous entities and the entity categories of the second query entity, the multiple ambiguous entities can be disambiguated to obtain the first query entity among the multiple ambiguous entities that has the same category as the second query entity.

[0090] Continuing with the query statement "What is the interest rate of Project XX?", if the second query entity is "Project XX" and its entity category is bonds, there are ambiguous entities: "Project XX" entity category is stocks, "Project XX" entity category is funds, and "Project XX" entity category is bonds. Therefore, the first query entity is "Project XX" entity category is bonds.

[0091] Specifically, determining the first query entity that is of the same category as the second query entity from multiple ambiguous entities may include: if the multiple ambiguous entities include multiple ambiguous entities of the same category as the second query entity, then determining the first query entity with the highest similarity to the statement to be retrieved from a preset ambiguous entity dictionary.

[0092] In this embodiment, the preset ambiguous entities in the preset ambiguous entity dictionary may include the ambiguous entities of all entities in the target entity dictionary, and the preset ambiguous entity dictionary can be pre-constructed based on all entities in the target entity dictionary.

[0093] When the entity categories of multiple ambiguous entities are different from the entity category of the second query entity, and comparing the entity categories of the ambiguous entities with the entity category of the second query entity cannot disambiguate the multiple ambiguous entities, then the query statement to be retrieved can be matched with all the preset ambiguous entities in the preset ambiguous entity dictionary for similarity, and the first query entity with the highest similarity to the query statement to be retrieved can be determined from the preset ambiguous entity dictionary.

[0094] In another specific implementation, step S600 may include: determining multiple nested entities from the target entity dictionary that match the query statement; and determining the first query entity with the longest entity length from the multiple nested entities.

[0095] In this embodiment, the query statement is matched with all entities in the target entity dictionary. If the target entity dictionary includes multiple nested entities of different lengths that match the query statement, the nested entity with the longest entity length among the multiple nested entities can be used as the first query entity based on the entity length.

[0096] For example, if multiple nested entities include "bond coupon rate" and "interest rate", then "bond coupon rate" will be the first query entity.

[0097] This embodiment provides a semantic retrieval method that can directly extract query entities from the query statement to be retrieved and combine them with target query attributes to generate a query statement and retrieve query results in a knowledge graph. The query statement generated by combining the target query attributes and query entities can directly reflect the query intent of the query statement to be retrieved, resulting in higher accuracy of the query statement. This improves the accuracy of semantic retrieval and solves the technical problem of low accuracy in existing semantic retrieval methods.

[0098] Furthermore, referring to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the semantic retrieval method of the present invention. As another implementation, after step S300, the method may further include:

[0099] Step S900: Determine whether the second query entity exists in the target entity dictionary.

[0100] In this embodiment, when it is determined that the second query entity does not exist in the target entity dictionary, the target knowledge graph can be expanded based on the search statement to continuously improve the completeness of the target knowledge graph.

[0101] Step S1000: If the second query entity does not exist in the target entity dictionary, then determine at least one similar natural language statement from multiple natural language statements that meets the preset conditions for similarity with the statement to be retrieved.

[0102] In this embodiment, the target knowledge graph can be expanded by combining the search query and multiple natural language statements. Typically, the search query can be matched with multiple natural language statements to obtain at least one similar natural language statement whose similarity score is higher than a preset condition.

[0103] Step S1100: Obtain the target extended triplet based on at least one similar natural language statement and the statement to be retrieved.

[0104] In this embodiment, after obtaining at least one similar natural language statement, the at least one similar natural language statement can be combined with the statement to be retrieved to extract at least one extended triplet, and the target extended triplet can be determined from the at least one extended triplet to generate an extended query statement.

[0105] Specifically, as one implementation, step S1100 includes: obtaining a splicing template based on the statement to be retrieved; splicing at least one similar natural language statement and the statement to be retrieved based on the splicing template to obtain at least one extended language statement; extracting at least one extended triplet from the at least one extended language statement; and determining the target extended triplet that is closest to the statement to be retrieved from the at least one extended triplet.

[0106] In this embodiment, the splicing template can be a prompt template, determined based on the retrieval statement. Extended triples can be obtained by extracting information from the extended language statement using a pre-trained language model (such as BERT-based). Then, a dual affine pointer network can be used to identify the nesting relationships among all entities in at least one extended triple, obtaining the target entity closest to the retrieval statement, and determining the target extended triple closest to the retrieval statement from at least one extended triple.

[0107] Step S1200: Generate an expanded query statement based on the target expanded triplet.

[0108] Step S1300: Add the extended query statement to the target knowledge graph and add the extended entities in the target extended triples to the target entity dictionary to obtain the extended knowledge graph and the extended entity dictionary.

[0109] In this embodiment, by adding the extended query statement to the target knowledge graph, the target knowledge graph can be expanded. Furthermore, by adding the extended entities from the target extended triples to the target entity dictionary, the target entity dictionary can be expanded accordingly. The extended entity is the main entity in the target extended triple.

[0110] This embodiment provides a semantic retrieval method. When the target entity dictionary does not include the query entity of the search statement, it extracts extended triples by combining the search statement and the natural language statement to generate an extended query statement. This extends the target entity dictionary and the target knowledge graph respectively. During the semantic retrieval process, the target entity dictionary and the target knowledge graph can be continuously enriched, so that the retrieval scope of the knowledge graph can automatically adapt to the retrieval needs, thereby improving the completeness of the knowledge graph.

[0111] Furthermore, this embodiment eliminates the need to design new prompt templates. The query statement is directly used as the prompt template and concatenated with similar natural language statements. The resulting expanded language statement requires no further fine-tuning and can be directly input into the pre-trained language model to extract expanded triples. This overcomes the few-sample constraint, achieving better pre-trained language model performance in few-sample scenarios and better unifying the pre-training task with other downstream tasks. A dual affine pointer network is used to perform nested recognition on the extracted triples, eliminating the influence of nested entities and making the expanded triples more closely match the query statement, thus improving the accuracy of graph expansion and dictionary expansion.

[0112] Based on the same inventive concept, embodiments of the present invention also provide a semantic retrieval device, referring to... Figure 5 , Figure 5 This is a schematic diagram of the modules of the first embodiment of the semantic retrieval device of the present invention; the device may include:

[0113] Module 10 is used to obtain the search query.

[0114] The intent recognition module 20 is used to identify the query intent of the query statement to be retrieved and obtain the target query attributes;

[0115] The entity matching module 30 is used to determine the first query entity that matches the statement to be retrieved from the target entity dictionary; wherein, all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain;

[0116] The statement generation module 40 is used to generate a first query statement based on the target query attribute and the first query entity;

[0117] The retrieval module 50 is used to perform a retrieval in the target knowledge graph based on the first query statement to obtain the first query result of the query statement to be retrieved; wherein, the target knowledge graph is constructed based on multiple natural language statements.

[0118] Furthermore, as one embodiment, the device may further include:

[0119] The entity extraction module is used to extract entities from the search statement and extract the second query entity.

[0120] The statement generation module 40 is also used to generate a second query statement based on the target query attribute and the second query entity;

[0121] The retrieval module 50 is also used to generate a second query statement based on the target query attribute and the second query entity;

[0122] The entity matching module 30 is used to determine the first query entity that matches the query statement from the target entity dictionary if the second query result is empty.

[0123] Furthermore, as another embodiment, the device may further include:

[0124] The feedback module is used to determine whether a second query entity exists in the target entity dictionary; if the second query entity does not exist in the target entity dictionary, at least one similar natural language statement that meets the preset conditions for similarity with the statement to be retrieved is determined from multiple natural language statements; based on at least one similar natural language statement and the statement to be retrieved, the target extended triplet is obtained.

[0125] The statement generation module 40 is also used to generate extended query statements based on target extended triples; add the extended query statements to the target knowledge graph, and add the extended entities in the extended triples to the target entity dictionary to obtain the extended knowledge graph and the extended entity dictionary.

[0126] For more details on the specific implementation of the semantic retrieval device described above, please refer to the description of the specific implementation of the semantic retrieval method in any one of Embodiments 1 to 3 above. For the sake of brevity, these details will not be repeated here.

[0127] Furthermore, embodiments of the present invention also propose a computer storage medium storing a computer program, which, when executed by a processor, implements the steps of the semantic retrieval method described above. Therefore, further details will not be repeated here. Additionally, the beneficial effects of employing the same method will not be repeated. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed to execute on a single computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0128] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0129] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A semantic retrieval method, characterized in that, The method includes: Retrieve the search query; Identify the query intent of the statement to be retrieved and obtain the target query attributes; The first query entity that matches the query statement is determined from the target entity dictionary; wherein, all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain; Based on the target query attribute and the first query entity, a first query statement is generated; Based on the first query statement, a search is performed in the target knowledge graph to obtain the first query result of the query statement; wherein, the target knowledge graph is constructed based on multiple natural language statements; After identifying the query intent of the query statement to be retrieved and obtaining the target query attributes, the method further includes: Entity extraction is performed on the statement to be retrieved to extract the second query entity; Determine whether the second query entity exists in the target entity dictionary; If the second query entity does not exist in the target entity dictionary, then at least one similar natural language statement that satisfies the preset condition of similarity with the statement to be retrieved is determined from the multiple natural language statements. Based on at least one of the similar natural language statements and the statement to be retrieved, a target extended triplet is obtained; Based on the target extended triplet, an extended query statement is generated; The extended query statement is added to the target knowledge graph, and the extended entities in the target extended triples are added to the target entity dictionary to obtain the extended knowledge graph and the extended entity dictionary.

2. The method as described in claim 1, characterized in that, After extracting the second query entity from the entity in the query statement to be retrieved, the method further includes: Based on the target query attribute and the second query entity, a second query statement is generated; Based on the second query statement, a search is performed in the target knowledge graph to obtain the second query result of the query statement to be searched. If the second query result is empty, then the step of determining the first query entity that matches the query statement from the target entity dictionary is executed.

3. The method as described in claim 2, characterized in that, The step of determining the first query entity that matches the search query from the target entity dictionary includes: Identify multiple ambiguous entities from the target entity dictionary that match the query statement; The first query entity that is of the same category as the second query entity is determined from among the multiple ambiguous entities.

4. The method as described in claim 3, characterized in that, The step of determining the first query entity that is of the same category as the second query entity from the plurality of ambiguous entities includes: If multiple ambiguous entities include multiple ambiguous entities of the same category as the second query entity, then the first query entity with the highest similarity to the query statement is determined from the preset ambiguous entity dictionary.

5. The method as described in claim 1, characterized in that, The step of determining the first query entity that matches the search query from the target entity dictionary includes: Identify multiple nested entities from the target entity dictionary that match the query statement; The first query entity with the longest entity length is determined from the multiple nested entities.

6. The method as described in claim 1, characterized in that, The step of obtaining the target extended triplet based on at least one of the similar natural language statements and the statement to be retrieved includes: Based on a preset splicing template, at least one of the similar natural language statements and the statement to be retrieved are spliced ​​together to obtain at least one extended language statement; Extract at least one extended triple from at least one extended language statement; The target extended triplet that is closest to the query statement is determined from at least one of the extended triplets.

7. A semantic retrieval device, characterized in that, The device includes: The retrieval module is used to retrieve the search query. The intent recognition module is used to identify the query intent of the query statement to be retrieved and obtain the target query attributes; An entity matching module is used to determine a first query entity that matches the statement to be retrieved from a target entity dictionary; wherein all entities in the target entity dictionary are extracted from multiple natural language statements in the target domain; The statement generation module is used to generate a first query statement based on the target query attribute and the first query entity; The retrieval module is used to perform a retrieval in the target knowledge graph based on the first query statement to obtain a first query result of the query statement to be retrieved; wherein, the target knowledge graph is constructed based on multiple natural language statements; The apparatus is further configured to: extract entities from the statement to be retrieved, extract a second query entity; determine whether the second query entity exists in the target entity dictionary; if the second query entity does not exist in the target entity dictionary, determine at least one similar natural language statement from the plurality of natural language statements that satisfies a preset similarity condition to the statement to be retrieved; obtain a target extended triplet based on at least one similar natural language statement and the statement to be retrieved; generate an extended query statement based on the target extended triplet; add the extended query statement to the target knowledge graph, and add the extended entities in the target extended triplet to the target entity dictionary, thereby obtaining an extended knowledge graph and an extended entity dictionary.

8. A semantic retrieval device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, configured by the computer program to implement the steps of the semantic retrieval method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the semantic retrieval method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Domain knowledge graph question answering method and system based on query path sorting

    CN115982338A