Information retrieval method and device

By finding semantic-related text block triplets in the intelligent question-answer system and obtaining the search results corresponding to the search text, the problem that LLM generates answers is not related to user questions is solved, and the relevance and answering ability of the answers is improved.

CN120216660APending Publication Date: 2025-06-27BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510151103.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the field of intelligent question and answers, the answers generated by LLM are sometimes unrelated to the user's questions and cannot answer the user's questions, which in turn cannot meet the user's needs.

Method used

By finding semantic related triples in triples in multiple text blocks, obtaining search results for search text, improving the integrity and relevance of text blocks used by the LLM model, thereby improving the correlation between answers and user's questions.

Benefits of technology

This improves the correlation between the answers generated by LLM and the user's questions, reduces the illusion phenomenon of the LLM model, and increases the probability that the LLM generation answers can answer user's questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216660A_ABST
    Figure CN120216660A_ABST
Patent Text Reader

Abstract

The invention provides an information retrieval method and device. And searching a first triple related to the retrieval text semantics in the triple of each text block in the plurality of text blocks. The triad comprises two entity words in the text block to which the triad belongs and an association relationship between the two entity words. And searching a text block with a second triad in the plurality of text blocks, the second triad comprising at least one of the following: an entity word in the first triad and an entity word having an association relationship with the entity word in the first triad. And obtaining a retrieval result corresponding to the retrieval text according to the text block with the second triple. By inputting the retrieval text and the retrieval result corresponding to the retrieval text into the LLM model, the integrity and relevance of the text block used by the LLM model when the LLM model is assisted to generate the answer can be improved, so that the relevance between the answer generated by the LLM model and the question of the user can be improved, the illusion phenomenon of the LLM model is reduced, and the user experience is improved. And the probability that the answer generated by the LLM can answer the question of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an information retrieval method and device. Background Art

[0002] LLM (Large Language Model) is an artificial intelligence model designed to understand and generate human language. It learns the ability to serve human language understanding and generation by training on a large amount of text data, and can perform a wide range of tasks, including text summarization, translation, intelligent question answering, and sentiment analysis, etc. The core idea of LLM is to learn the patterns and language structures of natural language through large-scale unsupervised training, which can to a certain extent simulate the human language cognition and generation process. LLM can better understand and generate natural text, and at the same time can also show certain logical thinking and reasoning abilities.

[0003] However, in the field of intelligent question answering, the answers generated by LLM are sometimes irrelevant to the user's question, unable to answer the user's question, and thus unable to meet the user's needs. Summary of the Invention

[0004] This application discloses an information retrieval method and device.

[0005] In a first aspect, this application discloses an information retrieval method, the method comprising:

[0006] Obtaining a retrieval text;

[0007] In the triples of each text block among multiple text blocks, searching for a first triple semantically related to the retrieval text; a text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words;

[0008] According to the first triple, searching for a second triple in the triples of each text block among multiple text blocks, the second triple including at least one of the following: the entity words in the first triple, and the entity words having an association relationship with the entity words in the first triple;

[0009] In the multiple text blocks, searching for the text blocks having the second triple;

[0010] According to the text blocks having the second triple, obtaining a retrieval result corresponding to the retrieval text.

[0011] In a second aspect, this application discloses an information retrieval device, the device comprising:

[0012] A first obtaining module, configured to obtain a retrieval text;

[0013] A first search module, configured to search for a first triple semantically related to the retrieved text from among the triples of each text block in multiple text blocks; a text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words;

[0014] A second search module, configured to search for a second triple from among the triples of each text block in multiple text blocks according to the first triple, where the second triple includes at least one of the following: the entity words in the first triple, and entity words having an association relationship with the entity words in the first triple;

[0015] A third search module, configured to search for the text blocks having the second triple from among the multiple text blocks;

[0016] A second obtaining module, configured to obtain a retrieval result corresponding to the retrieved text according to the text blocks having the second triple.

[0017] In a third aspect, the present application discloses an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the method according to any of the above aspects.

[0018] In a fourth aspect, the present application discloses a non-transitory computer-readable storage medium, which, when the instructions stored therein are executed by a processor of an electronic device, enables the electronic device to execute the method according to any of the above aspects.

[0019] In a fifth aspect, the present application discloses a computer program product, which, when the instructions in the computer program product are executed by a processor of an electronic device, enables the electronic device to execute the method according to any of the above aspects.

[0020] The technical solution provided by the present application may include the following beneficial effects:

[0021] In the present application, the retrieved text is obtained. A first triple semantically related to the retrieved text is searched for from among the triples of each text block in multiple text blocks. A text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words. According to the first triple, a second triple is searched for from among the triples of each text block in multiple text blocks, and the second triple includes at least one of the following: the entity words in the first triple, and entity words having an association relationship with the entity words in the first triple. The text blocks having the second triple are searched for from among the multiple text blocks. A retrieval result corresponding to the retrieved text is obtained according to the text blocks having the second triple.

[0022] In this application, the first triple is semantically related to the retrieval text. Since the second triple includes at least one of the following: the entity word in the first triple, and the entity word having an association relationship with the entity word in the first triple, therefore, the second triple is also semantically related to the retrieval text. There can be multiple text blocks having the second triple, and there is an association relationship in content among the multiple text blocks having the second triple. Thus, the multiple text blocks having the second triple can mutually confirm and complement each other in content. Thus, the content in the retrieval result corresponding to the retrieval text obtained from the multiple text blocks having the second triple (the retrieval result includes at least two text blocks having the second triple) is also related. Thus, it can play a role of mutual confirmation and complementation. In this way, inputting the retrieval text and the retrieval result corresponding to the retrieval text into the LLM model can improve the integrity and relevance of the text blocks used by the LLM model when assisting in generating answers, thereby improving the relevance between the answer generated by the LLM and the user's question, reducing the hallucination phenomenon of the LLM model, and further increasing the probability that the answer generated by the LLM can answer the user's question.

[0023] Especially when facing complex questions that require reasoning and summarization, or answers that need to traverse multiple content-related text blocks to provide comprehensive insights, it can more significantly improve the relevance between the answer generated by the LLM and the user's question, can more significantly reduce the hallucination phenomenon of the LLM model, and further more significantly increase the probability that the answer generated by the LLM can answer the user's question. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a schematic diagram of an information retrieval architecture of this application.

[0025] Figure 2 is a flowchart of the steps of an information retrieval method of this application.

[0026] Figure 3 is a schematic diagram of a knowledge graph of this application.

[0027] Figure 4 is a block diagram of the structure of an information retrieval device of this application.

[0028] Figure 5 is a block diagram of an electronic device of this application.

[0029] Figure 6 is a block diagram of an electronic device of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] In one example, in the field of intelligent question answering, the LLM can generate an answer for answering the user's question according to the user's question and provide the answer to the user.

[0032] However, the current LLM lacks new knowledge and domain-specific knowledge. As a result, when the user's question involves new knowledge or domain-specific knowledge, the answer generated by the LLM is sometimes irrelevant to the user's question, unable to answer the user's question, and thus unable to meet the user's needs.

[0033] For this problem, in one approach, RAG (Retrieval Augmented Generation) can be used for assistance.

[0034] Among them, retrieval augmented generation can obtain keywords for characterizing the theme of the user's question, then retrieve text chunks on the network according to the keywords, and then input the user's question and the retrieved text chunks into the LLM, so that the LLM generates an answer for answering the user's question according to the user's question and the retrieved text chunks and provides the generated answer to the user.

[0035] The text chunks retrieved by RAG contain the keywords. In this way, it is possible to reflect the knowledge of the theme for characterizing the user's question. In this case, the LLM can learn the relevant knowledge of the theme for characterizing the user's question according to the text chunks retrieved by RAG, thereby increasing the probability that the answer generated by the LLM is relevant to the user's question, and thus increasing the probability that the answer generated by the LLM can answer the user's question, and further increasing the probability that the answer generated by the LLM can meet the user's needs.

[0036] Among them, the current mainstream RAG frameworks include LangChain (language model integration framework), etc.

[0037] LangChain splits the documents on the network into multiple text chunks according to a specific character length. For example, the documents on the network are split into multiple text chunks according to a 250-character length.

[0038] Among them, in one approach, when LangChain retrieves text chunks according to the keywords, it retrieves the text chunks with the keywords.

[0039] However, through statistics, the inventors found that: the answers generated by the LLM based on the user's question and the text blocks retrieved through the above-mentioned method sometimes still have a low correlation with the user's question and still cannot answer the user's question well.

[0040] In view of this, the inventors analyzed the reasons specifically and found that: the text blocks retrieved through the above-mentioned method are text blocks with this keyword. However, text blocks without this keyword are not retrieved. However, sometimes there is an associative relationship in content between some text blocks without this keyword and the retrieved text blocks with this keyword. For example, they have a logical relationship and sometimes play a role of mutual confirmation and mutual supplementation in content. However, when the LLM generates answers, some text blocks without this keyword are not used by the LLM, resulting in incomplete text blocks used by the LLM, and further leading to a low correlation between the answers generated by the LLM and the user's question, and still unable to answer the user's question well.

[0041] Therefore, if some text blocks without this keyword can also be used by the LLM model to generate answers, then because there is an associative relationship in content between some text blocks without this keyword and the retrieved text blocks with this keyword, for example, they have a logical relationship, thus, they can play a role of mutual confirmation and mutual supplementation in content, thereby improving the integrity and relevance of the text blocks used by the LLM model when assisting the LLM model to generate answers, thereby improving the correlation between the answers generated by the LLM and the user's question, reducing the hallucination phenomenon of the LLM model, and further increasing the probability that the answers generated by the LLM can answer the user's question.

[0042] Especially when facing complex questions that require reasoning and summarization or answers that need to traverse multiple content-related text blocks to provide comprehensive insights, it can more significantly improve the correlation between the answers generated by the LLM and the user's question, more significantly reduce the hallucination phenomenon of the LLM model, and further more significantly increase the probability that the answers generated by the LLM can answer the user's question.

[0043] Specifically, for the convenience of understanding this application, first, a suitability description of the information retrieval architecture applicable to this application is given.

[0044] Refer to Figure 1 , which shows a schematic diagram of an information retrieval architecture of this application.

[0045] This architecture may include: a user device and an information retrieval system located on the server side.

[0046] A user can input a retrieval text (query) through a user device. Subsequently, the user device can send a retrieval request carrying the retrieval text to an information retrieval system, enabling the information retrieval system to perform a retrieval based on the retrieval text, obtain the retrieval results corresponding to the retrieval text, and return the retrieval results corresponding to the retrieval text to the user terminal for the user to view.

[0047] The user device can include: a smart mobile terminal, a smart home device, a wearable device, or a PC (Personal Computer), etc.

[0048] The smart mobile device can include: a mobile phone, a tablet computer, a laptop, a PDA (Personal Digital Assistant), or an Internet car, etc.

[0049] The smart home device can include: a smart TV or a smart refrigerator, etc.

[0050] The wearable device can include: a smart watch, smart glasses, a virtual reality device, an augmented reality device, or a mixed reality device (a device that can support virtual reality and augmented reality), etc.

[0051] The information retrieval system can adopt the method provided in the embodiments of this application to retrieve text blocks in the text block library based on the retrieval text and obtain the retrieval results.

[0052] The information retrieval system can be hosted on an independent physical machine on the server side, or it can also be in the cloud. For example, on a cloud server, which can also be referred to as a cloud computing server or a cloud host, and is a host product in the cloud computing service system to address the defects of high management difficulty and weak service scalability existing in traditional physical hosts and virtual private services (VPS, Virtual Private Server).

[0053] In addition to Figure 1 the architecture shown, the information retrieval system can also be set on a terminal device with strong computing power.

[0054] It can be understood that Figure 1 the numbers of the user device, the information retrieval system, the triple library, and the text block library in

[0055] are merely illustrative. According to the implementation requirements, there can be any number of user devices, information retrieval systems, triple libraries, text block libraries, etc.

[0056] The text block library is used to store text blocks, and the text blocks are used for retrieval.

[0057] Among them, during the retrieval process, the information retrieval system will involve the use of a triple library, a text block library, etc.

[0058] Refer to Figure 2 , which shows a step flowchart of an information retrieval method of the present application. This method can be applied to Figure 1 the information retrieval system shown. Among them, this method can specifically include the following steps:

[0059] In step S101, obtain the retrieval text.

[0060] In the present application, the user can input the retrieval text on the user device, control the user device to generate a retrieval request carrying the retrieval text, and control the user device to send the retrieval request to the information retrieval system.

[0061] The retrieval text can be a retrieval term (query), etc. Or, the retrieval text can also be a single character or a sentence, etc.

[0062] The information retrieval system can receive the retrieval request sent by the user device and extract the retrieval text in the retrieval request.

[0063] In one embodiment, the retrieval text can be the text directly input by the user in the text box on the interface displayed by the user device, etc., or it can be the text obtained after the user device performs speech recognition on the retrieval speech after inputting the retrieval speech to the user device by voice. Or, it can also be that after inputting the retrieval speech to the user device by voice, the user device sends a retrieval request carrying the retrieval speech to the information retrieval system. Then, the information retrieval system can extract the retrieval speech from the retrieval request and perform speech recognition on the retrieval speech to obtain the retrieval text.

[0064] In step S102, in the triples of each text block among multiple text blocks, find the first triple semantically related to the retrieval text. A text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words.

[0065] The multiple text blocks can be text blocks on the network, etc., or can be text blocks automatically collected from the network in advance and stored in the information retrieval system (such as the text block library in the information retrieval system).

[0066] In one example, multiple text blocks can be obtained from a data platform on the network. The data platform can include: encyclopedia data platforms such as Baidu Encyclopedia data platform, Sogou Encyclopedia data platform, Google Encyclopedia data platform, etc. Since the data in the encyclopedia data platform is highly accurate, rich in information, and the information layout is clear, it is convenient to extract entity words from the text blocks in the encyclopedia data platform later, which is beneficial to improving the accuracy of subsequent retrieval.

[0067] Among multiple text blocks, a part of the text blocks can belong to one document, another part of the text blocks can belong to another document... and yet another part of the text blocks can belong to yet another document, etc. That is to say, multiple text blocks can be obtained from multiple documents.

[0068] Documents can include online articles, papers, books, reviews, bullet screens or chat records, etc. A document can include more than two text blocks. The division of text blocks in a document can be done according to chapters, or according to headings, or according to the number of words, or according to paragraphs, etc., or it can also be manually divided according to actual needs or actual situations, etc.

[0069] There are many characters in a text block, which describe a piece of content. There are often multiple entity words in a text block. Any entity word in a text block can have an association relationship with at least one other entity word in the same text block to reflect the semantic relevance.

[0070] Association relationships include: ownership relationship (for example, Wang Wu owns a name), yes / no relationship (for example, Chen Liu is an athlete, or Zhang San is not the monitor), subordination relationship (for example, football belongs to sports, or basketball does not belong to music), and subject / object relationship (for example, the little dog has eaten the dog food, or the cat food has been eaten by the little cat), etc., which will not be elaborated here.

[0071] For example, in one example, assume a text block is: "Zhang San has a nickname 'Handsome Guy', Zhang San's date of birth is January 1, 1999, Zhang San's spouse is Li Si, Zhang San works in Company AA, and his position is a clerk."

[0072] Then the entity words in this text block can include: "Zhang San, nickname, Handsome Guy, date of birth, January 1, 1999, spouse, Li Si, Company AA, position, and clerk, etc."

[0073] The association relationships between these entity words include: Zhang San has a nickname, the nickname is Handsome Guy, Zhang San has a date of birth, the date of birth is January 1, 1999, Zhang San has a spouse, the spouse is Li Si, Zhang San has a company, the company is Company AA, Zhang San has a position in Company AA, and the position is a clerk, etc.

[0074] Thus, the triples of this text block can include:

[0075] Triple [Zhang San, has, alias]. In this triple, both "Zhang San" and "alias" are entity words, and "has" reflects the association relationship between the entity word "Zhang San" and "alias".

[0076] Triple [alias, is, handsome guy]. In this triple, both "alias" and "handsome guy" are entity words, and "is" reflects the association relationship between the entity word "alias" and "handsome guy".

[0077] Triple [Zhang San, has, date of birth]. In this triple, both "Zhang San" and "date of birth" are entity words, and "has" reflects the association relationship between the entity word "Zhang San" and "date of birth".

[0078] Triple [date of birth, is, January 1, 1999]. In this triple, both "date of birth" and "January 1, 1999" are entity words, and "is" reflects the association relationship between the entity word "date of birth" and "January 1, 1999".

[0079] Triple [Zhang San, has, spouse]. In this triple, both "Zhang San" and "spouse" are entity words, and "has" reflects the association relationship between the entity word "Zhang San" and "spouse".

[0080] Triple [spouse, is, Li Si]. In this triple, both "spouse" and "Li Si" are entity words, and "is" reflects the association relationship between the entity word "spouse" and "Li Si".

[0081] Triple [Zhang San, works at, company]. In this triple, both "Zhang San" and "company" are entity words, and "works at" reflects the association relationship between the entity word "Zhang San" and "company".

[0082] Triple [company, is, AA Enterprise]. In this triple, both "company" and "AA Enterprise" are entity words, and "is" reflects the association relationship between the entity word "position" and "AA Enterprise".

[0083] Triple [AA Enterprise, has, position]. In this triple, both "AA Enterprise" and "position" are entity words, and "has" reflects the association relationship between the entity word "(the AA Enterprise where Zhang San works)" and "position".

[0084] Triple [position, is, clerk]. In this triple, both "position" and "clerk" are entity words, and "is" reflects the association relationship between the entity word "(Zhang San's position in AA Enterprise)" and "clerk".

[0085] Among them, multiple triples of a text block can form the knowledge graph of the text block. A knowledge graph is a basic data structure in artificial intelligence and can be widely applied in fields such as search engines, social networks, and e-commerce. Generally, the knowledge graph of a text block can be formed according to multiple triples of the text block. A triple includes two different entity words and the association relationship between the two different entity words. Multiple triples of a text block can be connected through entity words, thereby obtaining the knowledge graph of the text block. For example, the knowledge graph of the above example can be seen as Figure 3 as shown.

[0086] For any one of multiple text blocks, the entity words in the text block and the association relationship between the entity words in the text block can be identified in advance, and then at least two triples of the text block can be established based on the entity words in the text block and the association relationship between the entity words in the text block.

[0087] A triple includes two entity words in the text block (the two entity words are different entity words) and the association relationship between the two entity words. Any two triples of the text block are different. For example, at least one entity word is different, or the two entity words are the same but the association relationship is different.

[0088] If the two entity words in two triples among multiple triples are the same but the association relationship is different, it means that there are multiple association relationships between the two entity words.

[0089] Among them, for identifying "the entity words in the text block and the association relationship between the entity words in the text block", the currently existing method can be used for identification. This application does not limit the specific identification method and will not elaborate here.

[0090] The same applies to each of the other text blocks among multiple text blocks.

[0091] Thus, at least two triples of each text block among multiple text blocks are obtained respectively. In this way, multiple triples are obtained, which are the triples of each text block among multiple text blocks.

[0092] The triples of each text block among multiple text blocks can be stored locally in the information retrieval system in advance (such as the triple library in the information retrieval system, etc.) to facilitate finding the first triple related to the text semantics for retrieval.

[0093] In an embodiment of this application, when finding the first triple related to the text semantics among the triples of each text block among multiple text blocks, it can be implemented through the following process, including:

[0094] 1021. Obtain the semantic similarity between the retrieval text and each triple in each of multiple text blocks.

[0095] In this application, the feature vector of the retrieval text can be obtained.

[0096] For example, the retrieval text can be encoded based on a pre-trained language model to obtain the feature vector of the retrieval text. For example, an embedding vector of the retrieval text can be obtained, etc.

[0097] The pre-trained language model can include: T5 (Transfer Text-to-Text Transformer) model, BERT (Bidirectional Encoder Representation from Transformers), XLNet (an autoregressive model that achieves bidirectional context information through a permutation language model), GPT (Generative Pre-Training) model, or bge-large-zh (a Chinese text representation model, part of a series of embedding large models open-sourced by BAAI, General Embedding-Large-zh, Beijing Academy of Artificial Intelligence), etc.

[0098] Secondly, the feature vector of each triple in each text block can be obtained separately.

[0099] Among them, for any text block and for any triple in that text block, the triple can be encoded based on the aforementioned pre-trained language model in advance to obtain the feature vector of the triple. For example, an embedding vector of the triple can be obtained, etc. Then, the triple and its feature vector can be combined into a corresponding table entry and stored in the correspondence between the triple and the feature vector of the triple (this correspondence can be located in the triple library). The same applies to each of the other triples in that text block. And the same applies to each triple in each of the other text blocks.

[0100] In this way, when it is necessary to obtain the feature vector of each triple in each text block separately, for any text block and for any triple in that text block, the feature vector corresponding to the triple can be found in the correspondence between the triple and the feature vector of the triple and used as the feature vector of the triple. The same applies to each of the other triples in that text block. And the same applies to each triple in each of the other text blocks.

[0101] By pre-setting the correspondence between triples and the feature vectors of the triples, the efficiency of obtaining the feature vectors of each triple in each text block separately can be improved.

[0102] Then, for any one of the triples in the triples of each text block among multiple text blocks, the vector similarity between the feature vector of the retrieval text and the feature vector of this triple can be calculated.

[0103] For example, calculate the cosine similarity between the feature vector of the retrieval text and the feature vector of this triple, and use it as the vector similarity between the feature vector of the retrieval text and the feature vector of this triple.

[0104] Or, for another example, calculate the reciprocal of the Euclidean distance between the feature vector of the retrieval text and the feature vector of this triple, and use it as the vector similarity between the feature vector of the retrieval text and the feature vector of this triple.

[0105] Among them, the smaller the Euclidean distance, the greater the similarity; the greater the Euclidean distance, the smaller the similarity. Therefore, the smaller the reciprocal of the Euclidean distance, the smaller the similarity; the greater the reciprocal of the Euclidean distance, the greater the similarity.

[0106] After that, the semantic similarity between the retrieval text and this triple can be obtained according to this vector similarity. For example, determine this vector similarity as the semantic similarity between the retrieval text and this triple.

[0107] The same applies to every other triple in the triples of each text block among multiple text blocks.

[0108] 1022. Among the triples of each text block among multiple text blocks, select K triples with the largest semantic similarity to the retrieval text, where K is a positive integer.

[0109] In one example, K can be 3, 4, 5, 6, 7, 8, etc., and can be determined according to the actual situation specifically, and this application does not limit this.

[0110] 1023. Obtain the first triple according to the K triples.

[0111] For example, the K triples can be determined as the first triple.

[0112] In step S103, according to the first triple, search for the second triple in the triples of each text block among multiple text blocks. The second triple includes at least one of the following: the entity word in the first triple, and the entity word having an association relationship with the entity word in the first triple.

[0113] Entity words that have an associated relationship with the entity words in the first triple will be located not only in the first triple but also in some other triples.

[0114] The second triple includes the entity words in the first triple, or entity words that have an associated relationship with the entity words in the first triple, or both the entity words in the first triple and entity words that have an associated relationship with the entity words in the first triple.

[0115] In an embodiment of the present application, this step can be implemented through the following process, including:

[0116] 1031. For any one of the entity words in the first triple, in the triples of each text block among multiple text blocks, search for the intermediate triples that have this entity word, and, in the triples of each text block among multiple text blocks, search for the associated triples that have another entity word, where the other entity word is the entity word other than this one in the intermediate triple.

[0117] In the present application, for any one of the entity words in each determined first triple, and for any one of the triples in the triples of each text block among multiple text blocks, if this entity word is located in this triple, then this triple can be determined as an intermediate triple, or, if this entity word is not located in this triple, then this triple may not be determined as an intermediate triple.

[0118] For each of the other entity words in each determined first triple, and for each of the other triples in the triples of each text block among multiple text blocks, the same applies. Thus, all intermediate triples that have this entity word are searched for in the triples of each text block among multiple text blocks.

[0119] Among them, each intermediate triple has another entity word in addition to having this entity word.

[0120] For the other entity word other than this entity word in any obtained intermediate triple, and for any one of the triples in the triples of each text block among multiple text blocks, if this other entity word is located in this triple, then this triple can be determined as an associated triple, or, if this other entity word is not located in this triple, then this triple may not be determined as an associated triple.

[0121] For each of the other obtained intermediate triples, for the other entity word in the intermediate triple other than the entity word, and for each of the other triples in the triples of each text block among multiple text blocks, the same applies, so as to find, in the triples of each text block among multiple text blocks, the associated triples having the other entity word, where the other entity word is the non-entity word in the intermediate triple.

[0122] 1032. Obtain a second triple according to the intermediate triple and the associated triple.

[0123] In an embodiment of the present application, each obtained intermediate triple and each obtained associated triple may be determined as the second triple.

[0124] Alternatively, in another embodiment of the present application, the semantic similarity between the retrieval text and each of the triples in the intermediate triple and the associated triple may be obtained.

[0125] For example, the feature vector of the retrieval text may be obtained.

[0126] In one example, the retrieval text may be encoded based on a pre-trained language model to obtain the feature vector of the retrieval text. For example, the embedding vector of the retrieval text, etc., may be obtained.

[0127] Secondly, the feature vectors of each of the triples in the intermediate triple and the associated triple may be obtained respectively.

[0128] For example, for any one of the triples in the intermediate triple and the associated triple, in the correspondence relationship between the triple and the feature vector of the triple, the feature vector corresponding to the triple may be found and used as the feature vector of the triple. The same applies to each of the other triples in the intermediate triple and the associated triple.

[0129] By pre-setting the correspondence relationship between the triple and the feature vector of the triple, the efficiency of respectively obtaining the feature vectors of each of the triples in the intermediate triple and the associated triple can be improved.

[0130] In addition, for any one of the triples in the intermediate triple and the associated triple, the vector similarity between the feature vector of the retrieval text and the feature vector of the triple may be calculated.

[0131] For example, calculate the cosine similarity between the feature vector of the retrieval text and the feature vector of the triple, and use it as the vector similarity between the feature vector of the retrieval text and the feature vector of the triple.

[0132] Alternatively, for another example, calculate the reciprocal of the Euclidean distance between the feature vector of the retrieval text and the feature vector of the triple, and use it as the vector similarity between the feature vector of the retrieval text and the feature vector of the triple.

[0133] After that, the semantic similarity between the retrieval text and the triple can be obtained according to the vector similarity. For example, determine the vector similarity as the semantic similarity between the retrieval text and the triple.

[0134] The same applies to each of the other triples in the intermediate triple and the associated triples.

[0135] After that, among the intermediate triple and the associated triples, select Q triples with the largest semantic similarity to the retrieval text, where Q is a positive integer. In one example, Q can be 10, 11, 12, 13, 14, or 15, etc., which can be determined according to the actual situation, and the present application does not limit this.

[0136] Then, the second triple can be obtained according to the Q triples. For example, the Q triples can be determined as the second triple.

[0137] In step S104, among the multiple text blocks, find the text block having the second triple.

[0138] The multiple text blocks can be located in the text block library.

[0139] In the present application, for any one of the multiple text blocks, after obtaining multiple triples of the text block in advance, for any one triple of the text block, the identification information of the text block and the triple of the text block can be combined into a corresponding table entry and stored in the corresponding relationship between the identification information of the text block and the triples of the text block (this corresponding relationship can be located in the triple library). The same applies to each of the other triples of the text block. Additionally, the same applies to each of the other text blocks among the multiple text blocks.

[0140] Therefore, in this step, in the corresponding relationship between the identification information of the text block and the triples of the text block, find the identification information corresponding to the second triple (this identification information is the identification information of a text block), and then obtain the text block corresponding to this identification information among the multiple text blocks, which is the text block having the second triple.

[0141] In step S105, obtain the retrieval result corresponding to the retrieval text according to the text block having the second triple.

[0142] In an embodiment of the present application, there is one text block with a second triple. In this case, one text block with a second triple can be determined as the retrieval result corresponding to the retrieval text.

[0143] Alternatively, in an embodiment of the present application, multiple text blocks are obtained. In this case, this step can be implemented through the following process, including:

[0144] 1051. Obtain the semantic similarity between the retrieval text and each of the obtained text blocks respectively.

[0145] In one embodiment, the feature vector of the retrieval text can be obtained.

[0146] For example, the retrieval text can be encoded based on a pre-trained language model to obtain the feature vector of the retrieval text. For example, an embedding vector of the retrieval text can be obtained, etc.

[0147] Secondly, the feature vectors of each of the obtained text blocks can be obtained respectively.

[0148] Among them, for any one text block, the text block can be encoded based on the aforementioned pre-trained language model in advance to obtain the feature vector of the text block. For example, an embedding vector of the text block can be obtained, etc. Then, the identification information of the text block and the feature vector of the text block can be combined into a corresponding table entry and stored in the corresponding relationship between the identification information of the triple and the feature vector of the text block. The same applies to each of the other text blocks.

[0149] In this way, when the feature vectors of each of the obtained text blocks need to be obtained respectively, for any one of the obtained text blocks, the feature vector corresponding to the identification information of the text block can be found in the corresponding relationship between the identification information of the text block and the feature vector of the text block and used as the feature vector of the text block. The same applies to each of the other obtained text blocks.

[0150] By setting in advance the corresponding relationship between the identification information of the text block and the feature vector of the text block, the efficiency of obtaining the feature vectors of each of the obtained text blocks respectively can be improved.

[0151] Then, for the feature vector of any one of the obtained text blocks, the vector similarity between the feature vector of the retrieval text and the feature vector of the text block can be calculated.

[0152] For example, calculate the cosine similarity between the feature vector of the retrieval text and the feature vector of the text block and use it as the vector similarity between the feature vector of the retrieval text and the feature vector of the text block.

[0153] Alternatively, for another example, calculate the reciprocal of the Euclidean distance between the feature vector of the retrieval text and the feature vector of the text block, and use it as the vector similarity between the feature vector of the retrieval text and the feature vector of the text block.

[0154] After that, the vector similarity between the retrieval text and the text block can be obtained according to the vector similarity. For example, determine the vector similarity as the vector similarity between the retrieval text and the text block.

[0155] The same applies to the feature vectors of each of the other obtained text blocks.

[0156] In another embodiment, for any obtained text block, determine the number of triples belonging to the second triples among the multiple triples of the text block. Obtain the semantic similarity between the retrieval text and the text block according to the number. The larger the number, the greater the semantic similarity between the retrieval text and the text block; or the smaller the number, the smaller the semantic similarity between the retrieval text and the text block.

[0157] 1052. Among multiple text blocks, select S text blocks with the largest semantic similarity to the retrieval text, where S is a positive integer.

[0158] In one example, S can be 1, 2, 3, 4, 5, or 6, etc., which can be determined according to the actual situation, and this application does not limit it.

[0159] 1053. Obtain the retrieval result corresponding to the retrieval text according to the S text blocks.

[0160] For example, the S text blocks can be determined as the retrieval result corresponding to the retrieval text.

[0161] Furthermore, the retrieval text and the retrieval result can be input into a model (such as an LLM model, etc.) so that the model uses the retrieval text and the retrieval result to generate an answer and output the answer, and the information retrieval system can obtain the answer output by the model.

[0162] In this application, obtain the retrieval text. Among the triples of each text block in multiple text blocks, search for the first triples semantically related to the retrieval text. A text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words. According to the first triples, search for the second triples in the triples of each text block in multiple text blocks. The second triples include at least one of the following: the entity words in the first triples, and the entity words having an association relationship with the entity words in the first triples. Among multiple text blocks, search for the text blocks having the second triples. Obtain the retrieval result corresponding to the retrieval text according to the text blocks having the second triples.

[0163] In this application, the first triple is semantically related to the retrieval text. Since the second triple includes at least one of the following: the entity word in the first triple, and the entity word having an association relationship with the entity word in the first triple, therefore, the second triple is also semantically related to the retrieval text. There can be multiple text blocks having the second triple, and there is an association relationship in content among the multiple text blocks having the second triple. For example, they have a logical relationship. Thus, the multiple text blocks having the second triple can serve to corroborate and complement each other in content. Thus, the content in the retrieval result corresponding to the retrieval text obtained from the multiple text blocks having the second triple (the retrieval result includes at least two text blocks having the second triple) is also related. For example, they have a logical relationship. Thus, it can serve to corroborate and complement each other. In this way, inputting the retrieval text and the retrieval result corresponding to the retrieval text into the LLM model can improve the integrity and relevance of the text blocks used by the LLM model when assisting the LLM model to generate answers, thereby improving the relevance between the answer generated by the LLM and the user's question, reducing the hallucination phenomenon of the LLM model, and further increasing the probability that the answer generated by the LLM can answer the user's question.

[0164] Especially when facing complex questions that require reasoning and summarization, or answers that require traversing multiple content-related text blocks to provide comprehensive insights, it can more significantly improve the relevance between the answer generated by the LLM and the user's question, can more significantly reduce the hallucination phenomenon of the LLM model, and further more significantly increase the probability that the answer generated by the LLM can answer the user's question.

[0165] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be in other sequences or carried out simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily essential to this application.

[0166] Refer to Figure 4 , which shows a structural block diagram of an information retrieval device of this application. The device includes:

[0167] A first acquisition module 11, configured to acquire a retrieval text;

[0168] A first search module 12, configured to search for a first triple that is semantically related to the retrieval text in the triples of each text block among multiple text blocks; a text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words;

[0169] A second search module 13, configured to search for a second triple in the triples of each text block among multiple text blocks according to the first triple, where the second triple includes at least one of the following: an entity word in the first triple, and an entity word having an association relationship with the entity word in the first triple;

[0170] A third search module 14, configured to search for a text block having the second triple among the multiple text blocks;

[0171] A second acquisition module 15, configured to acquire a retrieval result corresponding to the retrieval text according to the text block having the second triple.

[0172] In an optional implementation manner, the first search module includes:

[0173] A first acquisition unit, configured to acquire the semantic similarity between the retrieval text and each triple of each text block among multiple text blocks;

[0174] A first selection unit, configured to select the top K triples with the greatest semantic similarity between the retrieval text and the triples of each text block among multiple text blocks, where K is a positive integer;

[0175] A second acquisition unit, configured to acquire a first triple according to the K triples.

[0176] In an optional implementation manner, the first acquisition unit includes:

[0177] A first acquisition subunit, configured to acquire the feature vector of the retrieval text;

[0178] A second acquisition subunit, configured to acquire the feature vectors of each triple of each text block respectively;

[0179] A first calculation subunit, configured to calculate the vector similarity between the feature vector of the retrieval text and the feature vector of any triple among the triples of each text block among multiple text blocks, and a third acquisition subunit, configured to acquire the semantic similarity between the retrieval text and the triple according to the vector similarity.

[0180] In an optional implementation manner, the second search module includes:

[0181] A search unit, configured to, for any entity word in the first triple, search for an intermediate triple having the entity word in the triples of each text block among multiple text blocks, and search for an associated triple having another entity word in the triples of each text block among multiple text blocks, where the another entity word is the non-entity word in the intermediate triple;

[0182] A third obtaining unit, configured to obtain the second triple according to the intermediate triple and the associated triple.

[0183] In an optional implementation manner, the third obtaining unit includes:

[0184] A fourth obtaining subunit, configured to obtain the semantic similarity between the retrieved text and each triple in the intermediate triple and the associated triple;

[0185] A selecting subunit, configured to select Q triples with the largest semantic similarity with the retrieved text from the intermediate triple and the associated triple, where Q is a positive integer;

[0186] A fifth obtaining subunit, configured to obtain the second triple according to the Q triples.

[0187] In an optional implementation manner, the fourth obtaining subunit is specifically configured to:

[0188] Obtain the feature vector of the retrieved text;

[0189] Respectively obtain the feature vectors of each triple in the intermediate triple and the associated triple;

[0190] For any triple in the intermediate triple and the associated triple, calculate the vector similarity between the feature vector of the retrieved text and the feature vector of the triple, and obtain the semantic similarity between the retrieved text and the triple according to the vector similarity.

[0191] In an optional implementation manner, multiple text blocks are obtained;

[0192] The second obtaining module includes:

[0193] A fourth obtaining unit, configured to obtain the semantic similarity between the retrieved text and each obtained text block;

[0194] A second selecting unit, configured to select S text blocks with the largest semantic similarity with the retrieved text from multiple text blocks, where S is a positive integer;

[0195] A fifth acquisition unit, configured to obtain a retrieval result corresponding to the retrieval text according to the S text blocks.

[0196] In an optional implementation manner, the fourth acquisition unit includes:

[0197] A sixth acquisition subunit, configured to obtain a feature vector of the retrieval text;

[0198] A seventh acquisition subunit, configured to respectively obtain feature vectors of the acquired text blocks;

[0199] A second calculation subunit, configured to calculate a vector similarity between the feature vector of the retrieval text and the feature vector of any one of the acquired text blocks, and an eighth acquisition subunit, configured to obtain a semantic similarity between the retrieval text and the text block according to the vector similarity.

[0200] In an optional implementation manner, the apparatus further includes:

[0201] An input module, configured to input the retrieval text and the retrieval result into a model, so that the model uses the retrieval text and the retrieval result to generate an answer;

[0202] A third acquisition module, configured to obtain the answer output by the model.

[0203] In this application, a retrieval text is obtained. Among the triples of each text block in multiple text blocks, a first triple semantically related to the retrieval text is found. A text block has multiple triples, and a triple includes two entity words in the text block to which it belongs and the association relationship between the two entity words. According to the first triple, a second triple is found among the triples of each text block in the multiple text blocks, and the second triple includes at least one of the following: the entity words in the first triple, and the entity words having an association relationship with the entity words in the first triple. Among the multiple text blocks, the text blocks having the second triple are found. According to the text blocks having the second triple, a retrieval result corresponding to the retrieval text is obtained.

[0204] In this application, the first triple is semantically related to the retrieval text. Since the second triple includes at least one of the following: the entity word in the first triple, and the entity word having an association relationship with the entity word in the first triple, therefore, the second triple is also semantically related to the retrieval text. There can be multiple text blocks having the second triple, and there is an association relationship in content among the multiple text blocks having the second triple. For example, having a logical relationship. Thus, the multiple text blocks having the second triple can corroborate and complement each other in content. Thus, the content among the retrieval results corresponding to the retrieval text obtained from the multiple text blocks having the second triple (the retrieval results include at least two text blocks having the second triple) is also related. For example, having a logical relationship. Thus, it can play a role in corroborating and complementing each other. In this way, inputting the retrieval text and the retrieval results corresponding to the retrieval text into the LLM model can improve the integrity and relevance of the text blocks used by the LLM model when assisting the LLM model to generate answers, thereby improving the relevance between the answers generated by the LLM and the user's questions, reducing the hallucination phenomenon of the LLM model, and further increasing the probability that the answers generated by the LLM can answer the user's questions.

[0205] Especially when facing complex questions that require reasoning and summarization, or answers that need to traverse multiple content-related text blocks to provide comprehensive insights, it can more significantly improve the relevance between the answers generated by the LLM and the user's questions, can more significantly reduce the hallucination phenomenon of the LLM model, and further more significantly increase the probability that the answers generated by the LLM can answer the user's questions.

[0206] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment.

[0207] Optionally, the embodiment of the present application further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0208] The embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium, such as a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, or an optical disc, etc.

[0209] Figure 5 It is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0210] Referring to Figure 5 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0211] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0212] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0213] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0214] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0215] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0216] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0217] The sensor component 814 includes one or more sensors for providing an assessment of the various aspects of the status of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0218] The communication component 816 is configured to facilitate communication, either wired or wirelessly, between the electronic device 800 and other devices. The electronic device 800 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0219] In an exemplary embodiment, the electronic device 800 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0220] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided. The above instructions can be executed by the processor 820 of the electronic device 800 to complete the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, among others.

[0221] Figure 6 FIG. 10 is a block diagram of an electronic device 1900 shown in the present application. For example, the electronic device 1900 can be provided as a server.

[0222] Referring to Figure 6 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-described method.

[0223] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0224] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0225] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0226] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application, without departing from the purpose of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.

[0227] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0228] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0229] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0230] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0231] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0232] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0233] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An information retrieval method, characterized in that: The method comprises: Get the search text; Searching for a first triplet semantically related to the search text in the triplet of each text block in the plurality of text blocks; a text block has a plurality of triples, and a triplet includes two entity words in the text block to which it belongs and an association relationship between the two entity words; According to the first triple, searching for a second triple in the triples of each text block in the plurality of text blocks, wherein the second triple includes at least one of the following: an entity word in the first triple, and an entity word having an association relationship with the entity word in the first triple; Searching for a text block having the second triplet among the multiple text blocks; According to the text block having the second triplet, a search result corresponding to the search text is obtained.

2. The method according to claim 1, characterized in that: The step of searching, among the triples of each text block in the plurality of text blocks, for a first triple semantically related to the search text comprises: Obtaining semantic similarity between the search text and each triple of each text block in the plurality of text blocks; Selecting K triples with the greatest semantic similarity to the search text from the triples of each text block in the plurality of text blocks, where K is a positive integer; A first triplet is obtained according to the K triples.

3. The method according to claim 2, characterized in that The obtaining of the semantic similarity between the search text and each triple of each text block in the plurality of text blocks includes: Obtaining a feature vector of the search text; Obtain the feature vector of each triplet of each text block respectively; For any triple of the triplets of each text block in the plurality of text blocks, the vector similarity between the feature vector of the search text and the feature vector of the triple is calculated, and the semantic similarity between the search text and the triple is obtained according to the vector similarity.

4. The method according to claim 1, characterized in that: The step of searching, according to the first triplet, for a second triplet in the triplet of each text block in the plurality of text blocks comprises: For any entity word in the first triple, in the triples of each text block in a plurality of text blocks, an intermediate triple with the entity word is searched, and in the triples of each text block in a plurality of text blocks, an associated triple with another entity word is searched, where the other entity word is not the entity word in the intermediate triple; The second triplet is acquired according to the intermediate triplet and the associated triplet.

5. The method according to claim 4, characterized in that The acquiring the second triplet according to the intermediate triplet and the associated triplet includes: Obtaining semantic similarity between the search text and each of the intermediate triples and the associated triples; Select Q triples with the greatest semantic similarity to the search text from among the intermediate triples and the associated triples, where Q is a positive integer; A second triplet is obtained according to the Q triples.

6. The method according to claim 5, characterized in that The obtaining of the semantic similarity between the search text and each of the triples in the intermediate triples and the associated triples includes: Obtaining a feature vector of the search text; Obtaining feature vectors of each triple in the intermediate triple and the associated triple respectively; For any triplet in the intermediate triplet and the associated triplet, the vector similarity between the feature vector of the search text and the feature vector of the triplet is calculated, and the semantic similarity between the search text and the triplet is obtained according to the vector similarity.

7. The method according to claim 1, characterized in that There are multiple text blocks obtained; The step of obtaining the search result corresponding to the search text according to the obtained text block includes: Obtaining semantic similarities between the search text and each of the obtained text blocks; Selecting S text blocks with the greatest semantic similarity to the search text from among multiple text blocks, where S is a positive integer; The search results corresponding to the search text are obtained according to the S text blocks.

8. The method according to claim 7, characterized in that The obtaining of the semantic similarity between the search text and each of the obtained text blocks includes: Obtaining a feature vector of the search text; respectively obtaining feature vectors of each of the obtained text blocks; For the acquired feature vector of any text block, the vector similarity between the feature vector of the search text and the feature vector of the text block is calculated, and the semantic similarity between the search text and the text block is acquired according to the vector similarity.

9. The method according to claim 1, characterized in that: The method further comprises: Inputting the search text and the search results into a model so that the model generates an answer using the search text and the search results; Obtain the answer output by the model.

10. An information retrieval device, characterized in that: The device comprises: A first acquisition module is used to acquire a search text; A first search module is used to search for a first triplet semantically related to the search text in the triplet of each text block in the plurality of text blocks; a text block has a plurality of triples, and a triplet includes two entity words in the text block to which it belongs and an association relationship between the two entity words; A second search module is configured to search, according to the first triple, for a second triple in the triples of each text block in the plurality of text blocks, wherein the second triple includes at least one of the following: an entity word in the first triple and an entity word having an association relationship with the entity word in the first triple; A third search module, configured to search for a text block having the second triplet among the plurality of text blocks; The second acquisition module is used to acquire the search result corresponding to the search text according to the text block having the second triple.

11. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 9 when executed by the processor.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 9.