Information retrieval method and device based on large language model, medium and electronic equipment

By building a graph structure in a large language model to retrieve related document elements, the problem of contextual information fragmentation in existing technologies is solved, and more accurate response generation is achieved.

CN120763306APending Publication Date: 2025-10-10BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511287409.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the existing technology, the similarity recall-based method cannot effectively capture the semantic or logical association relationships between document elements, resulting in the fragmentation and loss of contextual information, affecting the response accuracy of the large language model.

Method used

By constructing a pre-built graph structure, we retrieve the first document element and its associated second document element that are highly similar to the target question, and use the hierarchical and associative relationships in the graph structure to determine the contextual information, supporting the large language model to generate more accurate responses.

Benefits of technology

It improves the completeness of contextual information, enhances the accuracy and relevance of responses generated by large language models, and solves the problem of fragmentation and missing contextual information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763306A_ABST
    Figure CN120763306A_ABST
Patent Text Reader

Abstract

The invention discloses an information retrieval method and device based on a large language model, a medium and electronic equipment, and relates to the technical field of large language models.The information retrieval method comprises the steps that a first document element with the similarity with a target problem larger than a preset similarity threshold value is obtained; second document elements associated with the first document elements are retrieved in a pre-constructed graph structure, the graph structure is obtained by constructing a connecting edge in a structure tree corresponding to the document, the structure tree is used for describing the hierarchical relation between nodes, each node uniquely corresponds to one document element in the document, and each node corresponds to one document element in the document; the connecting edge is used for describing a logic association relationship or a semantic association relationship between the nodes; and determining the first document element and the second document element as the context information of the target problem, thereby solving the problems of context information splitting, context information missing and the like caused by incapability of effectively capturing a semantic association relationship or a logic association relationship among the document elements. And the completeness of the context information on which the large language model is used for replying is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer technology, the technical field of large language model, in particular, relates to a large language model-based information retrieval method and device, medium, electronic equipment and program product. BACKGROUND

[0002] RAG (Retrieval Augmented Generation) refers to a technology that combines traditional information retrieval and generation model (e.g., LLM (Large Language Model)). RAG "injects" the retrieved context information into the input of the large language model, thereby enhancing the generation capability of the large language model and reducing the hallucination of the large language model, so that the large language model can generate more accurate, more relevant and more specific replies. Therefore, the context information affects the understanding of the large language model to the problem and the output of the large language model. SUMMARY

[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004] In a first aspect, the present disclosure provides a large language model-based information retrieval method, comprising: obtaining a first document element corresponding to a target question, wherein the similarity between the first document element and the target question is greater than a preset similarity threshold; retrieving a second document element having an association relationship with the first document element in a pre-constructed graph structure, wherein the graph structure is obtained by constructing a connection edge in a structure tree corresponding to a document, the structure tree is used to describe the hierarchical relationship between nodes, each node uniquely corresponds to a document element in the document, and the connection edge is used to describe the logical association relationship or semantic association relationship between nodes; determining the first document element and the second document element as context information of the target question, wherein the context information is used to support a large language model to generate a target reply to the target question.

[0005] In a second aspect, the present disclosure provides a large language model-based information retrieval device, comprising: a first obtaining module, configured to obtain a first document element corresponding to a target question, wherein the similarity between the first document element and the target question is greater than a preset similarity threshold; retrieving a second document element having an association relationship with the first document element in a pre-constructed graph structure, wherein the graph structure is obtained by constructing a connection edge in a structure tree corresponding to a document, the structure tree is used to describe a hierarchical relationship between nodes, each node uniquely corresponds to a document element in the document, and the connection edge is used to describe a logical association relationship or a semantic association relationship between nodes; determining the first document element and the second document element as context information of the target question, wherein the context information is used to support a large language model to generate a target reply of the target question.

[0006] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the information retrieval method of the first aspect.

[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the information retrieval method of the first aspect.

[0008] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the information retrieval method of the first aspect.

[0009] Through the above technical solution, the first document element having a similarity greater than a preset similarity threshold to the target question is obtained first, and then the second document element having an association relationship with the first document element is retrieved in the pre-constructed graph structure. This solves the problem of context information fragmentation and context information missing caused by the inability to effectively capture the semantic association relationship or logical association relationship between document elements in the mode of recall based only on similarity, improves the integrity of the context information, and further improves the accuracy of the target reply of the target question generated by the large language model relying on the context information.

[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. The same or similar components are denoted by the same or similar reference numerals throughout the drawings. It is to be understood that the drawings are schematic, and the original and elements are not necessarily drawn to scale. In the drawings: Figure 1is a flowchart of a large language model-based information retrieval method according to an embodiment of the present disclosure; Figure 2 is a schematic diagram of a structure tree according to an embodiment of the present disclosure; Figure 3 is a large language model-based information retrieval method according to an embodiment of the present disclosure; Figure 2 is a schematic diagram of a graph structure constructed according to the structure tree shown in FIG. 8; Figure 4 is a process schematic diagram of a large language model-based information retrieval method according to an embodiment of the present disclosure; Figure 5 is a block diagram of a large language model-based information retrieval device according to an embodiment of the present disclosure; Figure 6 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0012] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.

[0013] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.

[0014] The term "comprising" and variations thereof as used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0015] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0016] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0018] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0019] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.

[0020] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0021] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manners of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0022] At the same time, it can be understood that the data involved in the present technical solutions (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.

[0023] RAG (Retrieval Augmented Generation) is a technology that combines traditional information retrieval and generation models (such as LLM), which "injects" the retrieved contextual information into the input of the large language model, thereby enhancing the generation ability of the large language model and reducing the hallucination of the large language model, so that the large language model can generate more accurate, more relevant and more specific replies.

[0024] Since the retrieved information is obtained through sharding, due to the size limitation of the shards, the information is fragmented. For example, for a document containing 10 consecutive operation steps, it can be divided into 10 shards in the sharding process. This fragmentation of information makes each shard contain only part of the context, losing the coherence and integrity between the steps.

[0025] However, the related technology based only on similarity recall cannot effectively capture the semantic or logical association relationship between document fragments. It only recalls some similar fragments and omits other steps closely related to them. In other words, it is difficult to recall complete contextual information, which in turn affects the response of the large language model.

[0026] In view of this, the embodiments of the present disclosure provide an information retrieval method, apparatus, medium, electronic device, and program product based on a large language model, which solves problems such as the inability to effectively capture semantic or logical associations between document elements, resulting in fragmentation and loss of contextual information, and improves the integrity of the contextual information that the large language model relies on for responses.

[0027] The embodiments of the present disclosure are further explained and illustrated below with reference to the accompanying drawings.

[0028] Figure 1 This is a flowchart of an information retrieval method based on a large language model according to an embodiment of the present disclosure. The information retrieval method based on a large language model can be applied to electronic devices. The information retrieval method based on a large language model can be executed by an information retrieval device based on a large language model, wherein the information retrieval device based on a large language model can be implemented by software and / or hardware, and the software and / or hardware can be configured in an electronic device. Figure 1 The information retrieval method based on a large language model may include step 110, step 120 and step 130.

[0029] In step 110, a first document element corresponding to the target question is obtained, wherein the similarity between the first document element and the target question is greater than a preset similarity threshold.

[0030] The target question can be a question that the user inputs into the large language model. The large language model is used to understand the semantics of the target question and generate a target response to the target question.

[0031] The preset similarity threshold can be set according to actual conditions and is not limited in this embodiment.

[0032] Among them, the explanation and description of the first document element can refer to the following related embodiments, and this embodiment will not be described in detail here.

[0033] In step 120, a second document element having an association relationship with the first document element is retrieved in a pre-constructed graph structure, wherein the graph structure is obtained by constructing connecting edges in a structure tree corresponding to the document, the structure tree is used for the hierarchical relationship between nodes, each node uniquely corresponds to a document element in the document, and the connecting edges are used to describe the logical association relationship or semantic association relationship between the nodes.

[0034] The document can be subjected to a structured splitting process, and a structure tree corresponding to the document is obtained. A node in the structure tree uniquely corresponds to a document element in the document, which can be, for example, a title in the document, a sentence in a paragraph, a chart, and the like, and a root node in the structure tree is the document. Figure 2 FIG. 1 is a schematic diagram of a structure tree according to an embodiment of the present disclosure. Referring to FIG. 1, Figure 2 Figure 2 Each title and each section in FIG. 1 represents a different document element in the document. According to the structural relationship of the structure tree, the structure tree can describe the hierarchical relationship between nodes, for example, the first level, the second level, and the third level shown in FIG. 1, to provide support for subsequent hierarchical-based retrieval. Figure 2

[0035] The leaf node in the structure tree is a first type of document element, and the second document element and the first document element both belong to the first type of document element. The leaf node in the structure tree is a node without child nodes. The semantic information of the first type of document element is greater than that of the second type of document element. The second type of document element can be understood as an abstract of the first type of document element. For example, the second type of document element can be a title, and the first type of document element can be a sentence in a paragraph and a chart under the title, and the like.

[0036] Figure 3 FIG. 2 is a schematic diagram of a graph structure constructed based on the structure tree shown in FIG. 1 according to an embodiment of the present disclosure. Referring to FIG. 2, Figure 2 In the graph structure, the connection edges include directed edges 301 and bidirectional edges 302. The directed edges 301 are connection edges between document elements of the same level (for example, section 1 and title 1 shown in FIG. 2, and section 1 and title 2 shown in FIG. 2), and are directed from a first type of document element (for example, section 1 shown in FIG. 2) to a second type of document element (for example, title 2 shown in FIG. 2). The bidirectional edges 302 are connection edges between first type of document elements (for example, section 4 and section 5 shown in FIG. 2). Figure 3 Figure 3 Figure 3 Figure 3 Figure 3 For example, in a document covering operation steps, the logical association relationship can be, for example, the order relationship between operation steps; for another example, the semantic association relationship can be, for example, the relationship of referring to the same entity.

[0037] In step 130, the first document element and the second document element are determined as context information of the target question, wherein the context information is used to support the large language model to generate a target reply to the target question.

[0038]

[0039] ​​​​​​​By the technical solution, the first document element with a similarity greater than a preset similarity threshold to the target question is acquired first, and then the second document element having an association relationship with the first document element is retrieved in the pre-constructed graph structure, so as to solve the problem of context information fragmentation and context information loss caused by the inability to effectively capture the semantic association relationship or logical association relationship between document elements in the mode of similarity-based recall, improve the integrity of the context information, and further improve the accuracy of the target reply of the target question generated by the large language model depending on the context information.

[0040] In some embodiments, the step of retrieving the second document element having an association relationship with the first document element in the pre-constructed graph structure can include: taking a first target constraint as a target, retrieving the second document element having an association relationship with the first document element in the pre-constructed graph structure, wherein the first target constraint includes a specified level, and the level of the retrieved second document element is the specified level.

[0041] In the structure tree, the closer the level of the document element to the level of the first document element, the stronger the relevance of the document element to the first document element, and therefore, the level of the retrieved second document element can be set to the specified level, thereby ensuring that the retrieved second document element is strongly related to the first document element.

[0042] In some embodiments, the specified level is set by a user, and the specified level can be the level of the first document element or other levels. As an example, the user can set the specified level according to the scenario in which the large language model is used. For example, in a task scenario with high real-time requirements, the large language model needs to make a reply quickly, and therefore, fewer specified levels can be configured to reduce the amount of data that the large language model needs to process, and in a task scenario with low real-time requirements, the large language model does not need to make a reply quickly, and therefore, more specified levels can be configured to increase the amount of context information on which the large language model depends, to give a more accurate reply.

[0043] In some embodiments, the first target constraint further comprises a specified number, the specified number being used to indicate a number of the second document elements to be retrieved, the specified level comprises a level where the first document element is located, and the step of retrieving the second document elements associated with the first document element in the pre-constructed graph structure based on the first target constraint can be further implemented by: retrieving the second document elements associated with the first document element in the level where the first document element is located in the pre-constructed graph structure; in a case where the number of the second document elements retrieved in the level where the first document element is located does not reach the specified number, retrieving the second document elements associated with the first document element step by step upwards and downwards from the level where the first document element is located based on the second preset constraint until the number of the second document elements reaches the specified number, the second preset constraint being that the difference between the number of the second document elements above the level where the first document element is located and the number of the second document elements below the level where the first document element is located is minimized. The specified number is used to limit the number of the second document elements to be retrieved, so as to avoid the prompt words (including context) given to the large language model exceeding the limitation of the model.

[0044] It should be noted that the hierarchical relationship reflects the proximity of the document elements in the document, and in general, the closer the document elements, the higher the correlation. The upper and lower levels adjacent to the first document element are strongly related to the first document element, and therefore, in a case where the number of the second document elements retrieved in the level where the first document element is located does not reach the specified number, the retrieval of the second document elements is further performed step by step downwards and upwards from the level where the first document element is located.

[0045] In addition, for the upper level and the lower level of the level where the first document element is located, the degree of correlation can be considered consistent, and therefore, the second preset constraint is set to ensure that the difference between the number of the second document elements above the level where the first document element is located and the number of the second document elements below the level where the first document element is located is minimized.

[0046] In some embodiments, the step of obtaining the first document element corresponding to the target question can be implemented by: recalling third document elements based on the target question, wherein the similarity between the third document elements and the target question is greater than the preset similarity threshold; sorting the third document elements to obtain a sorting result; obtaining a preset number of third document elements from the sorting result, and taking the preset number of third document elements as the first document element corresponding to the target question.

[0047] Figure 4is a process schematic diagram of a method for information retrieval based on a large language model according to an embodiment of the present disclosure. Referring to Figure 4 , the document is subjected to a sharding module for sharding processing to obtain a first type of document element, and a connection edge is constructed in a structure tree corresponding to the document to obtain a graph structure. The summarization of the large language model realizes the summary of the first type of document element, and then generates a vector corresponding to each first type of document element, which is stored in a vector library.

[0048] Further, based on the target question, the third document element is recalled from the vector library. Then, the ranking large language model is used to rank the recalled third document element based on semantics. As an example, the ranking large language model can rank the third document element according to the relevance of the semantics and / or the coherence of the semantics, and the third document element with high semantic relevance and high semantic coherence is arranged in the front in the ranking result, and further, the first document element arranged in the front is obtained.

[0049] Based on the first document element, the second document element with an association is retrieved from the graph structure.

[0050] In the above manner, the retrieval of the second document element is performed based on the first document element with strong semantic relevance to the target question. In addition, through the cooperation of different large language models (such as the ranking large language model, the summarization large language model, and the question and answer large language model) shown in Figure 4 , the information retrieval is realized, and different tasks, i.e., ranking, summarization, and question and answer, are performed by different large language models, thereby facilitating the convergence and maintenance of each large language model.

[0051] In some embodiments, the above information retrieval method can further include the following steps: obtaining an evaluation result of the first document element, wherein the evaluation result is used to describe the accuracy of the first document element; and adjusting the value of the preset number based on the evaluation result.

[0052] In addition to the case where the number of contexts can be controlled by specifying the level and the specified number, the size of the prompt word can also be controlled by adjusting the value of the preset number.

[0053] Therefore, if the evaluation result is used to represent that the accuracy is greater than a preset accuracy, the value of the preset number can be reduced, so as to ensure that the prompt word given to the large language model does not exceed the limit of the model.

[0054] Figure 5 is a block diagram of a device for information retrieval based on a large language model according to an embodiment of the present disclosure. Referring to Figure 5 , the device for information retrieval based on a large language model 500 includes: The first obtaining module 501 is configured to obtain a first document element corresponding to a target question, wherein a similarity between the first document element and the target question is greater than a preset similarity threshold. The searching module 502 is configured to search, in a pre-constructed graph structure, a second document element having an association relationship with the first document element, wherein the graph structure is obtained by constructing a connection edge in a structure tree corresponding to a document, the structure tree is used to describe a hierarchical relationship between nodes, each node uniquely corresponds to a document element in the document, and the connection edge is used to describe a logical association relationship or a semantic association relationship between nodes. The determining module 503 is configured to determine the first document element and the second document element as context information of the target question, wherein the context information is used to support a large language model to generate a target reply to the target question.

[0055] Optionally, the searching module 502 includes: A searching sub-module is configured to search, in a pre-constructed graph structure, a second document element having an association with the first document element, with a first target constraint as a target, wherein the first target constraint includes a specified level, and a level of the searched second document element is the specified level.

[0056] Optionally, the first target constraint further includes a specified number, the specified number is used to indicate a number of the searched second document elements, and the specified level includes a level in which the first document element is located. The searching sub-module is further configured to: search, in a level in which the first document element is located in the pre-constructed graph structure, a second document element having an association with the first document element; in a case where a number of the second document elements searched in the level in which the first document element is located does not reach the specified number, search, in the level in which the first document element is located, a second document element having an association with the first document element, with a second preset constraint as a target, and search, on the basis of the level in which the first document element is located, a second document element having an association with the first document element upwards and downwards level by level, the second preset constraint being that, among the second document elements searched upwards and downwards level by level, a number of the second document elements located above the level in which the first document element is located and a number of the second document elements located below the level in which the first document element is located are minimized.

[0057] Optionally, the first obtaining module 501 includes: A recall sub-module is configured to recall a third document element based on the target question, wherein a similarity between the third document element and the target question is greater than the preset similarity threshold. The sorting submodule is configured to sort the third document elements to obtain a sorting result, wherein the sorting is based on semantic implementation between the target question and the third document elements. The obtaining submodule is configured to obtain a preset number of third document elements from the sorting result, and take the preset number of third document elements as the first document elements corresponding to the target question.

[0058] Optionally, the information retrieval apparatus 500 based on a large language model further includes: The second obtaining module is configured to obtain an evaluation result of the first document element, wherein the evaluation result is used to describe the accuracy of the first document element. The adjusting module is configured to adjust the value of the preset number based on the evaluation result.

[0059] Optionally, the leaf nodes in the structure tree are first-type document elements, the second document elements and the first document elements belong to the first-type document elements, the connection edges include directed edges and bidirectional edges, the directed edges are connection edges between document elements of the same level and from the first-type document elements to second-type document elements, and the bidirectional edges are connection edges between the first-type document elements.

[0060] The embodiments of the modules in the information retrieval apparatus 500 based on a large language model can refer to the related embodiments of the above method, and will not be repeated here.

[0061] The embodiments of the disclosure also provide a computer readable medium having a computer program stored thereon, and the computer program is executed by a processing device to implement the steps of the above information retrieval method based on a large language model.

[0062] The embodiments of the disclosure also provide a computer program product including a computer program, and the computer program is executed by a processor to implement the steps of the above information retrieval method based on a large language model.

[0063] The embodiments of the disclosure also provide an electronic device including: A storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the steps of the above information retrieval method based on a large language model.

[0064] The following will be described with reference to the accompanying drawings: Figure 6, which shows a schematic structural diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0065] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0066] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0067] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0068] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable medium or a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0069] In some embodiments, the electronic device can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0070] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device without being incorporated into the electronic device.

[0071] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a first document element corresponding to a target question, wherein the first document element has a similarity to the target question greater than a preset similarity threshold; retrieve a second document element having an association relationship with the first document element in a pre-constructed graph structure, wherein the graph structure is obtained by constructing a connection edge in a structure tree corresponding to a document, the structure tree is used to describe a hierarchical relationship between nodes, each node uniquely corresponds to a document element in the document, and the connection edge is used to describe a logical association relationship or a semantic association relationship between nodes; and determine the first document element and the second document element as context information of the target question, wherein the context information is used to support the large language model to generate a target reply to the target question.

[0072] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0073] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0074] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the first obtaining module can also be described as a module that obtains a first document element corresponding to a target question.

[0075] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0076] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0077] The above description is merely the preferred embodiments of the present disclosure and the explanation of the principles of the applied technology. Those skilled in the art should understand that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above features can be replaced with other technical features disclosed in the present disclosure (but not limited to) that have similar functions to form technical solutions.

[0078] Moreover, while operations have been depicted in a particular order, this should not be understood as requiring such an order nor limiting it to only those operations shown and described. One of ordinary skill in the art will recognize that many of the operations can be performed in a differing order, or be performed concurrently, that some operations can be performed in any order or omitted, and that some operations can be performed in parallel. Similarly, while several specific implementation details have been discussed in the context of the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0079] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.

Claims

1. An information retrieval method based on a large language model, characterized in that: include: Acquire a first document element corresponding to a target question, wherein a similarity between the first document element and the target question is greater than a preset similarity threshold; Retrieving a second document element associated with the first document element in a pre-constructed graph structure, wherein the graph structure is obtained by constructing connecting edges in a structure tree corresponding to the document, the structure tree being used to describe a hierarchical relationship between nodes, each node uniquely corresponding to a document element in the document, and the connecting edges being used to describe a logical or semantic association relationship between the nodes; The first document element and the second document element are determined as context information of the target question, wherein the context information is used to support a large language model to generate a target response to the target question.

2. The information retrieval method according to claim 1, wherein: The retrieving a second document element associated with the first document element in the pre-constructed graph structure includes: Taking a first target constraint as a target, a second document element associated with the first document element is retrieved in a pre-constructed graph structure, wherein the first target constraint includes a specified level, and the level of the retrieved second document element is the specified level.

3. The information retrieval method according to claim 2, wherein: The first target constraint further includes a specified number, where the specified number indicates the number of the retrieved second document elements. The specified level includes the level at which the first document element is located. The step of retrieving the second document elements associated with the first document element in the pre-constructed graph structure based on the first target constraint includes: Retrieving, in a level of the pre-constructed graph structure where the first document element is located, a second document element associated with the first document element; When the number of second document elements retrieved in the level where the first document element is located does not reach the specified number, with a second preset constraint as a target, second document elements associated with the first document element are retrieved step by step upward and downward on the basis of the level where the first document element is located until the number of second document elements retrieved reaches the specified number, and the second preset constraint is: among the second document elements retrieved step by step upward and downward, the difference between the number of second document elements located above the level where the first document element is located and the number of second document elements located below the level where the first document element is located is minimized.

4. The information retrieval method according to claim 1, wherein: The obtaining of the first document element corresponding to the target question includes: Recalling a third document element based on the target question, wherein the similarity between the third document element and the target question is greater than the preset similarity threshold; Sorting the third document element to obtain a sorting result, wherein the sorting is achieved based on the semantics between the target question and the third document element; A preset number of third document elements are obtained from the sorting results, and the preset number of third document elements are used as first document elements corresponding to the target question.

5. The information retrieval method according to claim 4, characterized in that: The information retrieval method further includes: Obtaining an evaluation result of the first document element, wherein the evaluation result is used to describe the accuracy of the first document element; Based on the evaluation result, the value of the preset number is adjusted.

6. The information retrieval method according to claim 1, wherein: The leaf nodes in the structure tree are document elements of the first type, the second document element and the first document element both belong to the first type of document elements, and the connecting edges include directed edges and bidirectional edges. The directed edges are connecting edges between document elements of the same level and from the first type of document element to the second type of document element. The bidirectional edges are connecting edges between document elements of the first type.

7. An information retrieval device based on a large language model, characterized in that: include: A first acquisition module is configured to acquire a first document element corresponding to a target question, wherein a similarity between the first document element and the target question is greater than a preset similarity threshold; a retrieval module configured to retrieve, from a pre-constructed graph structure, a second document element associated with the first document element, wherein the graph structure is obtained by constructing connecting edges in a structure tree corresponding to the document, the structure tree being configured to describe a hierarchical relationship between nodes, each node uniquely corresponding to a document element in the document, and the connecting edges being configured to describe a logical or semantic association between the nodes; A determination module is used to determine the first document element and the second document element as context information of the target question, wherein the context information is used to support a large language model to generate a target response to the target question.

8. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the information retrieval method according to any one of claims 1 to 6 are implemented.

9. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the information retrieval method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the information retrieval method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Automobile question answering method based on intelligent agent

    CN118051596A

  • Document segmentation method and device, computer equipment and storage medium

    CN119474250A

  • Multi-document question and answer method and device based on multi-head self-attention and hierarchical enhancement

    CN119537559A

  • Intelligent question answering method and system based on question index retrieval enhancement generation technology

    CN120124640A

  • Search system using hierarchical metadata based on retrieval augmented generation and method thereof

    KR102824126B1