Method and device for obtaining question reply through large language model
By generating a tree-like directory structure for professional documents and using a large language model to locate relevant chapter content to generate responses, the problems of poor response quality and resource waste in existing technologies are solved, achieving efficient and accurate responses to professional documents.
Patent Information
- Application Number
- CN202511473534.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies, when obtaining responses to professional documents through large language models, are prone to disrupting the hierarchical structure and logical integrity of the documents, resulting in poor response quality and excessive consumption of computational resources.
By generating a tree-structured document directory of professional documents, and using a large language model to infer the chapter directory nodes related to the question, answers can be extracted and generated, avoiding attention dilution and waste of computing resources caused by full-text input.
Without disrupting the document structure, it improves the accuracy and speed of responses while reducing computational resource consumption.
Smart Images

Figure CN120929580A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to one or more embodiments in the field of large language model technology, and more particularly to a method and apparatus for obtaining answers to questions through a large language model. Background Technology
[0002] Large Language Models (LLMs) are deep learning models based on natural language processing, trained on massive text corpora containing hundreds of millions or more parameters. Currently, LLMs are increasingly widely used in domain-specific question answering, especially in scenarios requiring high-precision answers to questions, such as medicine, law, and engineering. Users often hope to use LLMs to quickly obtain accurate answers to their questions based on lengthy and complex professional documents. Summary of the Invention
[0003] The embodiments in this specification aim to provide a method and apparatus for obtaining answers to questions through a large language model. This method can generate a global document directory structure, which guides the large model for global content location. This allows for more accurate location of relevant chapters from professional documents without disrupting their hierarchical structure and logical integrity. The large language model then generates answers to user questions based on the located chapters, improving the quality of the responses output by the large model and addressing the shortcomings of existing technologies.
[0004] Based on the first aspect, a method for obtaining question answers through a large language model is provided, including:
[0005] The system obtains the target question input by the user and a tree-structured document directory based on the target document sent by the user. The document directory includes multiple directory nodes and connections between the directory nodes. The directory nodes are used to indicate chapters in the target document, and the connections are used to indicate the hierarchical relationships between the chapters. Based on the target question and the document directory, and using a preset large language model, the system obtains the first directory node in the document directory that is related to the target question.
[0006] Extract the content of the target chapter pointed to by the first directory node from the target document, and obtain the target answer corresponding to the target question based on the target question and the content of the target chapter, using the large language model.
[0007] In one possible implementation, obtaining the target question input by the user and a tree-structured document directory based on the target document sent by the user includes:
[0008] Obtain the first structured text used to store the document directory;
[0009] Based on the target question and the document directory, and using a pre-defined big oracle model, the first directory node in the document directory related to the target question is obtained, including:
[0010] The target question, the first structured text, and the indications of the directory nodes related to the target question in the document directory determined based on the target question and the first structured text are input into the large language model to obtain the first directory node.
[0011] In one possible implementation, obtaining a first structured text for storing the document directory includes:
[0012] Obtain the target document sent by the user. If the target document is a structured document, the structured document contains preset hierarchical headings.
[0013] Extract the hierarchical headings contained in the target document, generate the document directory based on the hierarchical headings, and save the document directory in the first structured text.
[0014] In one possible implementation, the first structured text includes one of JavaScript object symbol text, Extended Markup Language text, and YXML text.
[0015] In one possible implementation, the structured document is one or more of Markdown, HTML, and LaTeX documents.
[0016] In one possible implementation, the method further includes:
[0017] If the target document does not belong to the structured document, the target document is converted into a first document, and the first document belongs to the structured document;
[0018] Extract the hierarchical headings contained in the first document, generate the document directory based on the hierarchical headings, and save the document directory in the first structured text. In one possible implementation, based on the target question and the content of the target chapter, and using the large language model, obtain the target answer corresponding to the target question, including:
[0019] The target question, the content of the target chapter, and the instruction to determine the answer corresponding to the target question based on the target question and the content of the target chapter are input into the large language model to obtain the target answer.
[0020] In one possible implementation, the target document is a professional document in a preset field, which may include one of the fields of medicine, law, or engineering.
[0021] According to the second aspect, an apparatus for obtaining question answers through a large language model is provided, comprising:
[0022] The acquisition unit is configured to acquire a target question input by the user and a tree-structured document directory obtained based on the target document sent by the user. The document directory includes multiple directory nodes and connections between the directory nodes. The directory nodes are used to indicate chapters in the target document, and the connections are used to indicate the subordinate relationships between the chapters. Based on the target question and the document directory, and using a preset large language model, the unit obtains a first directory node in the document directory that is related to the target question.
[0023] The response unit is configured to extract the content of the target chapter pointed to by the first directory node from the target document, and obtain the target response corresponding to the target question based on the target question and the content of the target chapter, using the large language model.
[0024] According to a third aspect, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in the first aspect.
[0025] According to a fourth aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method described in the first aspect.
[0026] By utilizing one or more of the methods, devices, computing equipment, and storage media mentioned above, a global document directory structure can be generated. This document directory structure guides a large model for global content location, thereby more accurately locating relevant chapters from the professional document without disrupting its hierarchical structure and logical integrity. Subsequently, the large language model generates answers to user questions based on the located chapters, improving the quality of the responses output by the large model. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This diagram illustrates a scheme for obtaining answers to professional questions using a large language model. Figure 2 This diagram illustrates another approach to obtaining answers to professional questions using a large language model. Figure 3 A schematic diagram illustrating a method for obtaining a question answer using a large language model according to an embodiment of this specification; Figure 4 A flowchart illustrating a method for obtaining a question answer using a large language model according to an embodiment of this specification is shown. Figure 5 This diagram illustrates a structural diagram of an apparatus for obtaining a question answer using a large language model according to an embodiment of this specification. Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0030] In this specification, the Large Language Model (LLM) may also be referred to simply as the Large Model. A Large Language Model is a natural language processing model based on deep learning techniques, typically with billions to hundreds of billions or even more parameters, possessing powerful language understanding and generation capabilities. Large Language Models can employ the Transformer architecture or its variants (such as GPT, BERT, etc.), which utilizes an attention mechanism to globally model sequential data, efficiently handling long-distance dependencies and thus performing exceptionally well in natural language tasks. Large Language Models learn the statistical features and semantic relationships of language through pre-training on large-scale corpora, giving them excellent generalization capabilities. The core capabilities of Large Language Models include, but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Its usage typically includes two modes: direct inference and fine-tuning. In direct inference mode, the user guides the Large Language Model to generate specific outputs by designing prompts. Cue words can be task descriptions or instructions in text form, used to stimulate the semantic understanding and generation capabilities of large language models. In fine-tuning mode, large language models are further trained on small-scale datasets in specific domains to optimize their performance on specific tasks. The powerful generalization ability and flexibility of large language models make them an important tool in the field of artificial intelligence, providing efficient and accurate solutions for automated text generation and understanding.
[0031] In some embodiments, large language models can also understand and generate data from other modalities (such as visual and audio data). In this case, large language models can also be called multimodal large language models (MLLMs). MLLMs provide a richer and more natural interactive experience by integrating multiple types of input and output, such as text, images, and sound. The core advantage of MLLMs lies in their ability to process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze an image and generate descriptive text, or generate a corresponding image based on a text description. This cross-modal understanding and generation capability makes MLLMs widely applicable across multiple fields.
[0032] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, published on March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), and will not be repeated here.
[0033] With the increasing application of large language models in professional domain knowledge question answering, especially in scenarios requiring high-precision answers to questions such as medicine, law, and engineering, users often hope to quickly obtain accurate answers to their questions from long and complex professional documents (e.g., clinical practice guidelines, industry technical standards, regulatory documents, etc.) using large language models. The current mainstream approach to obtaining answers to professional questions using large language models primarily involves pre-segmenting the professional document into fixed-length text chunks. Then, text retrieval is used to identify the text chunks most relevant to the user's question, and these chunks are input into the large language model to generate the answer. Figure 1 This diagram illustrates a scheme for obtaining answers to professional questions using a large language model. Figure 1 As shown, for example, a professional document input by a user can be segmented into multiple text segments based on a fixed segment length. Then, the user-input question is obtained, and relevance matching is performed between the question and each text segment to determine the text segment with the highest relevance to the question. Specifically, for example, TF-IDF (Term Frequency-Inverse Document Frequency) statistics can be used to determine the text segment with the highest relevance to the question. Alternatively, the question and each text segment can be converted into corresponding text vectors, and the text segment with the highest relevance to the question can be determined based on the similarity between the text vectors of multiple text segments and the text vector of the question. Subsequently, the text segment with the highest relevance to the question can be input into a large language model to obtain an answer to the question.
[0034] However, this approach also has the following problems: On the one hand, because professional documents are usually long, with clear internal hierarchical structures and rigorous content logic, simple fixed-length text segmentation can easily destroy the original hierarchical structure and logical integrity of the text, causing the text fragments input to the large language model to often lack key contextual information, resulting in poor quality of the large language model's output response. On the other hand, both word frequency relevance matching and vector similarity matching lack an understanding or processing of the overall document structure, thereby reducing the accuracy of locating the relevant chapters from the professional document, making the relevant text fragments input to the large language model often inaccurate, also leading to poor quality of the large language model's output response.
[0035] Figure 2 This diagram illustrates another approach to obtaining answers to specialized questions using a large language model. For example... Figure 2 As shown, both the full content of a professional document and the user's question can be input into a large language model to obtain a response to the question. However, this approach also has the following problems: First, although the information input into the large language model retains contextual integrity, the full content of professional documents is usually very long, often leading to input overflow due to the limitation of the input token length of the large model itself. Second, excessively long input content can cause the large model's inherent attention mechanism to become scattered, making it difficult to focus on the core content related to the question, resulting in a decrease in the accuracy of the response. Third, inputting the full content into the large model for each response not only slows down the speed at which the large language model completes reasoning and outputs a response, but also significantly increases the computational resources consumed by reasoning.
[0036] To address the aforementioned technical problems, this specification provides an embodiment of a method for obtaining answers to questions using a large language model. Figure 3 This diagram illustrates a method for obtaining a question answer using a large language model according to an embodiment of this specification. Figure 3 As shown, firstly, a tree-structured document directory can be generated based on the user-inputted professional document. The document directory includes multiple directory nodes indicating chapters within the target document, as well as connections between these nodes indicating the hierarchical relationships between chapters. Then, the document directory and the user-inputted question can be input into a large language model to obtain directory nodes in the document directory related to the user's question. Furthermore, the content of the chapter corresponding to that directory node can be extracted from the professional document, and this chapter content, along with the user's question, can be input into the large language model to obtain an answer to the user's question.
[0037] The advantages of this method are twofold: First, compared to methods that determine question answers based on fixed-length fragments of professional documents using a large model, this method generates a document directory tree for the professional document. The user's question and the document directory tree are then input into a large language model, which infers the directory nodes (entries) of the relevant chapters. The content of the corresponding chapters is then extracted from the professional document, and the large language model generates the question answer based on this content. Thus, with a global document directory structure, without disrupting the hierarchical structure and logical integrity of the original professional document, the large model is guided to perform global content location, significantly improving the completeness and accuracy of locating substantially relevant chapter content from the professional document. This, in turn, improves the quality of the answer output by the large language model based on more complete and accurate chapter content. Second, compared to methods that determine question answers based on the entire content of the professional document using a large model, this method inputs a lightweight document directory tree into the large language model, guiding it to infer the directory nodes of the relevant chapters. This allows large language models to perform global localization with a significantly smaller input volume than the full text, avoiding the attention dilution problem that often occurs when inputting the entire text. This improves the accuracy of subsequent question responses generated by the large language model based on the localized chapter content. Furthermore, since each response only requires the relevant chapters obtained after localization by the large model, it not only greatly reduces the probability of input overflow but also increases the speed at which the large language model outputs responses and reduces the computational resources consumed by the large model's inference.
[0038] The following section further elaborates on the detailed process of this method. Figure 4 A flowchart illustrating a method for obtaining a question answer using a large language model according to an embodiment of this specification is shown. Figure 4 The method includes at least the following steps:
[0039] Step S401: Obtain the target question input by the user and a tree-structured document directory obtained based on the target document sent by the user. The document directory includes multiple directory nodes and connections between the directory nodes. The directory nodes are used to indicate chapters in the target document, and the connections are used to indicate the subordinate relationships between the chapters. Based on the target question and the document directory, and based on a preset large language model, obtain the first directory node in the document directory that is related to the target question.
[0040] Step S403: Extract the content of the target chapter pointed to by the first directory node from the target document, and obtain the target answer corresponding to the target question based on the target question and the content of the target chapter, using the large language model.
[0041] First, in step S401, the target question input by the user and a tree-structured document directory based on the target document sent by the user are obtained. The document directory may include multiple directory nodes and connections between them. Directory nodes can be used to indicate chapters in the target document, and connections can be used to indicate the hierarchical relationships between chapters. In this step, the target question input by the user can be a question described in natural language. In different embodiments, the target question input by the user can be different specific questions, and this specification does not limit this.
[0042] In different embodiments, the target document sent by the user can be a different specific document. In one embodiment, the target document can be, for example, a professional document in a preset field. In a specific embodiment, the preset field may include one of the fields of medicine, law, or engineering.
[0043] Typically, a target document can contain multiple chapters. A chapter (or section) is a logically divided unit of content within a document, used to organize the hierarchical structure of information within the document. Multiple chapters within a target document can have a hierarchical relationship. For example, document D2 contains chapters A1, A11, and A12, where A11 and A12 are sub-chapters of A1, and the content of A11 and A12 belongs to the content of A1. The hierarchical relationship between chapters can also form a multi-level nested hierarchical relationship; for example, A12 contains sub-chapters A121 and A122, and the content of A121 and A122 also belongs to A1. In one embodiment, a chapter can have a title to indicate its content or a number to identify its hierarchical relationship with other chapters. The specific form of the chapter number may differ in different specific embodiments, and this specification does not limit this. In one example, the chapter number may include one or more of numbers, letters, punctuation marks, and text.
[0044] In different embodiments, the specific method of obtaining the document directory based on the target document can vary. In one embodiment, the target document sent by the user can be obtained. If the target document is a structured document, the structured document contains preset hierarchical headings. The hierarchical headings contained in the target document are extracted, and the document directory is generated based on the hierarchical headings. In one example, the structured document can be a Markdown document, a HyperText Markup Language (HTML) document, or a LaTeX document. Furthermore, the document directory can also be stored, for example, in a first structured text. In different specific embodiments, the representation of the hierarchical headings contained in the target document can vary. For example, in one specific embodiment, the document D1 sent by the user is, for example, a Markdown document. The hierarchical headings in the Markdown document can include one or more levels of headings such as Level 1 (represented by "#"), Level 2 (represented by "##"), Level 3 (represented by "###"), Level 4 (represented by "####"), Level 5 (represented by "#####"), and Level 6 (represented by "#####"). Furthermore, the headings at these hierarchical levels in D1 can be extracted to generate a document table of contents. In different specific embodiments, the specific algorithm for generating the document table of contents based on the hierarchical headings can vary. In one specific embodiment, for example, a recursive tree construction algorithm can be used to generate the document table of contents. In another specific embodiment, the directory nodes in the generated document table of contents can contain the title or number of the chapter they indicate.
[0045] In another embodiment, chapter summaries can be generated for each chapter based on the content of each chapter in the target document. Furthermore, the directory nodes in the generated document directory can also contain summaries of the chapters they indicate. In conventional document directory entries (nodes), only titles and codes are typically included to indicate chapters, without summaries. Adding summaries to each directory entry can lead to a less clear and less engaging user experience due to the complexity of the content. However, in this embodiment, by adding chapter summaries to each chapter node in the directory, the large model can more accurately identify chapters relevant to the user's question in subsequent steps. Since the document directory itself is used as input for the large model and differs from conventional human user browsing, it avoids the problem of a poor user experience.
[0046] Structured text refers to text whose content is organized in a predefined format. The structure of data within the text is typically defined by tags, marks, or hierarchical structures. In different specific embodiments, the text structure upon which the structured text is based can vary. In one specific embodiment, the first structured text may include one of the following: JavaScript Object Notation (JSON) text, Extensible Markup Language (XML) text, or YXML (YAML Ain't Markup Language) text.
[0047] In some scenarios, the document input by the user may not be a structured document, and a corresponding structured document can be generated based on the user input. Then, a tree-structured document directory is generated based on this structured document. Therefore, in one embodiment, if the target document does not belong to the structured document, the target document can be converted into a first document, which belongs to the structured document; the hierarchical headings contained in the first document are extracted, and the document directory is generated based on the hierarchical headings, and the document directory is saved in the first structured text.
[0048] It is important to note that the order in which the target question and the document directory are obtained may differ in different embodiments. For example, in one example, the target document input by the user can be obtained first, and after generating the document directory based on the target document, the target question input by the user can be obtained. In another example, both the user question input by the user and the target document can be obtained simultaneously. Then, the document directory is generated based on the target document.
[0049] After obtaining the target question and document directory, a first directory node related to the target question in the document directory can be obtained based on a preset large language model, according to the target question and document directory. In different embodiments, the specific method of obtaining the first directory node based on the large language model can vary. In the embodiment of obtaining the first structured text described above, the target question, the first structured text, and the indication of the directory node related to the target question in the document directory determined based on the target question and the first structured text can be input into the large language model to obtain the first directory node. In a specific example, the target question, the first structured text, and the indication of the directory node related to the target question in the document directory determined based on the target question and the first structured text can be used, for example, to construct a first prompt word, and then input the first prompt word into the large language model to obtain the first directory node. In different specific embodiments, the specific text content contained in the first prompt word can vary, and this specification does not limit this.
[0050] After obtaining the first directory node, in step S403, the content of the target chapter pointed to by the first directory node can be extracted from the target document. Then, based on the target question and the content of the target chapter, and using the large language model, the target answer corresponding to the target question is obtained.
[0051] In different embodiments, the specific method of obtaining the first directory node based on the large language model can vary. In one embodiment, the target question, the content of the target chapter, and an indication of the answer corresponding to the target question determined based on the target question and the content of the target chapter can be input into the large language model to obtain an answer to the target question (e.g., the target answer). In a specific example, an indication of the answer corresponding to the target question can be generated based on the target question, the content of the target chapter, and the content of the target question and the content of the target chapter, for example, by constructing a second prompt word, and then inputting the second prompt word into the large language model to obtain the target answer. In different specific embodiments, the specific text content contained in the second prompt word can also be different, and this specification does not limit this.
[0052] According to yet another embodiment, an apparatus for obtaining a question answer through a large language model is also provided. Figure 5 This diagram illustrates a structural diagram of an apparatus for obtaining question answers using a large language model according to an embodiment of this specification, such as... Figure 5 As shown, the system 500 includes:
[0053] The acquisition unit 501 is configured to acquire a target question input by the user and a tree-structured document directory obtained based on the target document sent by the user. The document directory includes multiple directory nodes and connections between the directory nodes. The directory nodes are used to indicate chapters in the target document, and the connections are used to indicate the subordinate relationships between the chapters. Based on the target question and the document directory, and based on a preset large language model, a first directory node related to the target question in the document directory is obtained.
[0054] The response unit 502 is configured to extract the content of the target chapter pointed to by the first directory node from the target document, and obtain the target response corresponding to the target question based on the target question and the content of the target chapter, using the large language model.
[0055] In another aspect, embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform any of the methods described above.
[0056] In another aspect, embodiments of this specification provide a computing device, including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement any of the methods described above.
[0057] It should be understood that the descriptions such as "first" and "second" in this article are merely for the sake of simplicity in description and to distinguish similar concepts, and do not have any other limiting function.
[0058] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0059] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0060] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0061] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.
[0062] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0063] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0064] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0065] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0066] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0067] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0068] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0069] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0070] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0071] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0072] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.
Claims
1. A method for obtaining question answers using a large language model, comprising: The system obtains the target question input by the user and a tree-structured document directory based on the target document sent by the user. The document directory includes multiple directory nodes and connections between the directory nodes. The directory nodes are used to indicate the chapters in the target document, and the connections are used to indicate the hierarchical relationships between the chapters. Based on the target question and the document directory, and using a preset large language model, the first directory node in the document directory that is related to the target question is obtained; Extract the content of the target chapter pointed to by the first directory node from the target document, and obtain the target answer corresponding to the target question based on the target question and the content of the target chapter, using the large language model.
2. The method according to claim 1, wherein, The system retrieves the target question input by the user and a tree-structured document directory based on the target document submitted by the user, including: Obtain the first structured text used to store the document directory; Based on the target question and the document directory, and using a pre-defined big oracle model, the first directory node in the document directory related to the target question is obtained, including: The target question, the first structured text, and the indications of the directory nodes related to the target question in the document directory determined based on the target question and the first structured text are input into the large language model to obtain the first directory node.
3. The method according to claim 2, wherein, Obtaining a first structured text for storing the document directory includes: Obtain the target document sent by the user. If the target document is a structured document, the structured document contains preset hierarchical headings. Extract the hierarchical headings contained in the target document, generate the document directory based on the hierarchical headings, and save the document directory in the first structured text.
4. The method according to claim 3, wherein, The first structured text includes one of JavaScript object symbol text, Extensible Markup Language text, and YXML text.
5. The method according to claim 1, wherein, The structured document is one or more of the following: Markdown document, HTML document, and LaTeX document.
6. The method according to claim 3, further comprising: If the target document does not belong to the structured document, the target document is converted into a first document, and the first document belongs to the structured document; Extract the hierarchical headings contained in the first document, generate the document directory based on the hierarchical headings, and save the document directory in the first structured text.
7. The method according to claim 1, wherein, Based on the target question and the content of the target chapter, and using the large language model, the target answer corresponding to the target question is obtained, including: The target question, the content of the target chapter, and the instruction to determine the answer corresponding to the target question based on the target question and the content of the target chapter are input into the large language model to obtain the target answer.
8. The method according to claim 1, wherein, The target document is a professional document in a preset field, which may include one of the fields of medicine, law, or engineering.
9. An apparatus for obtaining a question answer using a large language model, comprising: The acquisition unit is configured to acquire the target question input by the user and a tree-structured document directory obtained based on the target document sent by the user. The document directory includes multiple directory nodes and connections between the directory nodes. The directory nodes are used to indicate chapters in the target document, and the connections are used to indicate the hierarchical relationships between the chapters. Based on the target question and the document directory, and using a preset large language model, the first directory node in the document directory that is related to the target question is obtained; The response unit is configured to extract the content of the target chapter pointed to by the first directory node from the target document, and obtain the target response corresponding to the target question based on the target question and the content of the target chapter, using the large language model.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-8.
11. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-8.
Citation Information
Patent Citations
Chapter catalogue screening method and device
CN106294292A
Question response method and device based on large language model
CN117235226A
Question answering method and device based on large model, electronic equipment and medium
CN119357364A
Paper question answering method and system based on large language model
CN120124638A
Question-answering processing method, and device, product and storage medium
WO2025181561A1