Question and answer method, device, related equipment and computer program product
By building a structured knowledge base and performing two-stage fine-tuning training of large models, combined with RAG technology, the "illusion" problem of large models in specific domain Q&A applications and the problem of model fine-tuning and prediction separation between model fine-tuning and prediction are solved, significantly improving the Q&A performance.
Patent Information
- Application Number
- CN202411719722.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-11-28
AI Technical Summary
Existing large models have "illusion" problems in question-and-answer applications in specific fields, generating content that does not match the facts, and existing methods fail to make full use of original knowledge documents, resulting in the impact of the quality and relevance of the search information, and there is a separation between model fine-tuning and prediction.
By building a structured knowledge base, analyzing original knowledge documents in a specific field, generating a structured document tree, and converting knowledge fragments and their structural information into vector storage, using a pre-trained large model for two-stage fine-tuning training, combining RAG technology, integrating user query and target knowledge fragments to form prompt input data.
It significantly improves the Q&A application performance of large models in specific fields, solves the problem of separation between model fine-tuning and prediction, and improves the ability to deal with problems in specific fields.
Smart Images

Figure CN119227813B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a question-answering method, device, related equipment, and computer program product. Background Art
[0002] Large models, based on huge neural network architectures, possess powerful language understanding and generation capabilities and perform excellently in many natural language processing tasks. However, when dealing with specific domain knowledge, large models have the "hallucination" problem, that is, they may generate content inconsistent with facts. To solve this problem, the Retrieval-Augmented Generation (RAG) technology has emerged, which supplements the knowledge of large models by retrieving external knowledge bases to improve their accuracy. In addition, large model fine-tuning technology can be used to adapt pre-trained large models to downstream tasks.
[0003] However, existing large models still have some limitations in question-answering applications in specific domains. Some methods only rely on RAG technology and fail to fully utilize the original knowledge documents, affecting the quality and relevance of the retrieved information; other methods only use large model supervised fine-tuning and ignore the knowledge injected by RAG, resulting in the large model not learning how to utilize the injected knowledge during the fine-tuning stage, but relying on the injected knowledge during the prediction stage. This inconsistency causes the model to be unable to effectively utilize the retrieved injected knowledge during prediction, that is, it leads to the disconnection between model fine-tuning and model prediction. Neither of these two methods can fully improve the performance of large models in question-answering applications in specific domains. Summary of the Invention
[0004] In view of the above problems, the present application is proposed to provide a question-answering method, device, related equipment, and computer program product to improve the performance of large models in question-answering applications in specific domains. The specific solutions are as follows:
[0005] In a first aspect, a question-answering method is provided, including:
[0006] Obtaining a user query;
[0007] Retrieving, using a preset structured knowledge base, a target knowledge fragment related to the user query, where the structured knowledge base contains knowledge fragments in a specific domain;
[0008] Integrating the target knowledge fragment and the user query to form prompt input data;
[0009] Inputting the prompt input data into a pre-trained large model to obtain an answer matching the user query output by the large model;
[0010] Among them, the pre-trained large model is obtained through two-stage fine-tuning training: in the first stage, the initial large model is fine-tuned using training data in a specific domain to obtain a large model after the first-stage training, and the training data includes user query samples and matching answer labels; in the second stage, on the basis of the training data, knowledge fragments retrieved from the structured knowledge base and matching the user query samples are added as new training data to fine-tune the large model after the first-stage training to obtain a large model after the second-stage training.
[0011] Preferably, the construction process of the structured knowledge base includes:
[0012] Parse the text content in the original knowledge document in a specific domain into a structured document tree, where the leaf nodes in the structured document tree represent knowledge fragments, and the nodes on the path from the root node to the parent node of each leaf node represent the structural information of the knowledge fragment corresponding to the leaf node in the original knowledge document;
[0013] Combine the knowledge fragment corresponding to each leaf node and its structural information, convert the combined text into a vector, and store the knowledge fragment corresponding to each leaf node and the converted vector in the structured knowledge base correspondingly.
[0014] Preferably, the construction process of the structured knowledge base further includes:
[0015] Convert the tables in the original knowledge document into target formats that can be processed by the large model;
[0016] Use the pre-trained large model to summarize the target-format tables to obtain summary information;
[0017] Convert the summary information into a vector, and store the summary information as a knowledge fragment and the converted vector in the structured knowledge base correspondingly.
[0018] Preferably, the construction process of the structured knowledge base further includes:
[0019] Correspondingly store the structural information corresponding to the knowledge fragment in the structured knowledge base.
[0020] Preferably, the structured knowledge base includes: the corresponding relationship between knowledge fragments and vectors;
[0021] Correspondingly, the retrieval of the target knowledge fragment related to the user query using the pre-set structured knowledge base includes:
[0022] Vectorize the user query to obtain a user query vector;
[0023] Retrieve target knowledge fragments related to the user query vector from the structured knowledge base based on vector relevance.
[0024] Preferably, the structured knowledge base includes: the correspondence between knowledge fragments and vectors;
[0025] Accordingly, the process of retrieving target knowledge fragments related to the user query using the pre - established structured knowledge base includes:
[0026] Vectorize the user query to obtain a user query vector;
[0027] Based on vector relevance, retrieve candidate knowledge fragments related to the user query vector from the structured knowledge base;
[0028] Use a pre - trained relevance determination model to determine the relevance between each candidate knowledge fragment and the user query respectively, and screen the top N candidate knowledge fragments with the highest relevance as the target knowledge fragments, where the relevance determination model is trained using knowledge fragment samples and user query samples marked with relevance scores.
[0029] Preferably, the structured knowledge base includes: the structural information of knowledge fragments in the original knowledge document;
[0030] Then, after retrieving the target knowledge fragments related to the user query using the pre - established structured knowledge base, it further includes:
[0031] For the retrieved target knowledge fragments related to the user query, use the structural information of the target knowledge fragments in the original knowledge document to reconstruct the target knowledge fragments to obtain the reconstructed target knowledge fragments.
[0032] Preferably, the structural information includes: in the structured document tree parsed from the text content of the original knowledge document, the path information from the leaf node corresponding to each knowledge fragment to the root node;
[0033] Then, the process of reconstructing the target knowledge fragments using the structural information of the target knowledge fragments in the original knowledge document includes:
[0034] If two leaf nodes corresponding to two target knowledge fragments belong to the same parent node and are consecutive in the original knowledge document, then remove the duplicate characters in the two target knowledge fragments and splice them in the original text order of the two target knowledge fragments in the original knowledge document to obtain the reconstructed target knowledge fragments.
[0035] Preferably, the structure information includes: in the structured document tree parsed from the text content of the original knowledge document, the path information from the leaf node corresponding to each knowledge segment to the root node;
[0036] Then, the process of reconstructing the target knowledge segment by using the structure information of the target knowledge segment in the original knowledge document includes:
[0037] If there are at least two leaf nodes corresponding to the target knowledge segments sharing a common parent node, and the knowledge segments corresponding to more than a preset proportion of the leaf nodes under the parent node are included in each of the recalled target knowledge segments, then all the knowledge segments corresponding to the leaf nodes under the parent node are recalled and merged into a reconstructed target knowledge segment.
[0038] Preferably, the structure information includes: in the structured document tree parsed from the text content of the original knowledge document, the path information from the leaf node corresponding to each knowledge segment to the root node;
[0039] The process of reconstructing the target knowledge segment by using the structure information of the target knowledge segment in the original knowledge document includes:
[0040] For each target knowledge segment among the recalled target knowledge segments:
[0041] Judge whether the number of knowledge segments corresponding to all the leaf nodes under the parent node to which the leaf node corresponding to the target knowledge segment belongs is included in the recalled target knowledge segments exceeds a preset proportion. If not, further judge whether the relevance between the target knowledge segment and the user query is lower than a preset threshold. If so, delete the target knowledge segment.
[0042] Preferably, the specific field is the communication field.
[0043] In a second aspect, a question and answer device is provided, including:
[0044] A user query acquisition unit, configured to acquire a user query;
[0045] A target knowledge segment retrieval unit, configured to retrieve target knowledge segments related to the user query by using a preset structured knowledge base, where the structured knowledge base includes knowledge segments in a specific field;
[0046] A prompt input data generation unit, configured to integrate the target knowledge segment and the user query to form prompt input data;
[0047] A question and answer unit, configured to input the prompt input data into a pre-trained large model to obtain an answer matching the user query output by the large model;
[0048] Among them, the pre-trained large model is obtained through two-stage fine-tuning training: in the first stage, the initial large model is fine-tuned using training data in a specific domain to obtain a large model after the first-stage training. The training data includes user query samples and matching answer labels; in the second stage, knowledge fragments retrieved from the structured knowledge base and matching the user query samples are added to the training data as new training data to fine-tune the large model after the first-stage training, obtaining a large model after the second-stage training.
[0049] In a third aspect, an electronic device is provided, including: a memory and a processor;
[0050] The memory is used to store programs;
[0051] The processor is used to execute the program to implement the steps of the question-answering method described in any one of the foregoing first aspects of the present application.
[0052] In a fourth aspect, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the question-answering method described in any one of the foregoing first aspects of the present application are implemented.
[0053] In a fifth aspect, a computer program product is provided. The computer program product is tangibly stored in a non-transitory computer storage medium and includes machine-executable instructions. When the machine-executable instructions are executed by a device, the steps of the question-answering method described in any one of the foregoing first aspects of the present application are implemented.
[0054] By means of the above technical solutions, the present application retrieves target knowledge fragments related to user queries using a structured knowledge base containing knowledge fragments in a specific domain, and then integrates the user queries and the target knowledge fragments to form prompt input data. Finally, the prompt input data is input into a large model obtained through two-stage fine-tuning training. In the first stage of fine-tuning training of the large model, the use of training data in a specific domain improves the initial large model's ability to understand information in a specific domain and the model's instruction compliance. In the second stage of fine-tuning training of the large model, the RAG technology and the structured knowledge base are integrated, enabling the large model to effectively utilize the retrieved knowledge fragments and enhance its ability to handle problems in a specific domain; this method realizes the deep integration of the large model fine-tuning technology and the RAG technology, enabling the model to learn how to utilize the injected knowledge fragments during the model fine-tuning stage, so that the injected knowledge fragments can be effectively utilized during the prediction stage, solving the problem of the disconnection between model fine-tuning and prediction in the prior art, and significantly improving the performance of the large model in question-answering applications in a specific domain. Description of the Drawings
[0055] Upon reading the following detailed description of the preferred embodiments, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0056] Figure 1 It is a schematic diagram of an implementation system architecture of the Q&A method provided by an embodiment of the present application;
[0057] Figure 2 It is a schematic diagram of a terminal structure provided by an embodiment of the present application;
[0058] Figure 3 It is a schematic diagram of a server structure provided by an embodiment of the present application;
[0059] Figure 4 It is a schematic flowchart of the Q&A method provided by an embodiment of the present application;
[0060] Figure 5 It exemplifies a schematic diagram of a structured document tree;
[0061] Figure 6 It exemplifies a schematic diagram of the architecture of a Q&A system;
[0062] Figure 7 It is a schematic diagram of the structure of a Q&A device provided by an embodiment of the present application;
[0063] Figure 8 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0064] Before introducing the solution of the present application, first, the English terms involved in this article are explained:
[0065] prompt: An instruction. When interacting with an AI (such as an artificial intelligence model), an instruction that needs to be sent to the AI, which can be a text description, such as "Please recommend a pop music for me" when you interact with the AI, or a parameter description in a certain format, such as asking the AI to draw a picture according to a certain format and needing to describe the relevant drawing parameters.
[0066] Large Models: In the field of artificial intelligence, large models generally refer to large-scale pre-trained models. These models are called "large" because they are pre-trained on a large amount of data and can be transferred to a variety of downstream tasks. Their full English name is Large Pre-Trained Models or Large-Scale Pre-Training Models. The characteristics of large models are their huge scale, containing billions or even more parameters, which helps them learn complex patterns in the data. The emerging capabilities of large models include, but are not limited to: in-context learning, instruction following, code generation, step-by-step reasoning ability, etc. Large models can include large language models (LLMs), as well as multimodal large models. Large language models are mainly used to process text-modal data, and multimodal large models further integrate multimodal capabilities on the basis of large language models and can process information of multiple modalities, such as images, text, audio, etc.
[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0068] The present application provides a question-and-answer method, which can be applied to the Figure 1 system architecture as shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 illustrated by taking one server as an example).
[0069] Either the terminal 100 or the server 200 can be used alone to execute the question-and-answer method provided in the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the question-and-answer method provided in the embodiments of the present application.
[0070] Next, describe Figure 1 the product form of the terminal 100;
[0071] The terminal 100 in the embodiments of this application may be a mobile phone, a tablet computer, a teaching large screen, a wearable device, a vehicle-mounted device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of this application do not impose any restrictions on this.
[0072] Figure 2 Fig. 4 shows an optional schematic diagram of the hardware structure of the terminal 100.
[0073] Referring to Figure 2 As shown in Fig. 4, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160, a speaker 161, a microphone 162, a headphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190, and other components. Those skilled in the art can understand that Figure 2 This is merely an example of a terminal or a multifunctional device, and does not constitute a limitation on the terminal or the multifunctional device. It may include more or fewer components than shown in the figure, or combine certain components, or different components.
[0074] The input unit 130 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the portable multifunctional device. Specifically, the input unit 130 may include a touch screen 131 and / or other input devices 132. The touch screen 131 can collect touch operations of the user thereon or nearby (such as operations of the user using a finger, a joint, a stylus, or any suitable object on or near the touch screen), and drive corresponding connection devices according to a preset program. The touch screen can detect the touch actions of the user on the touch screen, convert the touch actions into touch signals and send them to the processor 170, and can receive commands sent by the processor 170 and execute them; the touch signals at least include contact coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch screen. In addition to the touch screen 131, the input unit 130 may further include other input devices. Specifically, the other input devices 132 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, and the like.
[0075] Among them, the other input device 132 can receive input data and so on.
[0076] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, an interactive interface, file display, and / or the playback of any multimedia file. In the embodiments of the present application, the display unit 140 can be used to display each interactive interface, processing result, etc. in the question-and-answer method.
[0077] The memory 120 can be used to store instructions and data. The memory 120 mainly includes a storage instruction area and a storage data area. The storage data area can store various data, such as multimedia files, texts, etc.; the storage instruction area can store software units such as an operating system, applications, instructions required for at least one function, or their subsets and extended sets. It can also include a non-volatile random access memory; it provides the processor 170 with management of the hardware, software, and data resources in the computing processing device, supports control software and applications. It is also used for the storage of multimedia files and the storage of running programs and applications.
[0078] The processor 170 is the control center of the terminal 100. It connects various parts of the entire terminal 100 through various interfaces and lines. By running or executing the instructions stored in the memory 120 and calling the data stored in the memory 120, it executes various functions of the terminal 100 and processes data, thereby performing overall control of the terminal device. Optionally, the processor 170 can include one or more processing units; preferably, the processor 170 can integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 170. In some embodiments, the processor and the memory can be implemented on a single chip. In some embodiments, they can also be separately implemented on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to corresponding components of the computing processing device, read and process data in the software, especially read and process the data and programs in the memory 120, so that each functional module therein executes corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.
[0079] Among them, the memory 120 can be used to store software codes related to the question-and-answer method. The processor 170 can execute the steps of the question-and-answer method and can also schedule other units (such as the above input unit 130 and display unit 140) to implement corresponding functions.
[0080] The radio frequency unit 110 (optional) can be used for receiving and transmitting information or signals during a call. For example, after receiving the downlink information from the base station, it is sent to the processor 170 for processing. Additionally, it sends the uplink data designed to the base station. Generally, the RF circuit includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0081] Among them, in the embodiment of this application, the radio frequency unit 110 can send data to the server 200 and receive the processing result sent by the server 200. Exemplarily, the radio frequency unit 110 sends the received user query to the server 200, and the server 200 obtains the answer matching the user query and returns the answer matching the user query to the terminal 100.
[0082] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.
[0083] The terminal 100 also includes a power supply 190 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 170 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0084] The terminal 100 also includes an external interface 180. This external interface can be a standard Micro USB interface or a multi-pin connector, which can be used to connect the terminal 100 to other devices for communication and can also be used to connect a charger to charge the terminal 100.
[0085] Although not shown, the terminal 100 may further include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here. Some or all of the methods described below can be applied to the terminal 100 as Figure 2 shown.
[0086] Next, the product form of the server 200 will be described; Figure 1 in
[0087] Figure 3 A schematic structural diagram of the server 200 is provided. As Figure 3 shown, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other through the bus 201.
[0088] The bus 201 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it in
[0089] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0090] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may further include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0091] Among them, the memory 204 can be used to store software codes related to the question-and-answer method, and the processor 202 can execute the steps of the question-and-answer method or schedule other units to implement corresponding functions.
[0092] It should be understood that the above terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware system without the function of executing instructions and the hardware system with the function of executing instructions.
[0093] An embodiment of the present application provides a question-and-answer method. Taking the application of this method to a computer device as an example, the computer device can specifically be Figure 1 the terminal 100 in or a system composed of the terminal 100 and the server 200. Referring to Figure 4 , the question-and-answer method specifically includes the following steps:
[0094] Step S100, obtain a user query.
[0095] Specifically, the system receives the query proposed by the user.
[0096] Optionally, this step can consider the diversity of user input, including spelling mistakes, ambiguous expressions, etc., and perform necessary preprocessing, such as spelling check, synonym replacement, etc., and can also record the user's historical queries for personalized recommendation and context understanding.
[0097] The user query can be in various forms, such as a query input by voice or a query in text form, etc. For user queries in other modalities, they can be uniformly converted into text modality for subsequent processing.
[0098] Step S110, retrieve a target knowledge fragment related to the user query by using a preset structured knowledge base, where the structured knowledge base contains knowledge fragments in a specific domain.
[0099] Specifically, the structured knowledge base contains knowledge fragments in a specific domain, and the structured knowledge base is used to retrieve knowledge fragments relevant to the user query. Optionally, the knowledge fragments can be obtained by parsing knowledge documents in a specific domain. Additionally, the structured knowledge base can be set to update dynamically, adding or updating new knowledge documents or knowledge fragments in a timely manner, and deleting some outdated knowledge documents or knowledge fragments. Optionally, incremental learning techniques can be considered. The design of the structured knowledge base can also consider supporting multimodal data and can include version control and metadata management (such as source, author, creation time, etc.) to enhance the credibility and traceability of knowledge documents.
[0100] The specific domain refers to the domain in which the question-answering method is applied or practiced within a certain specific industry, discipline, or topic range, which can be the communication field, financial field, education field, medical field, etc.
[0101] Step S120: Integrate the target knowledge fragment with the user query to form prompt input data.
[0102] Specifically, the user query is integrated with the knowledge fragments retrieved from the structured knowledge base that are relevant to the user query. Optionally, the process of forming the prompt input data can consider the organization method, quantity, and relevance degree of the knowledge fragments to the user query. The integration process can try different input templates and strategies, such as chain-of-thought prompting, few-shot prompting, etc., so as to guide the model to generate more accurate and comprehensive answers.
[0103] Step S130: Input the prompt input data into a pre-trained large model to obtain an answer that matches the user query output by the large model.
[0104] Among them, the pre-trained large model is obtained through two-stage fine-tuning training: In the first stage, the initial large model is fine-tuned using training data in a specific domain to obtain a large model after the first-stage training. The training data includes user query samples and matching answer labels. In the second stage, on the basis of the training data, knowledge fragments retrieved from the structured knowledge base that match the user query samples are added as new training data to fine-tune the large model after the first-stage training to obtain a large model after the second-stage training.
[0105] Specifically, the initial large model used in the first-stage fine-tuning can be a general large model, that is, a pre-trained large model. These general models are usually trained on a vast amount of text data and have certain language understanding and generation capabilities, such as the GPT series of large models. The initial large model is fine-tuned using domain-specific training data. The first-stage fine-tuning can enable the initial large model to better understand domain-specific terms, concepts, languages, knowledge, or relationships, and can also improve the initial large model's instruction compliance. The domain-specific training data includes user query samples and matching answer labels. These data can be carefully collected and annotated to ensure quality and coverage. Additionally, in order to minimize the difference between the answers predicted by the large model and the true answers, methods such as the cross-entropy loss function can be used to debug the large model.
[0106] In the second stage, the way to add the knowledge fragments matching the user query to the training data of the first stage can be to add the knowledge fragments as context information to the training data, or to use them as additional input features, so that the large model can learn how to utilize the injected knowledge fragments after the second-stage fine-tuning training.
[0107] In addition, the structured knowledge base and user queries may be constantly changing, and the large model can have the ability to continuously learn and update knowledge to adapt to new knowledge and queries.
[0108] The Q&A method provided in this embodiment enables the large model to not only improve its understanding ability and instruction compliance in a specific domain through two-stage fine-tuning training, but also enables the large model to learn how to utilize knowledge fragments, and can achieve a deep integration of the large model fine-tuning technology and the RAG technology, and can provide more accurate and comprehensive answers for user queries.
[0109] In some embodiments of the present application, the construction process of the structured knowledge base in the foregoing embodiment is introduced, and this process may include the following steps:
[0110] Step S200: Parse the text content in the original knowledge document in a specific domain into a structured document tree. The leaf nodes in the structured document tree represent knowledge fragments, and the nodes on the path from the root node to the parent node of each leaf node represent the structural information of the knowledge fragment corresponding to the leaf node in the original knowledge document.
[0111] Specifically, the original knowledge documents in a specific domain can be, for example, technical documents, research papers, book chapters, etc. These original knowledge documents can also be in various formats, such as plain text, HTML, PDF, etc. Optionally, a suitable parser can be used to parse the original document into a structured document tree. The parser can be selected according to the format and complexity of the document. For example, for an HTML document, an HTML parser can be used; for a PDF document, a PDF parser and OCR technology can be used; for a plain text document, specific parsing rules can be designed according to the structural characteristics of the original knowledge document (such as titles, paragraphs, lists, etc.).
[0112] The parent nodes from the root node to each leaf node in the structured document tree constitute the structural information of the leaf node. This path reflects the hierarchical structure of the knowledge fragment in the original document. For example, the structural information of a knowledge fragment may be "book / chapter 1 / section 1 / paragraph 1".
[0113] Step S210: Combine the knowledge fragment corresponding to each leaf node and its structural information, convert the combined text into a vector, and store the knowledge fragment corresponding to each leaf node and the converted vector in the structured knowledge base correspondingly.
[0114] Specifically, the knowledge fragment corresponding to each leaf node and the structural information corresponding to the knowledge fragment can be sequentially concatenated to obtain the combined text. A suitable text embedding model can be used to convert the combined text into a vector representation. Commonly used text embedding models include Word2Vec, GloVe, FastText, BERT, Sentence-BERT, etc. The selection of the model can depend on the specific application scenario and performance requirements. The design of the structured knowledge base can consider the storage efficiency, retrieval efficiency, and subsequent application scenarios of the data. Various database technologies can be used, such as relational databases, NoSQL databases, graph databases, etc.
[0115] Through the above steps of constructing the structured knowledge base, the text content in the original knowledge document can be structured and efficiently stored, improving the retrieval efficiency and accuracy of subsequent knowledge fragments.
[0116] In some possible implementations, the structural information corresponding to the knowledge fragment can also be stored in the structured knowledge base correspondingly.
[0117] Next, Figure 5 An example of a schematic diagram of a structured document tree is shown.
[0118] Decompose the document into smaller, semantically related knowledge fragments, and use the title structure of the original knowledge document to establish a hierarchical relationship to facilitate subsequent retrieval and utilization of knowledge fragments. Regular expressions, natural language processing tools, or specialized document parsing libraries can be used to implement the segmentation of the original knowledge document, and techniques such as XPath or CSS selectors can also be used to extract the titles and their hierarchical relationships.
[0119] Each leaf node is a knowledge fragment, and the corresponding path (the path from the root node to this leaf node) constitutes the structural information of a knowledge fragment. The finally formed document tree can use the document itself as the root node, information such as titles as intermediate nodes, and knowledge fragments as leaf nodes.
[0120] The above embodiments introduce the processing of the text content in the original knowledge document when constructing a structured knowledge base. In some cases, the original knowledge document may also contain other forms of content in addition to the text content, such as tables, images, etc. On this basis, taking the table as an example, this embodiment introduces how to process the tables in the original knowledge document when constructing a structured knowledge base, which can specifically include the following steps:
[0121] Step S300: Convert the tables in the original knowledge document into target formats that can be processed by large models.
[0122] Specifically, the target format can be LaTeX format, etc.
[0123] Step S310: Use a pre-trained large model to summarize the tables in the target format to obtain summary information.
[0124] Specifically, the length of the generated summary can be controlled by the model, such as limiting the number of words or sentences in the summary. The summary information can represent the content expressed in the tables in the original knowledge document.
[0125] Step S320: Convert the summary information into vectors, and store the summary information as a knowledge fragment corresponding to the converted vectors in the structured knowledge base.
[0126] Specifically, in this step, an embedding model can also be selected to convert the summary information into vectors, and this embedding model is the same as the embedding model introduced above. In addition, an efficient indexing strategy can be established when storing in the structured knowledge base to facilitate quick retrieval of relevant content.
[0127] Through the above steps of constructing the structured knowledge base, the table content in the original knowledge document is also processed into knowledge fragments and stored in the structured knowledge base. During the subsequent retrieval of knowledge fragments, the knowledge fragments represented by the tables can be retrieved, so that the content represented by the knowledge fragments can be more comprehensive.
[0128] In addition, the images in the original knowledge document can also be processed and stored in a similar manner.
[0129] The following embodiments introduce two specific implementation schemes for retrieving target knowledge fragments related to the user query by using the pre - set structured knowledge base mentioned above:
[0130] The first scheme specifically includes:
[0131] Step S400: Vectorize the user query to obtain a user query vector.
[0132] Specifically, a pre - trained embedding model can be used to vectorize the user query. This embedding model is the same as the one introduced above. In addition, an appropriate vector dimension can be selected.
[0133] Step S410: Based on the vector correlation, retrieve target knowledge fragments related to the user query vector from the structured knowledge base.
[0134] Specifically, calculate the correlation between the user query vector and the vectors stored in the structured knowledge base, and use the knowledge fragments corresponding to the vectors whose correlation exceeds the first threshold as target knowledge fragments. The correlation calculation methods can include methods such as cosine similarity, dot product, or Euclidean distance.
[0135] This scheme quickly compares the correlation between the user query and the knowledge fragments through vector representation, thus achieving fast retrieval. Compared with the traditional keyword - matching - based retrieval method, vector retrieval can better capture semantic information and improve the retrieval efficiency.
[0136] The second scheme specifically includes:
[0137] Step S500: Vectorize the user query to obtain a user query vector.
[0138] Step S510: Based on the vector correlation, retrieve candidate knowledge fragments related to the user query vector from the structured knowledge base.
[0139] Specifically, calculate the correlation between the user query vector and the vectors stored in the structured knowledge base, and use the knowledge fragments corresponding to the vectors whose correlation exceeds the second threshold as candidate knowledge fragments. The number of candidate knowledge fragments can be more than the final target knowledge fragments.
[0140] Step S520: Use a pre-trained relevance determination model to determine the relevance between each candidate knowledge fragment and the user query respectively, and screen the top N candidate knowledge fragments with the highest relevance as the target knowledge fragments. The relevance determination model is trained using knowledge fragment samples and user query samples annotated with relevance scores.
[0141] Specifically, the pre-trained relevance determination model can select various machine learning models as the relevance determination model, such as deep neural networks, support vector machines, gradient boosting trees, etc. Additionally, various features can be extracted to represent the user query and candidate knowledge fragments, such as word embeddings, word frequencies, TF-IDF, syntactic features, etc.
[0142] This solution can rank the knowledge fragments with higher relevance to the user query in the front by adding a step of re-ranking using the relevance determination model on the basis of the first solution, thereby improving the retrieval accuracy. This effectively solves the problem that the knowledge fragments with relatively weak relevance may be retrieved using vector similarity in the first solution.
[0143] The following embodiments introduce another question-answering method. After the process of retrieving the target knowledge fragments related to the user query using the structured knowledge base pre-set in the question-answering method described above, the method further includes the following steps:
[0144] For the target knowledge fragments retrieved and related to the user query vector, use the structural information of the target knowledge fragments in the original knowledge document to reconstruct the target knowledge fragments, and obtain the reconstructed target knowledge fragments.
[0145] On this basis, integrate the reconstructed target knowledge fragments with the user query to form prompt input data, and input the prompt input data into a pre-trained large model to obtain the answer matching the user query output by the large model.
[0146] Specifically, the structural information can record the position and structural information of each knowledge fragment in the original knowledge document, such as paragraph numbers, chapter titles, etc. According to different structural information, different reconstruction strategies can be adopted. Reconstruction refers to modifying and supplementing the knowledge fragments according to the retrieved knowledge fragments and their structural information in the original knowledge document to make them more complete, easier to understand, and better meet the user's query needs. For example, add the context information (such as the preceding and following paragraphs, chapter titles) of the target knowledge fragment to the reconstructed knowledge fragment. Optionally, if the target knowledge fragment comes from a table, the information of the table can be added to the reconstructed knowledge fragment, such as the title, rows, columns, original text, etc. Optionally, relevant text knowledge fragments and table knowledge fragments can be merged together.
[0147] The following embodiments introduce several possible ways of reconstructing using the structural information of the target knowledge fragments in the original knowledge document, which are introduced as follows:
[0148] 1. If the two leaf nodes corresponding to two target knowledge fragments belong to the same parent node and are consecutive in the original knowledge document, then remove the duplicate characters in the two target knowledge fragments and splice them according to the original order of the two target knowledge fragments in the original knowledge document to obtain the reconstructed target knowledge fragment.
[0149] Specifically, the duplicate characters can be sentences or phrases repeated between adjacent paragraphs. To check whether they are consecutive in the original knowledge document, it can be judged by comparing the positions, paragraphs or headings of the knowledge fragments in the original knowledge document.
[0150] The first reconstruction method above can effectively improve the coherence and integrity of the knowledge fragments and at the same time remove redundant information by removing duplicate content and splicing according to the original order.
[0151] 2. If at least two leaf nodes corresponding to the target knowledge fragments share a common parent node, and the knowledge fragments corresponding to more than a preset proportion of the leaf nodes under the parent node are included in each of the recalled target knowledge fragments, then recall all the knowledge fragments corresponding to the leaf nodes under the parent node and merge them into one reconstructed target knowledge fragment.
[0152] Specifically, the preset proportion can be dynamically adjusted according to different user queries and the original knowledge document structure. For example, for a document with a more complex structure and larger amount of information, the preset proportion can be appropriately increased; for a document with a simpler structure and smaller amount of information, the preset proportion can be appropriately decreased. The preset proportion can also be adjusted according to the specific content of the user query. For the recalled knowledge fragments, semantic analysis and other technologies can be used to identify the semantic relationships between the knowledge fragments and use appropriate conjunctions or phrases for merging.
[0153] The second reconstruction method above uses the structural information of the document to recall more relevant knowledge fragments more comprehensively and improve the integrity of the knowledge fragments. By setting a preset proportion, the recall range can be flexibly controlled to avoid recalling too much irrelevant information, so as to more completely display the target knowledge fragments related to the user query.
[0154] 3. For each target knowledge fragment in each of the recalled target knowledge fragments:
[0155] Determine whether the number of knowledge fragments corresponding to all leaf nodes under the parent node to which the leaf node corresponding to the target knowledge fragment belongs is included in each of the recalled target knowledge fragments exceeds a preset ratio. If not, further determine whether the relevance between the target knowledge fragment and the user query is lower than a preset threshold. If so, delete the target knowledge fragment.
[0156] Specifically, if the number of knowledge fragments corresponding to all leaf nodes under the parent node to which the leaf node corresponding to target knowledge fragment A belongs is not more than the preset ratio among the recalled target knowledge fragments, then further determine whether the relevance between target knowledge fragment A and the user query is lower than a third threshold. If so, delete this target knowledge fragment A.
[0157] The above-mentioned third reconstruction method can accurately screen out the knowledge fragments that are irrelevant to the user query and remove the irrelevant knowledge fragments, thereby improving the accuracy and conciseness of the selection of target knowledge fragments.
[0158] By combining the preset ratio and the relevance threshold for double selection and filtering, the selection granularity can be effectively controlled.
[0159] Figure 6 This is the architecture schematic diagram of the Q&A system of the present application.
[0160] As Figure 6 shown, the architecture of the Q&A system of the present application is divided into three modules: user-large model interaction module, structured knowledge base construction module, and large model two-stage fine-tuning module.
[0161] User-large model interaction module:
[0162] This module is the entry for user-system interaction, responsible for receiving user queries, understanding user intentions, and converting them into a format that the large model can understand. It acts as a bridge connecting users with the structured knowledge base and the large model. The core function of this module is to efficiently and accurately convert user queries into structured retrieval requests, and finally integrate the knowledge fragments retrieved from the structured knowledge base with the user queries and send them to the large model for the generation of the final answer.
[0163] First, the user-large model interaction module receives user queries. Considering the diversity of user inputs, this module can perform necessary preprocessing steps, such as vectorization and keyword extraction. To provide more personalized services, the module can also record the user's historical queries and use this information for personalized recommendation and context understanding.
[0164] Then, based on the preprocessed user query, the user-large model interaction module needs to select an appropriate retrieval strategy to retrieve target knowledge fragments from the structured knowledge base. The retrieval strategy can include the following possible methods: ① Based on vector relevance, retrieve target knowledge fragments from the structured knowledge base that match the user query. ② Based on vector relevance, retrieve candidate knowledge fragments from the structured knowledge base that match the user query, use a relevance determination model to determine the relevance of each candidate knowledge fragment to the user query respectively, and filter the top N candidate knowledge fragments with the highest relevance as the target knowledge fragments. After retrieving the target knowledge fragments, it is also possible to retrieve target knowledge fragments from the structured knowledge base that match the user query based on vector relevance, and then reconstruct the target knowledge fragments to obtain the reconstructed target knowledge fragments.
[0165] Finally, the user-large model interaction module needs to integrate the retrieved target knowledge fragments with the user query to form the final prompt input data. Different input templates and strategies can be tried to guide the large model to generate more accurate and comprehensive answers. Finally, the integrated prompt input data is sent into the large model, waiting for the model to generate the final answer and return it to the user.
[0166] Structured knowledge base construction module:
[0167] The structured knowledge base construction module is responsible for constructing a structured knowledge base to provide a high-quality knowledge source for the upstream user-large model interaction module.
[0168] The structured knowledge base construction module can parse the original knowledge document. During the construction process, the text content and tables in the original knowledge document can be processed separately. For the text content, it is parsed into a structured document tree according to the document structure features (such as titles, paragraphs, etc.). The leaf nodes of the structured document tree represent knowledge fragments, and the path from the root node to the leaf node represents the structure information of the knowledge fragment. The knowledge fragment and the corresponding structure information are combined and converted into vectors and stored in the structured knowledge base. For the tables in the original knowledge document, first, the table is converted into a format that can be processed by the large model (such as LaTeX format), the table is summarized by a pre-trained model, the summary information is converted into a vector, and the summary information is used as a knowledge fragment. The knowledge fragment and the vector are stored in the structured knowledge base.
[0169] Large model two-stage fine-tuning module:
[0170] This module is responsible for training and fine-tuning a large model to enable it to effectively utilize the injected knowledge for question answering. The core objective of the module is to incorporate the information from the structured knowledge base into the large model, enabling it to have stronger reasoning and question-answering capabilities in a specific domain. The fine-tuning process typically adopts a two-stage strategy: in the first stage, the initial large model is fine-tuned using domain-specific training data to adapt it to the language and knowledge of the specific domain, where the training data includes user query samples and matching answer labels; in the second stage, based on the first stage, knowledge fragments retrieved from the structured knowledge base that match the user query samples are added as new training data to further fine-tune the model so that it can better understand and utilize the structured knowledge. Finally, the fine-tuned model is used to receive user queries and retrieved knowledge fragments and generate answers that match the user queries.
[0171] The question-answering device provided by the embodiments of the present application will be described below. The question-answering device described below can be correspondingly referred to the question-answering method described above.
[0172] See Figure 7 , Figure 7 which is a schematic structural diagram of a question-answering device disclosed in the embodiments of the present application.
[0173] As Figure 7 shown, the device may include:
[0174] A user query acquisition unit 11, configured to acquire a user query;
[0175] A target knowledge fragment retrieval unit 12, configured to retrieve target knowledge fragments related to the user query by using a preset structured knowledge base, where the structured knowledge base includes knowledge fragments in a specific domain;
[0176] A prompt input data generation unit 13, configured to integrate the target knowledge fragments and the user query to form prompt input data;
[0177] A question-answering unit 14, configured to input the prompt input data into a pre-trained large model to obtain an answer that matches the user query output by the large model;
[0178] Wherein, the pre-trained large model is obtained through two-stage fine-tuning training: in the first stage, the initial large model is fine-tuned using domain-specific training data to obtain a large model after the first-stage training, where the training data includes user query samples and matching answer labels; in the second stage, knowledge fragments retrieved from the structured knowledge base that match the user query samples are added as new training data based on the training data to fine-tune the large model after the first-stage training to obtain a large model after the second-stage training.
[0179] Optionally, the device of the present application may further include:
[0180] A knowledge base text construction unit is configured to parse the text content in the original knowledge document in a specific field into a structured document tree. The leaf nodes in the structured document tree represent knowledge fragments, and the nodes on the path from the root node to the parent node of each leaf node represent the structural information of the knowledge fragment corresponding to the leaf node in the original knowledge document. Combine the knowledge fragment corresponding to each leaf node and its structural information, convert the combined text into a vector, and store the knowledge fragment corresponding to each leaf node and the converted vector in the structured knowledge base correspondingly.
[0181] Optionally, the device of the present application may further include:
[0182] A knowledge base table construction unit is configured to convert the tables in the original knowledge document into target formats that can be processed by a large model respectively;
[0183] Use a pre-trained large model to summarize the tables in the target format to obtain summary information;
[0184] Convert the summary information into a vector, and store the summary information as a knowledge fragment and the converted vector in the structured knowledge base correspondingly.
[0185] Optionally, the device of the present application may further include:
[0186] A knowledge base structure construction unit is configured to store the structural information corresponding to the knowledge fragment in the structured knowledge base correspondingly.
[0187] Optionally, the structured knowledge base includes: the corresponding relationship between knowledge fragments and vectors; correspondingly, the process of the target knowledge fragment retrieval unit using the preset structured knowledge base to retrieve the target knowledge fragment related to the user query includes:
[0188] Vectorize the user query to obtain a user query vector;
[0189] Based on the vector relevance, retrieve the target knowledge fragment related to the user query vector from the structured knowledge base.
[0190] Optionally, the structured knowledge base includes: the corresponding relationship between knowledge fragments and vectors; correspondingly, the process of the target knowledge fragment retrieval unit using the preset structured knowledge base to retrieve the target knowledge fragment related to the user query includes:
[0191] Vectorize the user query to obtain a user query vector;
[0192] Retrieve candidate knowledge fragments related to the user query vector from the structured knowledge base based on vector relevance;
[0193] Use a pre-trained relevance determination model to determine the relevance between each candidate knowledge fragment and the user query respectively, and screen the top N candidate knowledge fragments with the highest relevance as the target knowledge fragments, where the relevance determination model is trained using knowledge fragment samples and user query samples marked with relevance scores.
[0194] Optionally, the structured knowledge base includes: the structural information of the knowledge fragment in the original knowledge document; then, the device of the present application may further include:
[0195] A knowledge fragment reconstruction unit, configured to, after the target knowledge fragment retrieval unit retrieves the target knowledge fragment related to the user query using the preset structured knowledge base, for the retrieved target knowledge fragment related to the user query vector, use the structural information of the target knowledge fragment in the original knowledge document to reconstruct the target knowledge fragment to obtain a reconstructed target knowledge fragment.
[0196] Optionally, the structural information includes: the path information from the leaf node corresponding to each knowledge fragment to the root node in the structured document tree parsed from the text content in the original knowledge document. On this basis, the process of the knowledge fragment reconstruction unit reconstructing the target knowledge fragment using the structural information of the target knowledge fragment in the original knowledge document includes:
[0197] If there are two leaf nodes corresponding to two target knowledge fragments that belong to the same parent node and are consecutive in the original knowledge document, then remove the duplicate characters in the two target knowledge fragments and splice them in the original order of the two target knowledge fragments in the original knowledge document to obtain a reconstructed target knowledge fragment.
[0198] Optionally, the structural information includes: the path information from the leaf node corresponding to each knowledge fragment to the root node in the structured document tree parsed from the text content in the original knowledge document. On this basis, the process of the knowledge fragment reconstruction unit reconstructing the target knowledge fragment using the structural information of the target knowledge fragment in the original knowledge document includes:
[0199] If there are at least two leaf nodes corresponding to the target knowledge fragments that share a common parent node, and the knowledge fragments corresponding to more than a preset proportion of the leaf nodes under the parent node are included in the recalled target knowledge fragments, then recall all the knowledge fragments corresponding to the leaf nodes under the parent node and merge them into one reconstructed target knowledge fragment.
[0200] Optionally, the structure information includes: in the structured document tree parsed from the text content of the original knowledge document, the path information from the leaf node corresponding to each knowledge fragment to the root node. On this basis, the process of reconstructing the target knowledge fragment by the knowledge fragment reconstruction unit using the structure information of the target knowledge fragment in the original knowledge document includes:
[0201] For each of the recalled target knowledge fragments:
[0202] Judge whether the number of knowledge fragments corresponding to all the leaf nodes under the parent node to which the leaf node corresponding to the target knowledge fragment belongs is more than a preset ratio among the recalled target knowledge fragments. If not, further judge whether the relevance between the target knowledge fragment and the user query is lower than a preset threshold. If so, delete the target knowledge fragment.
[0203] Optionally, the specific field is the communication field.
[0204] An electronic device is further provided in an embodiment of the present application. Refer to Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, tablet computers, teaching large screens, wearable devices, and the like. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present application.
[0205] As Figure 8 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603 to implement the question-answering method in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0206] Typically, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0207] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device implements any one of the Q&A methods provided by the embodiments of the present application.
[0208] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any one of the Q&A methods provided by the embodiments of the present application.
[0209] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0210] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0211] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0212] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)), etc.
[0213] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0214] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0215] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A question-answering method, characterized in that: include: Get user query; Retrieving target knowledge fragments related to the user query using a preset structured knowledge base, wherein the structured knowledge base contains knowledge fragments in a specific field, and the specific field is the field to which the question-answering method is applied; Integrating the target knowledge fragment with the user query to form prompt input data; Inputting the prompt input data into a pre-trained large model to obtain an answer output by the large model that matches the user query; Among them, the pre-trained big model is obtained through two-stage fine-tuning training: in the first stage, the initial big model is fine-tuned using training data in a specific field to obtain a big model after one-stage training, the training data includes user query samples and matching answer labels, and the initial big model is a general big model; in the second stage, on the basis of the training data, knowledge fragments matching the user query samples retrieved using the structured knowledge base are added as new training data to fine-tune the big model after the one-stage training to obtain a big model after the two-stage training.
2. The method according to claim 1, characterized in that The construction process of the structured knowledge base includes: Parsing the text content in the original knowledge document of the specific field into a structured document tree, wherein the leaf nodes in the structured document tree represent knowledge fragments, and each node on the path from the root node to the parent node of each leaf node represents the structural information of the knowledge fragment corresponding to the leaf node in the original knowledge document; The knowledge fragments corresponding to each leaf node and their structural information are combined, and the combined text is converted into a vector, and the knowledge fragments corresponding to each leaf node and the converted vectors are correspondingly stored in the structured knowledge base.
3. The method according to claim 2, characterized in that The process of constructing the structured knowledge base also includes: Converting the tables in the original knowledge document into target formats that can be processed by the large model; Use the pre-trained large model to summarize the table in the target format to obtain summary information; The summary information is converted into a vector, and the summary information is stored in the structured knowledge base as a knowledge fragment corresponding to the converted vector.
4. The method according to claim 2, characterized in that: Also includes: The structural information corresponding to the knowledge fragments is correspondingly stored in the structured knowledge base.
5. The method according to claim 1, characterized in that The structured knowledge base includes: a correspondence between knowledge fragments and vectors; Accordingly, the method of using a preset structured knowledge base to retrieve a target knowledge fragment related to the user query includes: Vectorizing the user query to obtain a user query vector; Based on the vector relevance, a target knowledge segment related to the user query vector is retrieved from the structured knowledge base.
6. The method according to claim 1, characterized in that The structured knowledge base includes: a correspondence between knowledge fragments and vectors; Accordingly, the method of using a preset structured knowledge base to retrieve a target knowledge fragment related to the user query includes: Vectorizing the user query to obtain a user query vector; Based on the vector relevance, retrieving candidate knowledge fragments related to the user query vector from the structured knowledge base; Using the pre-trained relevance determination model, the relevance of each candidate knowledge fragment to the user query is determined respectively, and the top N candidate knowledge fragments with the highest relevance are selected as target knowledge fragments, wherein the relevance determination model is trained using knowledge fragment samples and user query samples labeled with relevance scores.
7. The method according to claim 1, 5 or 6, characterized in that: The structured knowledge base includes: structural information of knowledge fragments in original knowledge documents; Then, after using the preset structured knowledge base to retrieve the target knowledge fragment related to the user query, the method further includes: For the retrieved target knowledge fragments related to the user query, the target knowledge fragments are reconstructed using the structural information of the target knowledge fragments in the original knowledge document to obtain the reconstructed target knowledge fragments.
8. The method according to claim 7, characterized in that The structural information includes: path information from a leaf node to a root node corresponding to each knowledge fragment in a structured document tree parsed from the text content in the original knowledge document; Then, the process of reconstructing the target knowledge fragment by using the structural information of the target knowledge fragment in the original knowledge document includes: If there are two leaf nodes corresponding to two target knowledge fragments that belong to the same parent node and are continuous in the original knowledge document, then remove the repeated characters in the two target knowledge fragments, and splice the two target knowledge fragments according to the original text order in the original knowledge document to obtain the reconstructed target knowledge fragment.
9. The method according to claim 7, characterized in that: The structural information includes: path information from a leaf node to a root node corresponding to each knowledge fragment in a structured document tree parsed from the text content in the original knowledge document; Then, the process of reconstructing the target knowledge fragment by using the structural information of the target knowledge fragment in the original knowledge document includes: If there are at least two leaf nodes corresponding to the target knowledge fragments that share a common parent node, and the knowledge fragments corresponding to the leaf nodes under the parent node exceeding a preset proportion are included in the recalled target knowledge fragments, then all the knowledge fragments corresponding to the leaf nodes under the parent node will be recalled and merged into a reconstructed target knowledge fragment.
10. The method according to claim 7, characterized in that The structural information includes: path information from a leaf node to a root node corresponding to each knowledge fragment in a structured document tree parsed from the text content in the original knowledge document; The process of reconstructing the target knowledge fragment by using the structural information of the target knowledge fragment in the original knowledge document includes: For each target knowledge fragment among the recalled target knowledge fragments: Determine whether the number of knowledge fragments corresponding to all leaf nodes under the parent node to which the leaf node corresponding to the target knowledge fragment belongs contained in the recalled target knowledge fragments exceeds a preset ratio; if not, further determine whether the relevance of the target knowledge fragment to the user query is lower than a preset threshold; if so, delete the target knowledge fragment.
11. The method according to claim 1, characterized in that: The specific field is the communication field.
12. A question-answering device, characterized in that: include: A user query acquisition unit, used to acquire user queries; A target knowledge fragment retrieval unit, used to retrieve a target knowledge fragment related to the user query using a preset structured knowledge base, wherein the structured knowledge base contains knowledge fragments in a specific field, and the specific field is the field to which the question-answering method is applied; a prompt input data generating unit, used for integrating the target knowledge fragment and the user query to form prompt input data; A question-answering unit, used for inputting the prompt input data into a pre-trained large model to obtain an answer output by the large model that matches the user query; Among them, the pre-trained big model is obtained through two-stage fine-tuning training: in the first stage, the initial big model is fine-tuned using training data in a specific field to obtain a big model after one-stage training, the training data includes user query samples and matching answer labels, and the initial big model is a general big model; in the second stage, on the basis of the training data, knowledge fragments matching the user query samples retrieved using the structured knowledge base are added as new training data to fine-tune the big model after the one-stage training to obtain a big model after the two-stage training.
13. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the question-answering method according to any one of claims 1 to 11.
14. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the question-answering method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the question-answering method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Multi-field and multi-disciplinary science and technology policy resource retrieval method and device
CN115344668A
Knowledge base question and answer method and device, equipment and medium
CN117708300A
Academic conference question-answering system based on large language model
CN118377867A