Question answering method and device based on retrieval enhancement generation and storage medium
Through semantic slicing and vector knowledge base processing based on wiseflow, the problems of difficult and poor quality of information crawling are solved, and efficient and accurate question-and-answer generation is achieved.
Patent Information
- Application Number
- CN202510220901.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, the difficulty of information capture, poor quality, poor program generalization ability and unreasonable text processing lead to poor Q&A results.
By obtaining user concern information based on wiseflow, semantic slicing processing is performed, embedded vectors are generated and stored in the vector knowledge base, and using preset generation models to generate question-and-answer results.
It reduces the difficulty of information crawling, improves the quality of information crawling and the generalization ability of procedures, and improves the effectiveness of question and answers.
Smart Images

Figure CN120386876A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technologies, and in particular, to a question-answering method, device, and storage medium based on retrieval-augmented generation. Background Art
[0002] With the rapid increase in the amount of information on the Internet, the acquisition and processing of web page data have become an important task. However, the overloading of web page information makes it difficult for users to quickly filter out the required data. Especially when it is necessary to obtain relevant data from multiple web pages, how to efficiently crawl, parse, and store data has become an urgent problem to be solved. For example, in the cold site project, it is necessary to crawl and update the weather data of a specific location, in the intelligent pricing project, it is necessary to dynamically obtain competitor price information, and the consultation section of the Longhu official website involves obtaining data such as holiday arrangements from government websites.
[0003] The prior art usually uses traditional crawler tools to crawl web page information. However, this method has multiple problems. In the prior art, crawler tools often require manual specification of clear XPath and other information to extract web page data, which has a high threshold for ordinary users and lacks generality. Due to the structural differences of different websites and after website upgrades, crawler tools may not be able to work properly continuously, resulting in the need to constantly manually adjust and update the program. In addition, in the processing link after information extraction, traditional text processing methods often rely on chunking methods based on token length or special characters, ignoring the semantic information of the actual content. This processing method leads to unreasonable text segmentation, thereby affecting the accuracy and integrity of subsequent retrieval, and ultimately resulting in poor question-answering effects. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a question-answering method, device, and storage medium based on retrieval-augmented generation to solve the problems of difficult information crawling, poor information quality, poor program generalization ability, and unreasonable text processing existing in the prior art.
[0005] In the first aspect of the embodiments of this application, a question-answering method based on retrieval-augmented generation is provided, including: collecting document information according to user concerns, and performing semantic slicing processing on the document information to obtain a plurality of semantic chunks, where each semantic chunk contains text content related to user concerns; performing embedding processing on the semantic chunks to generate corresponding embedding vectors, and storing the embedding vectors in a vector knowledge base; in response to a question queried by a user, converting the question into a query vector, and retrieving semantic chunks related to the question from the vector knowledge base according to the query vector; using a preset generation model to process the retrieved semantic chunks related to the question to generate a question-answering result corresponding to the question, and returning the question-answering result to the user.
[0006] In the second aspect of the embodiments of the present application, a question-and-answer device based on retrieval-augmented generation is provided, including: a semantic slicing module, configured to collect document information according to user concerns, and perform semantic slicing processing on the document information to obtain a plurality of semantic blocks, where each semantic block contains text content related to user concerns; an embedding processing module, configured to perform embedding processing on the semantic blocks to generate corresponding embedding vectors, and store the embedding vectors in a vector knowledge base; a query retrieval module, configured to respond to a question queried by a user, convert the question into a query vector, and retrieve semantic blocks related to the question from the vector knowledge base according to the query vector; a question-and-answer generation module, configured to use a preset generation model to process the retrieved semantic blocks related to the question, generate a question-and-answer result corresponding to the question, and return the question-and-answer result to the user.
[0007] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0008] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0009] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects:
[0010] By collecting document information according to user concerns, performing semantic slicing processing on the document information to obtain a plurality of semantic blocks, where each semantic block contains text content related to user concerns; performing embedding processing on the semantic blocks to generate corresponding embedding vectors, and storing the embedding vectors in a vector knowledge base; responding to a question queried by a user, converting the question into a query vector, and retrieving semantic blocks related to the question from the vector knowledge base according to the query vector; using a preset generation model to process the retrieved semantic blocks related to the question, generating a question-and-answer result corresponding to the question, and returning the question-and-answer result to the user. The present application can reduce the difficulty of information scraping, improve the quality of information scraping, improve the generalization ability of the scraping program, and enhance the question-and-answer effect. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1It is a schematic diagram of the overall implementation process of the RAG retrieval question-answering method for obtaining user focus information based on wiseflow provided by an embodiment of the present application;
[0013] Figure 2 It is a schematic diagram of the process of the question-answering method based on retrieval-augmented generation provided by an embodiment of the present application;
[0014] Figure 3 It is a schematic diagram of the structure of the question-answering device based on retrieval-augmented generation provided by an embodiment of the present application;
[0015] Figure 4 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0016] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0017] Modern web pages contain a vast amount of information, and it is becoming increasingly difficult to obtain the required data. Especially in some specific application scenarios, such as weather data scraping in cold site projects, competitor price scraping in intelligent pricing projects, or scraping of government holiday schedules, users often need to dynamically obtain and update information from multiple websites. This information is usually presented in a structured or unstructured form on the web page.
[0018] The most commonly used method in the prior art is to use traditional web scraping tools. Such scraping tools can achieve data collection by scraping web page content and extracting relevant information. Specifically, the crawler usually locates the target data by parsing the HTML structure of the web page. In implementation, it is often necessary to manually provide explicit XPath information for each web page to guide the crawler tool to accurately scrape the required data.
[0019] The specific implementation method of the prior art is as follows:
[0020] First, the crawler tool needs to manually specify the XPath (i.e., the path in the web page DOM tree) to extract specific data. For example, in some web pages, weather data may be located within specific HTML tags, and the crawler needs to accurately locate the tag through the XPath path and then extract the data.
[0021] Then, after the data is captured, it usually needs to be subjected to document slicing and embedding, and the data is further processed using Retrieval Augmented Generation (RAG) retrieval technology.
[0022] Although traditional web crawler tools have been widely used, they also have the following several significant drawbacks, leading to many difficulties in practical applications.
[0023] 1) Difficulty in information capture
[0024] When using traditional web crawler tools, developers need to manually specify the XPath for each page, which is a challenge for ordinary users. Especially when faced with complex web page structures, it is difficult for ordinary users to provide accurate XPath information. In addition, when the website structure changes (for example, when the website is upgraded or redesigned), the original XPath information may become invalid, resulting in the crawler being unable to work properly.
[0025] 2) Poor quality of information capture
[0026] Due to the over-reliance of traditional web crawler tools on manually specified XPath, the quality of the captured data is relatively low. In some cases, changes in the website structure may cause some data to be unable to be captured, or the captured data is incomplete or inaccurate, thus affecting subsequent data analysis and processing.
[0027] 3) Poor generalization ability of the crawling program
[0028] Web crawler tools are usually customized for specific websites, which means that if data from different websites needs to be captured, the parsing rules must be rewritten. For example, if a website upgrades its page structure, the original crawling rules may become invalid, and the web crawler tool needs to be redebugged. This way of relying on manual customization makes the generalization ability of existing web crawler programs relatively poor and unable to easily adapt to new web page structures.
[0029] 4) Unreasonable text processing
[0030] In traditional text processing, the segmentation method based on token length or special characters is usually used, while ignoring the semantic information of the actual content. This approach is likely to lead to unreasonable text segmentation, unable to well retain the logical structure of the data, resulting in inaccurate subsequent retrieval and poor question-and-answer effects. For example, when processing natural language text, data segmented by characters or words often does not consider the semantic relationship between sentences, causing information loss and being unable to accurately match user needs.
[0031] In view of the problems existing in the prior art, this application proposes a RAG retrieval and question-answering method for obtaining user focus information based on wiseflow. It obtains the information data of user focus through wiseflow, parses the web page information, performs semantic slicing, etc., to construct a RAG retrieval and question-answering service based on points of interest, enabling users to quickly obtain valuable information they are interested in. The following will summarize the overall implementation process of the RAG retrieval and question-answering method for obtaining user focus information based on wiseflow in this application with reference to the accompanying drawings. Figure 1 It is a schematic diagram of the overall implementation process of the RAG retrieval and question-answering method for obtaining user focus information based on wiseflow provided by an embodiment of this application. As Figure 1 shown, the RAG retrieval and question-answering method for obtaining user focus information based on wiseflow may specifically include the following parts:
[0032] 1. Document information collection module: Use the wiseflow tool to obtain the information of user focus and incrementally store it in the pocketbase database.
[0033] 2. Document semantic slicing:
[0034] a) Obtain the data of each document through pocketbase.
[0035] b) Extract the meta information data of the document
[0036] c) Perform paragraph semantic splitting on the document content.
[0037] 3. Document content embedding: Embed the content of the semantic slice.
[0038] 4. Incrementally update the vector knowledge base data: Build a program to incrementally obtain data and update the vector knowledge base according to pocketbase.
[0039] 5. Build the RAG process: Use the Langchain framework to build the RAG retrieval process according to the LLM model, query, prompt, and retriever.
[0040] 6. Optimize the RAG question-answering effect: Optimize the prompt according to the question and the effect of the generated answer.
[0041] 7. Package the RAG question-answering service: Use flask and gradio to package the front-end and back-end services. The following will describe the content of the technical solution of this application in detail with reference to the accompanying drawings and specific embodiments.
[0042] Figure 2 It is a schematic diagram of the process of the question-answering method based on retrieval-augmented generation provided by an embodiment of this application.
[0043] As Figure 2 shown, the question-answering method based on retrieval-augmented generation may specifically include:
[0044] S201. Collect document information according to the user's focus, and perform semantic slicing on the document information to obtain multiple semantic chunks, where each semantic chunk contains text content related to the user's focus;
[0045] S202. Embed the semantic chunks to generate corresponding embedding vectors, and store the embedding vectors in the vector knowledge base;
[0046] S203. In response to the question queried by the user, convert the question into a query vector, and retrieve the semantic chunks related to the question from the vector knowledge base according to the query vector;
[0047] S204. Use a preset generation model to process the retrieved semantic chunks related to the question, generate the question-and-answer result corresponding to the question, and return the question-and-answer result to the user.
[0048] In some embodiments, collecting document information according to the user's focus includes:
[0049] Configure a focus list and a list of source website addresses according to the user's focus, and set the frequency and time interval of web page crawling;
[0050] Crawl relevant web pages according to the configuration, and parse the crawled relevant web pages to extract web page information, and incrementally store the web page information in the database.
[0051] Specifically, first, configure a focus list and a list of source website addresses according to the user's task requirements. The focus list includes specific topics or questions that the user is interested in, such as weather data, competitor prices, holiday information, etc.; the list of source website addresses includes the website URLs that provide relevant data, such as weather websites, price query websites, etc. The user can select the fields of interest and corresponding data sources according to different needs.
[0052] In actual operation, the user can define this information through a configuration interface and specify the URLs of relevant websites. For example, if the user is interested in the weather data of a specific city, the user can add the city to the focus list and list the URLs of multiple weather websites in the list of source website addresses.
[0053] Further, after configuring the focus points and source websites, the next step is to set the frequency and time interval of web page scraping. Users can set an appropriate scraping frequency, such as scraping once per hour, per day, or per week, to ensure the timeliness of data. At the same time, the scraping time interval can be set to avoid imposing an unnecessary burden on the website by frequently scraping the same web page.
[0054] For example, a user may wish to scrape weather data at a fixed time every morning and set the scraping time interval to 1 hour to obtain the latest weather changes in a timely manner.
[0055] Further, to achieve efficient web page scraping and parsing, it is first necessary to download and deploy the wiseflow tool. wiseflow is an open-source tool dedicated to web page information scraping and parsing. During deployment, users need to install wiseflow on the server and perform necessary configurations to ensure that it can work in conjunction with parameters such as the configured list of source website addresses and scraping frequencies.
[0056] After deployment, the wiseflow service will start and enter the waiting-to-scrape state, ready to scrape and parse web page data.
[0057] Further, the wiseflow tool starts scraping web page data according to the configured list of source website addresses. Each web page will be parsed according to preset rules to extract key web page information. Common web page information includes but is not limited to:
[0058] Author: The author of the article or data on the web page.
[0059] Title: The title of the web page, usually containing the theme of the article.
[0060] Publication Time: The publication date or update time of the web page content.
[0061] Associated Tag: The classification tag of the web page content, facilitating subsequent content retrieval and classification.
[0062] Web Page URL: The URL address of the web page.
[0063] Web Page Content: The main content of the web page, including text, tables, pictures, etc.
[0064] In specific implementation, wiseflow uses specific XPath expressions to extract this information. For each scraped web page, wiseflow will extract relevant web page information according to the configured XPath rules.
[0065] For example, if the target website is a weather forecast website, wiseflow will scrape information such as the city name, weather forecast, temperature, and humidity on the page.
[0066] Furthermore, the web page information captured and parsed will be incrementally stored in the database. To ensure continuous data updates, wiseflow checks whether the data captured each time is different from the data already stored in the database, and only new or updated data will be stored in the database, thus avoiding duplicate storage.
[0067] Specifically, the stored data will be saved to the article collection in the pocketbase database. Each piece of data represents the content of a captured web page and includes all the information extracted from the web page (such as author, title, publication time, etc.). The pocketbase database generates a unique identifier for each web page record to ensure that the web page content can be tracked and updated during each capture and storage.
[0068] For example, assume that the content of a certain web page has been stored in the article collection, and the content of this web page has been updated during the next capture. Wiseflow will detect the content change and store the updated data in the database with a new timestamp.
[0069] Furthermore, to ensure that the captured data is always up-to-date, wiseflow periodically checks and updates the web page content. After each capture, wiseflow will run and update the data in the database at regular intervals according to the set capture frequency and time interval. Users can access the latest stored data through the database query interface.
[0070] For example, users can query the weather data of a specific city, and the system will return the latest captured data, including information such as temperature and weather forecast, to ensure the timeliness of the data.
[0071] According to the method of the present embodiment described above, by configuring the list of focus points and the list of source website addresses, setting the capture frequency and time interval, and using the wiseflow tool to capture and parse web page information, users can efficiently and automatically obtain and store the required web page data from the specified website. This method can solve the problem of information overload and provide a more accurate and real-time data acquisition and update solution. Through the incremental storage technology, it ensures that the information in the database is always up-to-date, avoids the storage of duplicate data, and improves the data processing efficiency.
[0072] In some embodiments, before performing semantic slicing processing on the document information, the method further includes:
[0073] Obtain the stored web page information from the database, extract the web page information to obtain the metadata information in the web page information, and use the metadata information as the additional information of the document.
[0074] Specifically, first, obtain the stored web page data through the API of the pocketbase database. This data includes the web page content scraped and parsed from the source website, such as information like the author, title, publication time, tags, web page URL, and article content. Each piece of web page data has been stored in the article collection, and each article has a unique identifier. Through the query interface of pocketbase, the system can efficiently retrieve and obtain the web page data that users are interested in.
[0075] Furthermore, next, save each piece of obtained web page data as a Markdown (MD) document in a specified format. Each data item of the web page content (such as the author, publication time, title, tags, web page URL, and article content) needs to be extracted and formatted into a complete document according to the Markdown syntax. The specific operations are as follows:
[0076] Title: Represent the title using the # syntax of Markdown.
[0077] Author, publication time, tags, web page URL: Place them as additional information at the top of the document, formatted as plain text.
[0078] Article content: Serve as the main content of the Markdown document, retaining the original paragraphs and content structure.
[0079] At this time, each scraped web page will be stored as a separate Markdown file, with the file name being the title of the web page. After saving, the Markdown file contains all the basic information and content of the article, facilitating subsequent semantic processing and analysis.
[0080] Furthermore, during the process of saving the web page data as a Markdown document, in addition to saving the main content of the article, it is also necessary to extract and retain the metadata information of the document. Specifically, the metadata of the document includes:
[0081] Author: The author of the article, which can be extracted from the web page through meta tags or specified fields.
[0082] Publish_date: Record the publication time or update date of the web page, which can usually also be extracted through the meta tags or timestamps of the web page.
[0083] Title: The title of the article, usually on the web page <title>in the label or extracted through the page structure. < / title>
[0084] Tags: The classification tags of the article, usually extracted from the tags on the web page or the classification list near the text. <meta> in the tags or the classification list near the text.
[0085] Web page URL (url): The unique URL for each article, facilitating tracking and citation of the original source.
[0086] This metadata information will be stored as additional information at the top part of the MD document, ensuring the integrity of the document structure and making it easy for subsequent processing and querying. This information is very important for subsequent semantic slicing, retrieval, and processing based on the question-answering generation model.
[0087] Furthermore, to ensure the continuous update and tracking of data, any new web page data or updated data will be incrementally stored in the pocketbase database. If the web page content changes, such as the article updates the publication time or tag information, the system will re-scrape and store the latest metadata information.
[0088] Each time new web page information is scraped, by checking the unique identifier of the article, it is ensured that the data in the database remains up-to-date. If there is an update, the wiseflow tool will automatically update the article content and append the new metadata information to the corresponding document.
[0089] In this embodiment, first, the stored web page data is obtained from the pocketbase database, and the metadata information of the document is extracted. Then, by saving the web page content as a Markdown-formatted document, the basic information of the article (such as author, publication time, tags, web page URL) and the main content are preserved. Through this process, the system can efficiently organize and store web page content while ensuring the attachment of metadata information, providing strong support for subsequent semantic slicing, information retrieval, and question-answering generation.
[0090] In some embodiments, semantic slicing processing is performed on the document information to obtain multiple semantic blocks, including:
[0091] When the number of characters in the document information is less than the preset number of characters, the document information is taken as a complete semantic block;
[0092] When the number of characters in the document information is greater than the preset number of characters, the document information is split into multiple paragraphs, and the cosine similarity between adjacent paragraphs is calculated in turn. When the cosine similarity is less than the similarity threshold, the adjacent paragraphs are used as the cut-off points, and according to the determined cut-off points, the document information is divided into multiple semantic blocks;
[0093] When the document information contains a title, the title and the corresponding content in the document information are extracted using regular expressions, and the content corresponding to each title is divided into a semantic block.
[0094] Specifically, in this embodiment, it is determined whether to segment according to the character length. The processing method for short articles is as follows: For short articles with less than 500 characters, it is considered that the article expresses a complete semantic information. Therefore, the whole article is processed as a semantic block without further segmentation. For example, if the number of characters in the input article is 450, at this time, the content of the article is directly used as a semantic block without segmentation.
[0095] The processing method for long articles is as follows: For long articles with more than a preset number of characters (such as 500), it is necessary to determine whether to segment according to the similarity between paragraphs.
[0096] In some examples, first, the long article is split into multiple paragraph blocks according to the line break character. For example, if the article content contains multiple paragraphs and is split into several paragraph blocks through the line break character, at this time, each paragraph is independently used as a block to start calculating the similarity.
[0097] For every two adjacent paragraphs after splitting, the cosine similarity between them is calculated using the sentence embedding method. The cosine similarity is used to measure the similarity between two paragraphs in the semantic space, and the higher the value, the greater the semantic difference between the two paragraphs.
[0098] By calculating the vector representation of each paragraph and applying the cosine similarity formula, the similarity between each pair of adjacent paragraphs is obtained.
[0099] Furthermore, in order to determine when to segment paragraphs, the 95th percentile value of the cosine similarity is used as the threshold. Only when the cosine similarity between paragraphs is lower than this threshold, are they considered to have a significant semantic difference and used as segmentation points.
[0100] For example, in some examples, if the similarity between two paragraphs is 0.85 and the threshold is set to 0.90, it is considered that they express different topics and can be used as segmentation points.
[0101] After calculating the cosine similarity between all adjacent paragraphs, mark all the points where the similarity is lower than the threshold and record the indices of these points. These indices are the positions of the segmentation points, marking the separation positions of the text.
[0102] According to the calculated threshold, mark the segmentation points of each paragraph in the document. Through the segmentation points, the original document is divided into multiple semantic blocks. Each semantic block contains paragraphs with high similarity, while there are significant semantic differences between blocks.
[0103] For example, in some examples, the original article is split into paragraph blocks and the similarity is calculated to obtain multiple segmentation points. Semantic blocks are generated according to the segmentation points. For example, the following semantic blocks can be generated:
[0104] Semantic block 1: includes paragraphs 1 to 3;
[0105] Semantic block 2: includes paragraphs 4 to 6;
[0106] Semantic block 3: includes paragraphs 7 to 9.
[0107] Furthermore, for an article containing subheadings, extract the headings and their corresponding content in the article through regular expressions. Based on the semantic similarity between the headings and the content, divide the content corresponding to each heading into a separate semantic block.
[0108] In some examples, use regular expressions to extract the subheadings in the article. For example, the following method can be used to extract chapter headings: when there is a subheading such as "Section 1: Introduction" in the article, extract "Section 1: Introduction" as the heading.
[0109] Furthermore, regard the content corresponding to each heading as an independent semantic block. If the content under the heading contains multiple paragraphs, further split it according to semantic similarity. For example, split it in the following way: for example, for the heading "Section 1: Introduction", which introduces part of the content with multiple paragraphs, regard this part of the content as a semantic block.
[0110] Furthermore, use the text-embedding-v2 model to embed each generated semantic block, and convert each semantic block into its corresponding vector representation.
[0111] For each generated semantic block, use the text-embedding-v2 model to embed the semantic block to generate a vector representation. Store the embedding vector of each semantic block in the Chroma vector knowledge base for subsequent retrieval and question-answer generation.
[0112] Use the Chroma vector library to store the embedding vectors and label the vectors according to the metadata information of each semantic block. The embedding vectors of each semantic block are stored in the vector knowledge base for subsequent retrieval and use by the question-answer system.
[0113] In this embodiment, by according to the length of the article, the similarity between paragraphs, and the heading information, the document content is effectively segmented into multiple semantic blocks. For long documents, it is judged whether to segment according to the cosine similarity between paragraphs, while articles containing subheadings are divided into different semantic blocks according to the headings. Each generated semantic block is embedded and stored in the vector knowledge base to support subsequent retrieval and question-answering. This method not only improves the accuracy of text segmentation but also ensures the efficiency and accuracy of subsequent processing.
[0114] In some embodiments, retrieving semantic chunks related to a question from a vector knowledge base according to a query vector includes:
[0115] Constructing a retriever using a retrieval framework and setting the recall number of the retriever to determine the number of semantic chunks returned during retrieval; according to the recall number, using the retriever to retrieve semantic chunks related to the user's query question from the vector knowledge base.
[0116] Specifically, when a user submits a question, the system needs to generate a query vector according to the user's query. For this purpose, the system will use a preset language model (such as an LLM model) to transform the user's query question into a high-dimensional vector representation. This query vector represents the semantic features of the user's question and will be used for subsequent matching and retrieval in the vector knowledge base.
[0117] For example, in some examples, for the user's query question "What is the current weather in Beijing", the query question is converted into a query vector using a language model, and this vector represents the semantic content of the question, such as the geographical location (Beijing) and the requirement (weather query).
[0118] Furthermore, use the as_retriever method in the Langchain framework to construct a retriever that can find the most relevant semantic chunks to the query question from the vector knowledge base through the query vector.
[0119] For example, the specific implementation is as follows: First, use the as_retriever method in Langchain, passing in the vector knowledge base and the query vector. By setting the recall number, control the number of semantic chunks returned during retrieval. For example, the user can set to recall 5 relevant semantic chunks.
[0120] In the above example, by setting k = 5, the retriever is restricted to recall the 5 most relevant semantic chunks to the query question.
[0121] Furthermore, once the query vector is generated and the retriever is configured, the system can use the retriever to retrieve the most relevant semantic chunks to the user's query question from the vector knowledge base. The retriever will select the most relevant several semantic chunks according to the similarity between the query vector and the vectors stored in the knowledge base.
[0122] For example, in some examples, for the query vector (such as: "What is the current weather in Beijing"), the retriever finds the 5 most relevant semantic chunks to the query question from the vector knowledge base based on similarity calculation, 5 semantic chunks most relevant to the query question.
[0123] When the user asks "What is the current weather in Beijing", the retriever returns 5 weather information chunks related to "Beijing", for example including:
[0124] Semantic Block 1: Contains weather data in Beijing for the past week.
[0125] Semantic Block 2: Contains real-time information on the current temperature and humidity in Beijing.
[0126] Semantic Block 3: Contains future weather forecasts for Beijing.
[0127] Semantic Block 4: Contains weather data from other locations but related to Beijing.
[0128] Semantic Block 5: Contains a comparison of weather information between Beijing and other major cities.
[0129] Through the above process, the system retrieves the semantic blocks most relevant to the user's query from the vector knowledge base. Next, the system can return these semantic blocks to the user for subsequent question-and-answer generation or other applications. The system will return 5 most relevant semantic blocks and provide them to the downstream question-and-answer generation module to generate the final question-and-answer results.
[0130] According to the method of this embodiment above, first generate a query vector based on the user's query and construct a retriever using the as_retriever method in the Langchain framework. By setting the recall number of the retriever, the system can retrieve the semantic blocks most relevant to the user's query problem from the vector knowledge base. In this way, the system can ensure that the most relevant data is returned, thus providing accurate information support for subsequent question-and-answer generation. This method improves the efficiency and accuracy of retrieval and ensures the efficient operation of the question-and-answer system.
[0131] In some embodiments, use a preset generation model to process the retrieved semantic blocks related to the question to generate the question-and-answer results corresponding to the question, including:
[0132] Determine the generation model used to generate the question-and-answer results and input the semantic blocks related to the question into the generation model;
[0133] Combine the user's question, the input prompt of the generation model, the retriever, and the selected generation model to generate a question-and-answer link;
[0134] Use the question-and-answer link to process the semantic blocks related to the question to generate the final question-and-answer results.
[0135] Specifically, this embodiment describes how to use a preset generation model to process the retrieved semantic blocks related to the question to generate the question-and-answer results corresponding to the question. It describes how to select a suitable language generation model (such as qwen-plus) and construct a complete RAG (Retrieval-Augmented Generation) question-and-answer link through the Langchain framework to generate accurate question-and-answer results.
[0136] First, according to the nature and requirements of the question, select a suitable language generation model (LLM), such as qwen-plus. The qwen-plus model has strong natural language understanding and generation capabilities and can generate high-quality answers.
[0137] For example, when the user queries a question (such as, "What's the weather like in Beijing today"), select an appropriate generation model, such as qwen-plus, as part of the subsequent question-and-answer link.
[0138] Furthermore, based on the user's query question, the retrieved relevant semantic chunks, and the selected generation model, combined with the constructed input prompt, use the Langchain framework to create a complete question-and-answer link. The question-and-answer link will integrate the retriever, the generation model, and the user's question to ensure that each module can cooperate with each other to generate accurate answers.
[0139] According to the user's query question, the input prompt of the generation model needs to incorporate the retrieved semantic chunk information to provide relevant context for the model.
[0140] Use the create_stuff_documents_chain function in Langchain to combine the retrieved semantic chunks with the generation model to generate a complete question-and-answer link.
[0141] Use create_retrieval_chain to combine the retriever and the generation model to ensure effective matching of the retrieved semantic chunks with the user's question.
[0142] Furthermore, after constructing the question-and-answer link, the system will process the semantic chunks related to the question through the question-and-answer link and generate the final question-and-answer result. The question-and-answer link will combine the retrieved semantic chunks, the input prompt, and the generation model to generate an answer according to the user's query question.
[0143] Input the user's question and the relevant semantic chunks into the generation model through the question-and-answer link. The generation model (such as qwen-plus) will generate the final answer according to the input prompt and context.
[0144] Furthermore, according to the initial question-and-answer effect, conduct iterative optimization. In practice, the generated answers may need further optimization. By adjusting the input prompt and the threshold of semantic segmentation, the accuracy and relevance of the question-and-answer results can be improved.
[0145] Adjust the prompt according to the preliminary generation results to ensure that the generation model can understand and accurately answer questions. By adjusting the threshold of semantic segmentation, ensure that the retrieved semantic chunks are more relevant and improve the retrieval accuracy. Finally, the system will generate the Q&A results according to the optimized Q&A link and return them to the user.
[0146] This embodiment describes how to use a preset generation model (such as qwen-plus) to process the retrieved relevant semantic chunks and generate the final answer to the question. By using the Langchain framework to build a complete Retrieval-Augmented Generation (RAG) Q&A link, the system can combine the user's question, input prompt, retriever, and generation model to generate accurate Q&A results. In practical applications, the system can also continuously optimize the prompt and semantic segmentation strategy according to the Q&A effect, thereby improving the accuracy and relevance of the Q&A.
[0147] In some embodiments, the method further includes:
[0148] According to the data update time of the web page information in the database, use a timing scheduling tool to record the latest time after each data update, query the incremental data according to the latest time, and store the incremental data in the vector knowledge base.
[0149] Specifically, to ensure that the data in the vector knowledge base can be continuously updated, it is first necessary to record the latest time after each data update according to the update time of the web page information in the pocketbase database. The data update time in the database can be regularly checked through a timing scheduling tool (such as cron or a custom scheduler).
[0150] In some examples, by querying the web page data stored in pocketbase, the update time field of each record is extracted. After each new data is crawled, the latest data update time is recorded. To ensure the efficiency of incremental updates, the maximum update time needs to be maintained so that only updated data can be queried later.
[0151] Furthermore, according to the recorded latest update time, the timing scheduling tool will use this time to query the incremental data in the database since the last update time. The timing task will be triggered regularly to ensure that each query is based on the latest timestamp to retrieve new or updated data.
[0152] Furthermore, after obtaining the incremental data, the next step is to process and store the incremental data in the vector knowledge base. Use the Chroma vector database to store the incremental data, convert the data into vectors through the embedding representation of the text, and store it in the knowledge base.
[0153] For each piece of incremental data, extract the article content and generate a vector representation through the text-embedding-v2 model. Store the generated embedding vectors in the Chroma vector knowledge base for subsequent retrieval.
[0154] Furthermore, through a timing scheduling tool, the system will regularly check the latest update time in the database and store the incremental data in the vector knowledge base. This incremental update mechanism ensures that the data in the vector knowledge base is always up-to-date, reduces the storage of duplicate data, and improves the retrieval efficiency.
[0155] For example, in some examples, trigger the data update task according to the set time interval (such as every hour, every day). After each update, record the new maximum update time and prepare for the next update.
[0156] In this embodiment, by using the update time of the web page data stored in pocketbase and combining with a timing scheduling tool, query the incremental data regularly and store it in the vector knowledge base. Through this incremental update mechanism, the system can ensure that the data in the vector knowledge base is always up-to-date, while avoiding repeated processing and storage. The processing and storage process of incremental data is efficient and automated, providing timely and relevant semantic information for subsequent retrieval and question answering generation.
[0157] In some embodiments, this embodiment also provides how to encapsulate a question answering service based on the RAG (Retrieval-Augmented Generation) model, encapsulate the backend retrieval service using Flask, and combine with Gradio to build a front-end question answering page to provide an efficient retrieval service based on user focus information.
[0158] First, use Flask to create a simple backend service responsible for receiving the user's query questions, calling the front-end retriever, obtaining the semantic chunks related to the questions, and returning the query results to the front end. If Flask is not installed, install the framework first. Then, create a basic Flask application, encapsulate the retrieval service into an interface for the front-end page to call.
[0159] For example, create a simple application through Flask and expose a / query interface that allows users to submit query questions through a GET request, load data from the already stored Chroma vector knowledge base, use OpenAIEmbeddings as the embedding model, utilize the retriever in the Chroma vector knowledge base to retrieve relevant semantic chunks from the knowledge base according to the user's question, and return the query results.
[0160] Next, a simple front-end interface is provided in combination with Gradio, allowing users to input questions and receive answers through the Flask back-end service.
[0161] For example, the front-end interface sends an HTTP request to the Flask back-end through the query_api function, passing the user's query question. A simple web interface is created through gr.Interface to receive user input and display the returned query results. The query results are obtained from the Flask back-end and the relevant semantic chunks are displayed on the front-end interface.
[0162] During actual operation and interaction, first, run the Flask back-end service in the terminal and run the Gradio front-end interface in the terminal. After starting Gradio, a user interface will automatically open in the browser, where users can input questions, such as: "What's the weather like in Beijing today?" The Gradio interface will call the / query interface of the Flask back-end to initiate a query to the vector knowledge base and retrieve relevant semantic chunks. The Flask back-end returns the retrieved relevant results and displays them to the user through the Gradio interface.
[0163] In this embodiment, the back-end retrieval service is encapsulated through Flask, enabling the system to receive the user's query questions and return relevant semantic chunks. A front-end Q&A page is provided in combination with Gradio, allowing users to directly input questions through a simple interface, and the system returns relevant answers through the back-end retrieval service. In this way, the system can efficiently and real-time provide Q&A services based on the RAG model, ensuring the optimization of query efficiency and user experience.
[0164] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the present application.
[0165] Figure 3 It is a schematic structural diagram of a question and answer device based on retrieval-augmented generation provided by an embodiment of the present application.
[0166] As Figure 3 shown, the question and answer device based on retrieval-augmented generation includes:
[0167] A semantic slicing module 301, configured to collect document information according to user concerns and perform semantic slicing processing on the document information to obtain a plurality of semantic chunks, where each semantic chunk contains text content related to user concerns;
[0168] An embedding processing module 302, configured to perform embedding processing on the semantic chunks to generate corresponding embedding vectors and store the embedding vectors in a vector knowledge base;
[0169] The query retrieval module 303 is configured to, in response to a question queried by a user, convert the question into a query vector, and retrieve semantic chunks related to the question from the vector knowledge base according to the query vector;
[0170] The question-answer generation module 304 is configured to process the retrieved semantic chunks related to the question by using a preset generation model, generate a question-answer result corresponding to the question, and return the question-answer result to the user.
[0171] In some embodiments, Figure 3 the semantic slicing module 301 configures a list of focus points and a list of source website addresses according to the user's focus points, and sets the frequency and time interval of web page scraping; scrapes relevant web pages according to the configuration, and parses the scraped relevant web pages to extract web page information, and incrementally stores the web page information in a database.
[0172] In some embodiments, Figure 3 before performing semantic slicing processing on the document information, the semantic slicing module 301 obtains the stored web page information from the database, extracts the web page information to obtain the metadata information in the web page information, and uses the metadata information as additional information of the document.
[0173] In some embodiments, Figure 3 when the number of characters in the document information is less than a preset number of characters, the semantic slicing module 301 regards the document information as a complete semantic chunk; when the number of characters in the document information is greater than the preset number of characters, the document information is split into multiple paragraphs, and the cosine similarity between adjacent paragraphs is calculated in turn. When the cosine similarity is less than the similarity threshold, the adjacent paragraphs are used as split points, and according to the determined split points, the document information is divided into multiple semantic chunks; when the document information contains a title, the title and the corresponding content in the document information are extracted by using a regular expression, and the content corresponding to each title is divided into a semantic chunk.
[0174] In some embodiments, Figure 3 the query retrieval module 303 of uses a retrieval framework to construct a retriever, and sets the recall number of the retriever to determine the number of semantic chunks returned during retrieval; according to the recall number, the retriever is used to retrieve semantic chunks related to the user's query question from the vector knowledge base.
[0175] In some embodiments, Figure 3 the question-answer generation module 304 determines a generation model for generating a question-answer result, and inputs the semantic chunks related to the question into the generation model; combines the user's question, the input prompt of the generation model, the retriever, and the selected generation model to generate a question-answer link; uses the question-answer link to process the semantic chunks related to the question to generate a final question-answer result.
[0176] In some embodiments, Figure 3 The data update module 305 records the latest time after each data update by using a timing scheduling tool according to the data update time of the web page information in the database, queries the incremental data according to the latest time, and stores the incremental data in the vector knowledge base.
[0177] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0178] Figure 4 is a schematic structural diagram of the electronic device 4 provided by the embodiments of the present application. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0179] Exemplarily, the computer program 403 can be divided into one or more modules / units. One or more modules / units are stored in the memory 402 and executed by the processor 401 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 403 in the electronic device 4.
[0180] The electronic device 4 can be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4 do not constitute a limitation to the electronic device 4. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0181] The processor 401 may be a Central Processing Unit (CPU), or may be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0182] The memory 402 may be an internal storage unit of the electronic device 4, for example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk equipped on the electronic device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 402 may also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 may also be used to temporarily store data that has been output or is to be output.
[0183] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0184] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0185] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0186] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0187] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0188] In addition, the functional units in each embodiment of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0189] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0190] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A question-answering method based on retrieval-augmented generation, characterized in that, include: Collect document information based on user focus, and perform semantic slicing on the document information to obtain multiple semantic blocks, where each semantic block contains text content related to the user focus; Embedding the semantic block to generate a corresponding embedding vector, and storing the embedding vector in a vector knowledge base; In response to a question queried by a user, converting the question into a query vector, and retrieving a semantic block related to the question from the vector knowledge base according to the query vector; The retrieved semantic blocks related to the question are processed using a preset generation model to generate a question-and-answer result corresponding to the question, and the question-and-answer result is returned to the user.
2. The method according to claim 1, characterized in that, The collecting of document information according to user focus includes: Configure the focus list and source website address list based on user focus, and set the frequency and time interval of web page crawling; Relevant web pages are captured according to the configuration, and the captured relevant web pages are parsed to extract web page information, and the web page information is incrementally stored in a database.
3. The method according to claim 2, wherein Before performing semantic slicing processing on the document information, the method further includes: The stored web page information is acquired from the database, the web page information is extracted to obtain metadata information in the web page information, and the metadata information is used as additional information of the document.
4. The method according to claim 1, characterized in that The semantic slicing process is performed on the document information to obtain a plurality of semantic blocks, including: When the number of characters in the document information is less than a preset number of characters, the document information is regarded as a complete semantic block; If the number of characters in the document information is greater than a preset number of characters, the document information is split into multiple paragraphs, and the cosine similarity between adjacent paragraphs is calculated in sequence. When the cosine similarity is less than a similarity threshold, the adjacent paragraphs are used as segmentation points, and the document information is divided into multiple semantic blocks according to the determined segmentation points. In the case that the document information contains a title, the title and the corresponding content in the document information are extracted using a regular expression, and the content corresponding to each title is divided into a semantic block.
5. The method according to claim 1, wherein Retrieving a semantic block related to the question from the vector knowledge base according to the query vector includes: A retriever is constructed using a retrieval framework, and a recall quantity of the retriever is set to determine the number of semantic blocks returned during retrieval; based on the recall quantity, the retriever is used to retrieve semantic blocks related to the user query question from the vector knowledge base.
6. The method according to claim 5, wherein The process of processing the retrieved semantic blocks related to the question using a preset generation model to generate a question-answer result corresponding to the question includes: Determining a generative model for generating question-answering results, and inputting semantic blocks related to the question into the generative model; Combining the user's question, the input prompt of the generative model, the retriever, and the selected generative model to generate a question-answer link; The question-answer link is used to process the semantic blocks related to the question to generate a final question-answer result.
7. The method according to claim 1, characterized in that, The method further comprises: According to the data update time of the web page information in the database, use a timing scheduling tool to record the latest time after each data update. Query the incremental data according to the latest time, and store the incremental data in the vector knowledge base.
8. A question-answering device based on retrieval-augmented generation, characterized in that, Including: A semantic slicing module, configured to collect document information according to user concerns, and perform semantic slicing processing on the document information to obtain a plurality of semantic blocks, where each semantic block contains text content related to user concerns; An embedding processing module, configured to perform embedding processing on the semantic blocks to generate corresponding embedding vectors, and store the embedding vectors in the vector knowledge base; A query and retrieval module, configured to, in response to a question queried by a user, convert the question into a query vector, and retrieve semantic blocks related to the question from the vector knowledge base according to the query vector; An answer generation module, configured to use a preset generation model to process the semantic blocks retrieved and related to the question, generate an answer result corresponding to the question, and return the answer result to the user.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Internet information acquisition method and system based on artificial intelligence large model, and medium
CN121030072A