A method of matching a user problem corpus
Patent Information
- Application Number
- CN202311208130.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-18
AI Technical Summary
[0003]一般情况下,知识库的结构为常见问题标题向量与答案一一对应,实际应用中,由于用户的语言习惯不同,不同的人对同一件事情的提问的表达用词也会不同,因此,用户实际提出的问题与知识库的问题标题存在差异,在提示工程中将用户问题和知识库进行关联时,如果直接将用户提问生成向量,去知识库进行检索关联,可能会因问题标题与实际提问的向量存在偏差太大而匹配不到答案;而如果从问题答案中提取关键词后生成向量,又不能适应用户提问的多种场景变化
[0028]根据本发明,可以适应用户提问的多种场景变化,准确地把用户提问和知识库里的知识内容匹配起来,提高用户提问与知识库的关联准确性和全面性。
Smart Images

Figure CN117435717B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method for matching user question corpora. Background Technology
[0002] Current large-scale language models have made significant progress in the field of natural language processing, but they also have some limitations and shortcomings. Generally, these models are trained on massive amounts of data, so their answers are often based on existing knowledge and samples. For unknown areas or knowledge gaps in specialized domains, these models may provide incorrect or no answers. Relying solely on the output of large-scale language models can lead to misleading or inaccurate answers for users. As the number of parameters in large-scale language models continues to increase, their reading comprehension capabilities are also gradually improving, enabling them to answer questions in areas of knowledge gaps with the help of prompting engineering. Knowledge bases can integrate knowledge from vertical domains to form a structured knowledge base, and prompting engineering allows large-scale language models to understand and apply this knowledge to answer user questions.
[0003] In general, the structure of a knowledge base consists of a one-to-one correspondence between common question title vectors and answers. In practical applications, due to different users' language habits, different people may use different words to express the same question. Therefore, there is a difference between the actual question asked by the user and the question title in the knowledge base. When linking user questions and the knowledge base in the prompting project, if the user's question is directly generated into a vector and retrieved from the knowledge base, the answer may not be matched because the deviation between the question title and the actual question vector is too large. On the other hand, if the vector is generated by extracting keywords from the question and answer, it cannot adapt to the various scenarios of user questions.
[0004] In this context, a solution is needed to match user question corpora, accurately matching user questions with knowledge content in the knowledge base, thereby improving the accuracy and comprehensiveness of the association between user questions and the knowledge base. Summary of the Invention
[0005] To achieve the above objectives, this application provides a method for matching user question corpora, comprising the following steps:
[0006] Obtain the text of the user's question, vectorize the question text, and generate a question vector;
[0007] Based on the question vector, perform an approximation search in the title vector library to obtain the N most similar corpus titles and their corresponding content identifiers;
[0008] Based on content identifiers, corpus information is extracted from the corpus to form model corpus information, which includes: document content;
[0009] The model's corpus information is sent to the language model, which then outputs the corresponding response to the question text.
[0010] Furthermore, the corpus information also includes internal association attribute IDs and external association attributes;
[0011] Among them, the internal association attribute ID is the file identifier of the corpus, and the file content corresponding to the internal association attribute ID is obtained based on the internal association attribute ID;
[0012] Furthermore, the model corpus information also includes the file content corresponding to the internal association attribute ID and the external association attribute.
[0013] The corpus includes file identifiers, file content, internal association attribute IDs, and external association attributes;
[0014] The corpus data sources include PDF files;
[0015] The method for importing the content of a PDF file into a corpus includes: obtaining the content of the PDF document as the original content; dividing the original content into chapters to generate multiple independent chapter contents, with the title corresponding to each chapter content being the chapter title, and generating a unique identifier for each chapter content as a content identifier;
[0016] Import the chapter content and content identifiers into the corpus, corresponding to the file content and file identifiers of the corpus.
[0017] When importing chapter content and content identifiers into the corpus, it is determined whether the chapter content contains a reference to a chapter title. If so, the content identifier corresponding to the referenced chapter title is obtained as an internal association attribute identifier, and the internal association attribute identifier is imported into the corpus, corresponding to the internal association attribute ID in the corpus. The internal association attribute ID can be empty or can be multiple content identifiers.
[0018] Furthermore, when importing chapter content and content identifiers into the corpus, it is determined whether there are business-related keywords in the chapter content. If so, the keywords are imported into the corpus, corresponding to the external association attributes in the corpus. The external association attributes can be empty or multiple words.
[0019] The title vector library includes corpus title vectors and file identifiers;
[0020] The data sources for the title vector library are chapter content, chapter titles, and content identifiers;
[0021] Generate title vectors from the data source, and store the title vectors and the corresponding content identifiers of the chapter content into the title vector library, corresponding to the corpus title and file identifier;
[0022] The title vector includes data generated by vectorizing chapter titles.
[0023] Furthermore, title vectors also include data generated by extracting content titles from chapter content and performing title vectorization operations on the content titles.
[0024] The question vectorization and title vectorization operations adopt a unified processing model, including the text2vec-large-chinese model.
[0025] Here, dividing the original content by chapter refers to dividing it according to the document's table of contents, including:
[0026] Use regular expressions to obtain the list of document headings and their hierarchical relationships;
[0027] To obtain the remaining part of a document after deleting the document directory, locate the main headings and subheadings starting from the beginning according to the order of the document directory, and then segment the document according to the heading positions to obtain the segmentation result.
[0028] According to the present invention, it can adapt to various scenarios of user questions, accurately match user questions with knowledge content in the knowledge base, and improve the accuracy and comprehensiveness of the association between user questions and the knowledge base. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating the steps of a method for matching user question corpora according to an embodiment of the present invention;
[0030] Figure 2 This is a schematic diagram illustrating the conversion between the original content constituting a corpus and a title vector library according to an embodiment of the present invention;
[0031] Figure 3 This is a schematic diagram of the original content chapter titles provided according to an embodiment of the present invention;
[0032] Figure 4 This is a flowchart of a method for matching user question corpora according to an embodiment of the present invention. Detailed Implementation
[0033] This invention starts with the source of the knowledge base, constructing a corpus by splitting information from original files such as PDFs. It then seeks relationships between the information extracted from the original files, establishing connections between different chapter contents through internal association attributes. Simultaneously, keywords are extracted from the chapter content as external association attributes. Combining these internal and external association attributes enriches the corpus content for each user question. Based on the corpus, a title vector library is built to improve the coverage of user-asked questions. By combining the title vector library with the corpus, connections are established for information related to user questions, ultimately providing a complete corpus. This corpus is then used for further training of a large-scale language model.
[0034] The specific implementation of the present invention will now be described in detail with reference to the accompanying drawings.
[0035] The method for matching user question corpora provided by this invention is based on a title vector library and a corpus constructed from information extracted from original files such as PDF files. This prepares the data for matching complete and accurate corpora. The data conversion process from original files to the title vector library and corpus is as follows: Figure 2 As shown:
[0036] First, the original file contains content that can be used as reference material to answer user questions, and its format supports PDF, images, Word documents, etc. For example, when operating on a PDF file, the original file can be split into a batch of independent documents. The splitting boundaries depend on project requirements; for example, a document directory can naturally divide the text according to different functional content. Therefore, a method such as... Figure 3 The document directory shown is used to segment the content. The segmentation steps are as follows:
[0037] 1) Use regular expressions to obtain the entire document's table of contents, including the headings and hierarchical structure, such as... Figure 3 In the example, Chapter 1 has 6 sections, Chapter 2 has 4 sections, and so on;
[0038] 2) Obtain the remaining part of the document after deleting the document directory. Starting from the beginning, locate the positions of the main headings and subheadings according to the order of the document directory (string matching). Then, divide the document according to the position of the headings (the index of the first and last characters of the heading in the document). This will give you the result of the division, namely: a batch of independent documents (doc).
[0039] After segmentation, chapter titles and chapter content are extracted from the original files. Then, step S200 is performed, which assigns a content identifier to each chapter title. The content identifier is the chapter content itself, meaning it's a unique identifier for each individual document.
[0040] After this step is completed, the chapter titles, chapter content, and content identifiers corresponding to the original files are temporarily stored in a cache or other medium to facilitate the subsequent import of data into the corpus and title vector library.
[0041] The structure of a corpus includes file identifiers, file content, internal association attribute IDs, and external association attributes;
[0042] As shown in step S210, the chapter content and content identifiers are first imported into the corpus, corresponding to the file content and file identifiers of the corpus; wherein, the internal association attribute ID in the corpus structure can be empty or multiple content identifiers; the external association attribute can also be empty or multiple words.
[0043] When importing chapter content and content identifiers into the corpus, it is determined whether there is a reference to the chapter title in the chapter content. If there is, the content identifier corresponding to the referenced chapter title is obtained as the internal association attribute identifier. The internal association attribute identifier is imported into the corpus and the corresponding internal association attribute ID of the chapter content and content identifier in the corpus is matched.
[0044] For example:
[0045]
[0046]
[0047] The section titled "Mechanism of Action and Animal Experiment Results" and the content identifier "b9c7d8f1-3a4b-4f2a-8bea-3f7f7c2d9e32" corresponds to the section whose internal association attribute ID stores the content identifier "7b3b2b8f-5b7a-4a85-8f3e-5e0c171e1e2a" for the section on "Pharmacology and Toxicology".
[0048] On the other hand, it determines whether there are business-related keywords in the chapter content. If so, the keywords are imported into the corpus, and the corresponding content is identified by the corresponding external association attribute in the corpus.
[0049] In this invention, named entity recognition (NER) is used to extract disease and drug entity words from chapter content.
[0050] For example, the keyword "paclitaxel" appears in the statement, "Based on the mechanism of action and animal studies, use of this product by pregnant women may cause harm to the fetus. In animal reproductive toxicity studies, when pregnant rats were given a dose of this product below the maximum recommended human dose based on body surface area, embryo-fetal toxicity occurred, including intrauterine embryonic death, increased drug absorption, reduced live birth rate, and malformations (see [Pharmacology and Toxicology])." This record has a total of four attributes: document identifier, document content, internal association attribute ID, and external association attribute. See the table below:
[0051]
[0052]
[0053] The title vector library includes corpus title vectors and file identifiers;
[0054] The data for the title vector library comes from temporarily stored chapter content, chapter titles, and content identifiers;
[0055] like Figure 2 As shown in step S220, content titles are extracted from the chapter content, and title vectors are generated by vectorizing the content titles. The title vectors and the content identifiers corresponding to the chapter content are stored in the title vector library, which corresponds to the corpus title and the file identifier. The data stored in the title vector also includes the data generated by vectorizing the chapter titles.
[0056] In the example above, the two sections are titled "Mechanism of Action and Animal Experiment Results" and "Pharmacology and Toxicology," respectively. However, the titles do not clearly indicate which drug's mechanism of action, animal experiment results, or pharmacology and toxicology they pertain to. Therefore, this invention also provides a method for extracting titles from the section content. For example, the statement "Paclitaxel has a protein binding rate of 89%–98%, is mainly metabolized in the liver, enters the intestine with bile, and is excreted in feces (>90%). Renal clearance accounts for only 1%–8% of total clearance, and its metabolism in patients with hepatic or renal insufficiency is unclear" can be refined into the title "Paclitaxel Metabolism in the Liver." Combining this with the section title "Pharmacology and Toxicology" can broaden the scope of issues that this content can address.
[0057] The chapter titles and content titles are vectorized using the text2vec-large-chinese model (1000 dimensions) in this invention.
[0058] This invention provides two methods for extracting titles from chapter content: a neural network model for text generation and a language big data model for reading comprehension. In practical applications, these two methods can be selected according to the environment.
[0059] 1) Neural Network Model for Generating Titles: This case uses the T5 model, with manually labeled titles corresponding to a batch of chapter content. Both the chapter content and the manually labeled titles serve as the training dataset for the T5 model. The chapter content acts as the input prompt for the T5 model, while the manually labeled titles serve as its output. The trained T5 model possesses the ability to deduce the title from the given answer. The prompt template for training and inference is as follows:
[0060] Disease: {disease_name}, Content: {chapter content}, Title:
[0061] Where: {disease_name} and {section content} are the input content, and the output title is under this template;
[0062] 2) The language model-based reading comprehension approach can be used, and the prompt template is as follows:
[0063]
[0064] Both of the above methods can generate content titles, but the difference is that T5 can only generate one title at a time, while the language model can generate multiple titles at once, as shown in the following example:
[0065]
[0066]
[0067] The number of content titles generated for each document can be set according to actual needs, generally from one to five.
[0068] It's important to note that the file content in the corpus doesn't have corresponding titles. Instead, it's linked by file identifiers in the corpus and file identifiers in the title vector library. This maps a file's content in the corpus to multiple title vectors in the title vector library. Therefore, for a user's question, as long as its vectorization approximates multiple title vectors in the title vector library, a matching file content can be obtained. In the example in the table above, the title of this file content is "Mechanism of Action and Animal Experiment Results." If the user's question is not similar to this title, it only needs to approximate one of the extracted titles to match that file content for the user.
[0069] Once the title vector library and corpus are completed, the corpus matching method provided by this invention can be implemented.
[0070] Figure 1 This is a flowchart illustrating the method steps for matching user question corpora provided by the present invention, as shown in the figure, including the following steps:
[0071] S100: Obtain the text of the user's question, vectorize the question text, and generate a question vector;
[0072] The process of vectorizing questions is the same as that of vectorizing titles, using the text2vec-large-chinese model to convert text into question vectors.
[0073] S110: Based on the question vector, perform an approximation search in the title vector library to obtain the content identifiers corresponding to the N most similar corpus title vectors;
[0074] The approximity search can be implemented using the cosine similarity formula. The number of results N returned can be set according to the actual situation, for example, N is usually set to a value within 5.
[0075] S120: Extract corpus information from the corpus based on the content identifier to form model corpus information. The corpus information includes: file content, internal association attribute ID, and external association attribute.
[0076] Since the internal association attribute ID in the corpus is actually the file identifier of the corpus, the corresponding file content can be obtained based on the internal association attribute ID.
[0077] External related attributes are keywords in the document content, i.e., conceptual text content, such as "paclitaxel" in the example above.
[0078] The file content and the file content corresponding to the internal association attribute ID, along with keywords, are combined to generate model corpus information. This information is used as input prompts for the language model, enabling the language model to read and understand the content of the prompts and output answers to the user.
[0079] S130: Send the model corpus information to the language model and output the response answer corresponding to the question text.
[0080] This invention provides a complete solution for matching user question corpora, namely, the structural design and data generation scheme of the title vector library for vector matching, the structural design and data generation scheme of the corpus for finally generating training corpora, and the processing scheme for responding to user questions; the collaborative process of these three is as follows: Figure 4 As shown:
[0081] The data foundation of this invention is a title vector library and a corpus, with a unified data source and support for PDF files, etc.
[0082] The problem this invention aims to solve is: processing user questions and matching user question corpora. The process of processing user questions includes step S410, obtaining question vectors through vectorization of user questions, and step S420, performing document similarity search and obtaining associated corpus data, combining them into corpus data for the next step of obtaining answers from a large language model. In the above process, step S420 interacts with the data foundation involved in this invention: in step S421, a request is initiated to the title vector library to find the most similar question vector; the query result returned from the title vector library is obtained in step S422; the document identifier extraction request is organized in step S423; the corresponding information is queried from the corpus library in step S424; and the queried corpus results are obtained from the corpus library in step S425. The corpus results involved in this invention support chapter content from the original document, other chapter content corresponding to the titles referenced in the chapter content, and conceptual keywords involved in the chapter content. In practical applications, the type and quantity of corpus results can be flexibly selected in the request information to the corpus library in step S423.
[0083] This invention provides a complete solution for obtaining the optimal model corpus by matching user question corpora. From the perspective of building a knowledge base, it integrates knowledge from vertical domains using internal and external relational attributes to form a structured, network-connected knowledge base. From the perspective of matching, it establishes multiple titles with the highest relevance to user questions and achieves question matching through vector approximation. From the perspective of corpus preparation, it provides comprehensive information related to users to the language model. This complete solution can adapt to various scenarios of user questioning, accurately matching user questions with knowledge content in the knowledge base, and improving the accuracy and comprehensiveness of the association between user questions and the knowledge base.
[0084] The above-disclosed embodiments are merely a few specific examples of the present invention. However, the present invention is not limited thereto, and any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A method for matching user question corpora, characterized in that, Includes the following steps: Obtain the text of the user's question, and vectorize the question text to generate a question vector; Based on the question vector, perform an approximation search in the title vector library to obtain the N most similar corpus titles and their corresponding content identifiers; Based on the content identifier, corpus information is extracted from the corpus to form model corpus information, wherein the corpus information includes: file content; The model corpus information is sent to the language model, which then outputs the response answer corresponding to the question text. The corpus includes file identifiers, file content, internal association attribute IDs, and external association attributes; the data source of the corpus includes PDF files; the method for importing the content of the PDF files into the corpus includes: obtaining the content of the PDF document as the original content; dividing the original content into chapters to generate multiple independent chapter contents, with the title corresponding to each chapter content being the chapter title, and generating a unique identifier for each chapter content as a content identifier; importing the chapter content and content identifier into the corpus, corresponding to the file content and file identifier of the corpus; The process of segmenting the original content by chapter refers to segmenting it according to the document directory, including: using regular expressions to obtain the title list and hierarchical relationship of the document directory; obtaining the remaining part of the document after deleting the document directory; locating the positions of the main headings and subheadings from the beginning according to the order of the document directory; and then segmenting the document according to the heading positions to obtain the segmentation result. When importing the chapter content and content identifier into the corpus, it is determined whether there is a reference to the chapter title in the chapter content. If there is, the content identifier corresponding to the referenced chapter title is obtained as the internal association attribute identifier, and the internal association attribute identifier is imported into the corpus, corresponding to the internal association attribute ID in the corpus. When importing the chapter content and content identifier into the corpus, it is determined whether there are business-related keywords in the chapter content. If they exist, the keywords are imported into the corpus, corresponding to the external related attributes in the corpus. The title vector library includes corpus title vectors and file identifiers; the data sources of the title vector library are chapter content, chapter titles, and content identifiers; title vectors are generated from the data sources, and the title vectors and the content identifiers corresponding to the chapter content are stored in the title vector library, corresponding to corpus titles and file identifiers; wherein, the title vectors include data generated by performing title vectorization operations on chapter titles; The title vector also includes: data generated by extracting content titles from chapter content and performing title vectorization operations on the content titles.
2. The method for matching user question corpora according to claim 1, characterized in that, The question vectorization and the title vectorization adopt a unified processing model, including the text2vec-large-chinese model.
3. The method for matching user question corpora according to claim 1, characterized in that, The approximity search uses cosine similarity calculation.
4. The method for matching user question corpora according to claim 1, characterized in that, The business-related keywords are extracted from the chapter content using named entity recognition.
5. The method for matching user question corpora according to claim 1, characterized in that, The content identifier is a unique identifier corresponding to each chapter.
Citation Information
Patent Citations
Document directory automatic generation method and device and computer readable storage medium
CN110852079A
Question and answer method and device, terminal equipment and readable storage medium
CN116483962A
Event query method and device, computer equipment, storage medium and program product
CN116561288A