Large model question and answer method and device based on multi-party private domain long text information

CN120030131AInactive Publication Date: 2025-05-23BEIJING BIG DATA ADVANCED TECH RES INST

Patent Information

Application Number
CN202510503661.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for big models to accurately capture and acquire factual knowledge during the Q&A process, and there are problems with hallucinations and insufficient information value content.

Method used

By obtaining the question text entered by the user, vector search is performed in multiple private domain vector databases and local vector databases, related search information fragments are aggregated and sorted, and recall fragment sequences are generated using template splicing, and the knowledge question and answer model is input to answer.

Benefits of technology

It effectively enhances the ability of large models to obtain factual knowledge related to the question, improves the accuracy and accuracy of knowledge questions and answers, and breaks the limitation of insufficient value content of a single local information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030131A_ABST
    Figure CN120030131A_ABST
Patent Text Reader

Abstract

The invention discloses a large model question and answer method and device based on multi-party private domain long text information, and belongs to the technical field of large model data processing. According to the question text, performing vector retrieval in the plurality of private domain vector databases and a local vector database to obtain a plurality of related retrieval information fragments; wherein vector representations obtained by converting long text information through a vector model are stored in each private domain vector database and the local vector database; the long text information refers to text information with the length exceeding the length of the maximum fragment text; gathering the plurality of related retrieval information fragments, de-duplicating the related retrieval information fragments with the similarity higher than a threshold value, and sorting by using a re-sorting strategy to obtain a recall fragment sequence consisting of a plurality of recall fragments; and splicing the recall fragment sequence and the question text by using the template, and inputting the recall fragment sequence and the question text into the knowledge question-answer large model for answering to obtain an answer text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of large-model data processing, and specifically relates to a large-model question-answering method and device based on multi-party private domain long text information. Background Art

[0002] At present, large models such as ChatGPT and LLama have achieved good results in fields such as knowledge question answering due to their emergence ability and generalization. Large models such as ChatGPT have hundreds of billions or tens of billions of parameters, and their representation ability has also been greatly improved. The initial training cost enables large models to have better question-answering capabilities, and also have better downstream task migration capabilities.

[0003] However, large models still have some limitations for question answering. They are usually unable to accurately capture and obtain factual knowledge. Moreover, the local knowledge content they store is relatively generalized, making it difficult to accurately obtain factual knowledge for a question. Summary of the invention

[0004] The embodiments of the present application provide a large-model question-answering method and device based on multi-party private domain long text information, which is conducive to improving the ability of the large model to acquire factual knowledge related to the question.

[0005] In a first aspect, a large model question answering method based on multi-party private domain long text information is provided, the method comprising: Get the question text entered by the user; According to the question text, vector retrieval is performed in multiple private domain vector databases and local vector databases to obtain multiple relevant retrieval information fragments; wherein each of the private domain vector database and the local vector database stores a vector representation obtained by converting long text information through a vector model; the long text information refers to text information whose length exceeds the maximum fragment text length; Aggregating the multiple related search information fragments, removing duplicates from the related search information fragments whose similarity is higher than a threshold, and sorting them using a reordering strategy to obtain a recall fragment sequence consisting of multiple recall fragments; The recall segment sequence and the question text are spliced ​​together using a template, and are input into the knowledge question-answering model to answer, thereby obtaining an answer text.

[0006] In a second aspect, a large model question-answering device based on multi-party private domain long text information is provided, wherein the device is used to execute the steps in the large model question-answering method described in the first aspect; wherein the device comprises: The question acquisition module is used to obtain the question text input by the user; A retrieval module, configured to perform vector retrieval in a plurality of private domain vector databases and a local vector database according to the question text, and obtain a plurality of relevant retrieval information fragments; wherein each of the private domain vector database and the local vector database stores a vector representation obtained by converting long text information through a vector model; the long text information refers to text information whose length exceeds the maximum fragment text length; A sorting module is used to aggregate the multiple related search information fragments, remove duplicates of the related search information fragments with similarity higher than a threshold, and sort them using a re-sorting strategy to obtain a recall fragment sequence composed of multiple recall fragments; The knowledge question-answering model is used to use a template to splice the recall segment sequence and the question text, and to answer the question to obtain an answer text.

[0007] According to a third aspect, an electronic device is provided, including a processor and a memory, wherein the memory stores programs or instructions that can be executed on the processor, and when the programs or instructions are executed by the processor, the steps of the large model question-answering method based on multi-party private domain long text information as described in the first aspect are implemented.

[0008] In a fourth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the large model question and answer method based on long text information in the private domain of multiple parties as described in the first aspect are implemented.

[0009] In a fifth aspect, a computer program / program product is provided, which is stored in a storage medium and executed by at least one processor to implement the steps of the large model question-answering method based on multi-party private domain long text information as described in the first aspect.

[0010] The beneficial effect of the present application is that the present application embodiment can effectively enhance the ability of large-model question-answering by using multiple local private domain knowledge (i.e., private domain vector databases) and multiple local highly reliable, timely, and professional knowledge information, and help to break the defect of insufficient value content of a single local information, and accurately answer questions in related fields. In addition, the present application embodiment converts long text information into vectors after segmentation and stores them in the vector database, so that subsequent large-model retrieval can obtain highly relevant recall fragments, so as to improve information retrieval capabilities, accurately obtain factual knowledge related to the problem, and thus improve knowledge question-answering capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 It is a flowchart of the steps of a large model question-answering method based on multi-party private domain long text information in an embodiment of the present application; Figure 2 is a flowchart of a knowledge question answering method in an embodiment of the present application; Figure 3 is a schematic diagram of a process of fragment aggregation in an embodiment of the present application; Figure 4 This is a schematic diagram of a storage process of long text information in an embodiment of the present application; Figure 5 is a schematic diagram of a generation process of a recall fragment sequence in an embodiment of the present application; Figure 6 It is a schematic diagram of the execution flow of a direct interception strategy in an embodiment of the present application; Figure 7 It is a schematic diagram of the execution flow of a scoring re-ranking strategy in an embodiment of the present application; Figure 8 It is a schematic diagram of an execution flow of a related judgment reordering strategy in an embodiment of the present application; Fig. 9 It is a schematic diagram of the execution flow of a bubble reordering strategy in an embodiment of the present application; Fig.10 It is a schematic diagram of the execution flow of a keyword reordering strategy in an embodiment of the present application. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of this application.

[0014] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first" and "second" are generally of the same type, and the number of objects is not limited. For example, the first object can be one or at least two. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.

[0015] At present, large models such as ChatGPT and LLama have achieved good results in fields such as knowledge question answering due to their emergence ability and generalization. Therefore, large model question answering enhancement based on local knowledge may be the next research direction for large model related tasks.

[0016] Large models such as ChatGPT have hundreds of billions or tens of billions of parameters, and their representation capabilities have also been greatly improved. And as the number of model parameters increases, the training cost required also increases, including more training resources, more unsupervised training data, and more sophisticated supervised training data. The initial cost enables the model to have better question-answering capabilities, as well as better downstream task migration capabilities. Generally, models with more than one billion parameters are called large models, and their model frameworks have not changed much. Most large models are improved on the basis of Transformer.

[0017] However, big models still have some limitations for question answering. First, since big models are black box models, they are usually unable to accurately capture and obtain factual knowledge. In addition, big models are pre-trained with knowledge from multiple sources, and the knowledge content they store is relatively generalized. For a question, it may not be possible to accurately obtain factual knowledge.

[0018] In addition, large models may have hallucination problems. Large models may generate content that is inconsistent with user input, contradictory to previously generated content, or inconsistent with known world knowledge. The hallucination problem of large models may include the following reasons: insufficient amount of relevant data. When the amount of training data related to the large model field or problem is insufficient, no matter how much computing resources are invested, better results will not be achieved. Insufficient data may cause the model to learn false or wrong knowledge, thus generating hallucinations; insufficient data quality. When the amount of data increases, the quality of the data is difficult to guarantee. The model may learn wrong information and memorize it, and then may generate content that is inconsistent with the facts. At the same time, data inconsistency can also cause hallucinations; the reason for over-generalization of the model. During training, the large model may learn some rules or patterns that are inconsistent with actual knowledge, which is called over-generalization. When faced with new data, the rules or patterns learned during training may be incorrectly applied to the new data, thus generating hallucinations.

[0019] In addition, large models are generally pre-trained with corpus from a period of time, and the knowledge trained may be outdated and incomplete, so new knowledge cannot be used to answer questions. Constantly updating parameters with new knowledge to improve the timeliness of large models will incur high costs. When the knowledge content is constantly changing and may conflict with the original knowledge, there is uncertainty as to whether the new knowledge can be correctly acquired.

[0020] In view of the above problems, this application proposes a large-model question-answering method and device based on multi-party private domain long text information to improve the ability of the large model to acquire factual knowledge related to the question.

[0021] In the first aspect, the embodiment of the present application provides a large model question answering method based on multi-party private domain long text information. Below, the large model question answering method based on multi-party private domain long text information proposed in the first aspect of the present application is specifically described through Sections 1.1-1.7.

[0022] 1.1 Overview of the method proposed in this application: See also Figure 1 As shown, it is a flowchart of a large model question-answering method based on multi-party private domain long text information provided by an embodiment of the present application. The method may include the following steps: Step S101, obtaining the question text input by the user.

[0023] Among them, the question text refers to the question in text form input by the user. When the user asks a relevant question, such as "What important contributions has person A made in philosophy?", the knowledge question and answer model can search based on the information stored locally (for example, massive long texts), obtain factual knowledge related to the question (for example, person A's personal information and main experiences), and answer based on the retrieved factual knowledge to give an accurate answer.

[0024] Step S102, based on the question text, vector search is performed in multiple private domain vector databases and local vector databases to obtain multiple relevant search information fragments; wherein each of the private domain vector database and local vector database stores a vector representation obtained by converting long text information through a vector model; the long text information refers to text information whose length exceeds the maximum fragment text length.

[0025] Among them, the private domain vector database refers to a private database that stores text information through vector encoding. In this embodiment, it can refer to a vector database in any field. In this embodiment, the field to which the private domain vector database belongs is not limited. The local vector database refers to a local database for information storage that is provided by the large model itself. Information is stored in the form of vectors in this database. Unstructured text information is an important storage form of local knowledge. The massive local long texts generally include multiple long text information (i.e., long text information). Each long text exceeds the length limit of the general BERT model (such as the maximum fragment text length) and varies in length, generally including a description in text form. Taking a news article introducing person A as an example, the description may include person A’s personal information, main experiences, etc. Reference Figure 2 , Figure 2 A flow chart of a knowledge question answering method is shown, such as Figure 2 As shown in the figure, after obtaining the question text input by the user, the vector model is used to vector encode the question text, and according to the encoded vector (the vector representation corresponding to the question text), a search is performed in various vector databases to find multiple vectors (i.e., relevant search information fragments) with high similarity to the vector. The multiple relevant search information fragments come from multiple private domain vector databases and local vector databases, thereby realizing the aggregation of fragment information from multiple databases, enriching the information source, and further improving the accuracy of the acquired factual knowledge.

[0026] Step S103 , aggregating the multiple related search information segments, removing duplicates of the related search information segments with similarity higher than a threshold, and sorting them using a reordering strategy to obtain a recall segment sequence consisting of multiple recall segments.

[0027] Specifically, refer to Figure 3 , Figure 3A schematic diagram of a process of fragment aggregation is shown, such as Figure 3 As shown, the fragments (i.e., the relevant retrieval information fragments, such as Figure 3 The fragments 1, 2, ..., and n shown in the figure are aggregated, and the fragments (i.e., relevant search information fragments) with similarity higher than the threshold are deduplicated (i.e., similarity is calculated for any two relevant search information fragments, and if the similarity between the two exceeds the threshold, one of the relevant search information fragments is deleted), and search information from multiple private domains is obtained to support the large model for knowledge question answering. The specific value of the threshold can be set according to actual application requirements and is not limited in this embodiment.

[0028] Step S104, using a template to splice the recall segment sequence and the question text, and inputting the knowledge question and answer model to answer, thereby obtaining an answer text.

[0029] Specifically, the embodiment of the present application uses a large model to answer questions that are spliced ​​with knowledge. Use a template to splice knowledge and questions, and then input them into the large model to get answers. In this embodiment, the parameters of the original question-answering model can be used, and it is hoped that the large model can combine the knowledge of the knowledge base (private domain vector database) and its own knowledge (local vector database) to give answers to domain questions.

[0030] The embodiment of the present application uses multiple local private domain knowledge (i.e., private domain vector databases) and multiple local highly reliable, timely, and professional knowledge information to effectively enhance the ability of large model question answering, and helps to overcome the defect of insufficient value content of a single local information, and accurately answer questions in related fields. In addition, the embodiment of the present application converts long text information into vectors after segmentation and stores them in the vector database, which facilitates subsequent large model retrieval to obtain recall fragments with high relevance, so as to improve information retrieval capabilities, accurately obtain factual knowledge related to the problem, and thus improve knowledge question answering capabilities.

[0031] 1.2. Split the long text information in advance and store it in the database in the form of vectors.

[0032] Reference Figure 4 , Figure 4 A schematic diagram of the storage process of long text information is shown. Figure 4 As shown, the embodiment of the present application proposes to first segment the long text information whose length exceeds the maximum segment text length, and then encode it into vectors. Therefore, the present embodiment selects a suitable vector database method and vector method to store the massive long text in the form of vectors in a vector database (multiple private domain vector databases and local vector databases).

[0033] In a possible implementation, the method further includes: Follow the steps below to convert long text information into vector representation using the vector model: Step S201, obtaining the text to be stored, and preprocessing the text to be stored.

[0034] Step S202, determining whether the length of the text to be stored exceeds the maximum segment text length.

[0035] Step S203, when the length of the text to be stored exceeds the maximum length of the text segment, segment the text to be stored to obtain a plurality of segmented text segments.

[0036] In a possible implementation, step S203, segmenting the text to be stored to obtain a plurality of segmented text segments, including: According to the maximum segment text length and the character overlap length, the text to be stored is segmented to obtain a plurality of segmented segment texts, wherein the maximum segment text length is the maximum length of the vector model, and the character overlap length is the length of the overlapping characters between the segment texts.

[0037] Specifically, segmentation needs to be based on two parameters. One parameter is the maximum segment text length. This parameter sets the maximum length of each segment of text. It is generally set according to the maximum length of the vector model (for text processing) and cannot exceed the maximum length of the vector model. The other parameter is the character overlap length. In order to maintain the semantic coherence of the segment, there are some overlapping characters between segments. The length of the overlapping characters is the overlap length. Based on these two parameters, the embodiment of the present application divides a long text into multiple segments, and there is overlap between the segments. The length of the segment does not exceed the maximum segment text length.

[0038] Step S204: convert the segmented text segments into vector representations to be stored through a vector model.

[0039] Step S205 , when the length of the text to be stored does not exceed the maximum segment text length, convert the text to be stored into a vector representation to be stored through a vector model.

[0040] Step S206: store the vector representation to be stored in a vector database.

[0041] Specifically, the embodiment of the present application uses a vector model to slice and vectorize a long text. First, a section of text is sequentially taken out from a large amount of long text, and then preprocessed, including text numbering, removing blank characters, merging multiple paragraphs in the text, and other preprocessing processes, and then it is determined whether the length of the text exceeds the preset segment length (i.e., the maximum segment text length). The specific value of the maximum segment text length can be set according to actual application requirements. If it exceeds the maximum segment text length, the text belongs to long text information and needs to be segmented.

[0042] After the segmentation is completed, Figure 4 As shown, the fragments need to be vector-encoded according to the vector model, and the embodiment of the present application is based on the selected vector model. It should be noted that during the storage process of the vector database, the method of the vector model cannot be changed, and the vector encoding of the same vector model needs to be stored in the same vector database.

[0043] After the vector encoding is completed, the vector is stored in the corresponding vector database (multiple private domain vector databases and local vector databases). Different vector databases have different storage methods and performances. After the storage is completed, the next text is processed sequentially (according to the above steps S201-S206) until all texts are converted into vectors and stored in the vector database. Therefore, each private domain vector database and local vector database stores the vector representation of the fragment text obtained by converting the long text information through the vector model.

[0044] 1.3 Text compression based on vector model.

[0045] In this embodiment, before segmentation, all possible titles and full texts can be combined into a complete long text message, and then segmented according to the length of the vector model. Taking BGE-Large as the vector model, the segmentation length is 512 tokens. Then, vector encoding can be performed, and the BGE-Large model can be used as the vector model. Its overall structure uses the BERT structure, which is similar to the LLama model and also includes a multi-layer encoder.

[0046] The overall method of the bge-large vector model is as follows: The first is the pre-training phase. In simple terms, it is to first randomly mask the text X, then encode it, and then train an additional light-weight decoder (such as a single-layer transformer) for reconstruction. Through this process, the encoder is forced to learn good embeddings.

[0047] Next is the comparative learning stage, which focuses on: using the in-batch negative sample method and using a large batch_size. For example, the size can be 19200. This stage focuses on simplicity and efficiency. As long as the batch is large enough, it is enough to find hard negative samples in the batch.

[0048] The final stage is the task-oriented fine-tuning stage, which mainly involves hard negative sampling. During the training process, an ANN-style sampling strategy is adopted to globally sample a hard negative sample with the closest embedding representation from the corpus of the task.

[0049] 1.4 Vector database method.

[0050] The embodiment of the present application uses a vector database-based method to store and search vectors (i.e., searching in the vector database in step S102 to obtain multiple relevant search information fragments). In this embodiment, the proposed vector databases can use the vector database method of 1.4.1 Chroma or the vector database method of 1.4.2 Lancedb to store and retrieve vector data.

[0051] 1.4.1 chroma.

[0052] Specifically, the first method is chroma, and the main method is the working principle of the (Hierarchical Navigable Small World graphs, HNSW) index.

[0053] The construction process of using HNSW clustering in Chroma: 1. Vector insertion: When building the HNSW index, each vector is inserted from high level to low level in turn. On each level, the closest node is found and edges are added according to the connection rules of the small world graph.

[0054] 2. Random level selection: Each newly inserted node will be randomly assigned a level, and will be inserted from the highest level to the corresponding level step by step.

[0055] 3. Connection maintenance: In order to maintain the properties of the small-world graph, HNSW maintains a limited number of connections for each node and ensures efficient search performance by optimizing the selection of adjacent nodes.

[0056] Search process using HNSW clustering in Chroma: 1. Start from the top level: The search starts from the top level of the graph, and the initial node is usually selected randomly.

[0057] 2. Layer-by-layer approximation: At each layer, by traversing the nodes of the current layer, find the node closest to the query vector and continue searching in its direction.

[0058] 3. Final result: In the bottom-level graph, find the node set closest to the query vector as the final search result 1.4.2 lancedb.

[0059] The second method is lancedb. Since the chroma loss vector recall is too much, this method is relatively light and fast. Even when accelerated by using non-clustering methods, the performance is still excellent.

[0060] Lancedb is built on the lance data format, an innovative columnar data format for machine learning. Lancedb uses an embedded serverless architecture. Lancedb has a relatively low resource utilization rate and is lighter and faster to retrieve.

[0061] For example, Lancedb can support storage in the float16 data format, which reduces storage space. At the same time, float16 quantization also speeds up the retrieval calculation speed, which is also an acceleration method.

[0062] Lancedb is quite different from other vector databases. It uses a new columnar data format in the data storage layer and a serverless architecture in the infrastructure. Therefore, Lancedb greatly reduces the complexity of the infrastructure and improves the performance of vector database retrieval.

[0063] 1.5 Reordering strategy.

[0064] Reference Figure 5 , Figure 5 A schematic diagram of the generation process of a recall fragment sequence is shown, such as Figure 5 As shown, after retrieving multiple relevant search information fragments (such as Figure 5 After the fragments 1-5 shown in the figure, the embodiment of the present application can adopt a variety of different reordering strategies to select a part of the multiple related search information fragments as the recall fragment (such as Figure 5 3, 2, 4) and sort them in a certain order ( Figure 5 The rerank in refers to the use of a reordering strategy to sort), forming a recall segment sequence. Figure 5 As shown, the dark-colored segments 1 and 5 represent segments that are determined not to be recalled segments.

[0065] In a possible implementation, the step S103 uses a reordering strategy to perform sorting to obtain a recall segment sequence consisting of a plurality of recall segments, including: Step S1031, obtaining the plurality of related search information segments arranged in a first order from high to low in similarity after deduplication.

[0066] Specifically, the question text is vector-encoded to obtain the vector represented by the question text, and then the similarity between the vector and each relevant search information fragment is calculated, and the first N relevant search information fragments are taken in descending order of similarity. The arrangement order of the first N relevant search information fragments is the first order (similarity from high to low).

[0067] Step S1032: sort the relevant search information segments respectively according to a plurality of reordering strategies to obtain a plurality of candidate recall segment sequences.

[0068] Step S1033: determining the recalled segment sequence according to the multiple candidate recalled segment sequences.

[0069] In a possible implementation, determining the recalled segment sequence according to the multiple candidate recalled segment sequences includes: Based on a preset scoring rule, each recalled segment is scored according to its position in the plurality of candidate recalled segment sequences, and all recalled segments are sorted according to the scores to obtain the recalled segment sequence.

[0070] For example, the pre-set scoring rules may include that when a recall segment is the first in a candidate recall segment sequence A, it can get 5 points, and when the recall segment is also the second in another candidate recall segment sequence B, it can get 3 points, and the total score is 8. By calculating the total score of each recall segment and then sorting them from high to low according to the score, the final recall segment sequence can be determined.

[0071] In a possible implementation, the relevant search information segments are sorted respectively according to a plurality of reordering strategies to obtain a plurality of candidate recall segment sequences, including at least one of the following: A-1: Based on the direct interception strategy, the N related search information segments with the highest similarity to the vector represented by the question text are sorted from high to low according to the similarity to obtain the first candidate recall segment sequence.

[0072] Specifically, the number of windows is set to N, and N segments (as recalled segments) with the highest similarity to the vector represented by the question text are selected from the multiple recalled relevant retrieval information segments, and they are sorted from high to low according to the similarity to generate a first candidate recalled segment sequence.

[0073] Exemplarily, the first candidate recall segment sequence is directly used as the final recall segment sequence, and the first candidate recall segment sequence is used as the input of knowledge. Figure 6 , Figure 6 A schematic diagram of the execution flow of a direct interception strategy is shown, such as Figure 6 As shown, N fragments are concatenated with the question text through the template (such as Figure 6 As shown, N is 3, and the first three segments are spliced ​​with the question. The dark-colored segments 4 and 5 are not spliced ​​this time), and the knowledge question-answering model is input. The template example is: "Please use the provided knowledge to answer the question. The following is the known context information: {text1}, {text2}, ···, {textN}. Given the above background information, please answer the question: {query}, answer: ". Among them, text is the text information (that is, the recalled segment), and query is the question. The direct interception strategy does not use the rerank method, but relies on the recall effect of the vector model to directly answer the question, which may cause the loss of recall. At the same time, this method only uses the large model once for question answering, so the speed is relatively fast.

[0074] A-2: Based on the relevance judgment reordering strategy, the relevant retrieval information segments are sorted to obtain a second candidate recall segment sequence.

[0075] A-3: Based on the score re-ranking strategy, the knowledge question and answer model is used to determine the relevance score between each of the relevant retrieval information segments and the question text, and the N relevant retrieval information segments with the highest relevance scores are sorted from large to small according to the relevance scores to obtain the third candidate recall segment sequence.

[0076] Specifically, refer to Figure 7 , Figure 7 A schematic diagram of the execution flow of a scoring re-ranking strategy is shown in FIG. Figure 7 As shown in the figure, for a question and a given number of relevant search information fragments (for example, a total of K fragments are retrieved), each fragment is first taken from the recalled relevant search information fragments in the first order, and then the single relevant search information fragment is spliced ​​with the question text through the template, and then input into the knowledge question answering model to determine the relevance score of the relevant search information fragment and the question text. An example of a template for determining relevance is shown in Figure 7As shown in the figure: "Please judge the relevance of the provided knowledge to the question and give a score. The score range is an integer from 1 to 10, with 10 being the highest score and 1 being the lowest score. The following is the knowledge information: {text1}, and the following is the question: {query}. Please give a score: ". Among them, text is the relevant search information fragment, and query is the question text.

[0077] Exemplarily, the third candidate recall fragment sequence is directly used as the final recall fragment sequence, and the third candidate recall fragment sequence is used as the knowledge input. The K relevant retrieval information fragments are reordered according to the relevance score results judged by the knowledge question and answer big model. First, the score results judged by the knowledge question and answer big model are taken out through the rules. If the knowledge question and answer big model does not give a specific relevance score, it is processed as 1 point, and the specific score of each fragment is stored in the score list. Then, according to the results of the relevance scores from high to low, the serial numbers of the fragments are reordered. It should be noted that the original order (i.e. the first order) should be retained for the same relevance score. After the sorting is completed, the fragments are reordered according to the serial number sorting results, and the first N fragment groups are selected as recall fragments to form the third recall sequence. According to the direct interception strategy method, after obtaining the third recall sequence, it is spliced ​​with the question text and input into the knowledge question and answer big model, and then a large model question and answer is performed (such as Figure 7 As shown in the figure, the base strategy is used to obtain the final question-answering result. The score re-ranking strategy uses a re-ranking method, relying on the large model to judge the relevance score between the fragment and the question for re-ranking. This method uses a large model (K+1) times for question-answering, where K refers to the total number of relevant search information fragments, K times for judging the relevance score, and 1 time for the final answer.

[0078] A-4: Based on the bubble reordering strategy, the relevant retrieval information fragments are sorted to obtain a fourth candidate recall fragment sequence.

[0079] A-5: Based on the keyword reordering strategy, the relevant search information segments are sorted to obtain a fifth candidate recall segment sequence.

[0080] A-6: Based on the machine learning sorting strategy, the relevant retrieval information segments are reordered using the reordering model to obtain a sixth candidate recall segment sequence.

[0081] Specifically, in the process of re-ranking, it is not necessary to use the knowledge question answering large model for understanding and re-ranking. A more sophisticated small model (i.e., the re-ranking model) can be used to re-rank the recalled text, and then the first N segments after re-ranking are selected as the recalled segments to form the sixth candidate recalled segment sequence) and input into the knowledge question answering large model for answering. Exemplarily, the bge-large-reranker model can be selected as the re-ranking model to re-rank the recalled text (i.e., the relevant retrieval information segment). Specifically, the bge-reranker-large model is a re-ranking model designed for text retrieval tasks, aiming to improve the accuracy and effectiveness of the retrieval system. It can re-rank the top documents returned by the retrieval system to improve the accuracy of retrieval. The core principle of this model is based on the structure of the cross-encoder, and the retrieval results are optimized by learning the interactive information between documents and queries. The bge-reranker-large model adopts the training strategy of contrastive learning, which is a method of training models to learn the representation of data by comparing positive examples and negative examples. During the training process, the model accepts data in the form of triples as input, including a query, a positive example, and a negative example. The goal of the model is to learn a representation space that makes the positive example closer to the query representation and the negative example farther away from the query representation.

[0082] 1.5.1 Related judgment reordering strategy.

[0083] This embodiment proposes that the relevant search information fragments can be sorted based on the relevance judgment re-ranking strategy using the knowledge question and answer big model to obtain a second candidate recall fragment sequence.

[0084] In a possible implementation, the relevant search information segments are sorted based on a relevance judgment reordering strategy to obtain a second candidate recall segment sequence, including: Step S301: splice each of the relevant search information fragments with the question text in accordance with the first order to obtain a first spliced ​​text.

[0085] Step S302: input the first concatenated text into the knowledge question and answer model to determine whether the relevant search information fragment is relevant to the question text.

[0086] Step S303, taking out the judgment result through the rule, and storing it in the first sequence or the second sequence according to the first order, wherein the first sequence includes: relevant retrieval information fragments judged by the large model to be related to the question text, and the second sequence includes: relevant retrieval information fragments judged by the large model to be irrelevant to the question text, and the arrangement order of the fragments in the first sequence and the second sequence are both maintained in the first order.

[0087] Step S304: concatenate the first sequence and the second sequence, with the first sequence in front and the second sequence in the back, to obtain a first concatenated sequence.

[0088] Step S305 , taking the first N relevant search information segments in the first concatenated sequence, and forming the second candidate recall segment sequence according to the arrangement order in the first concatenated sequence.

[0089] Specifically, first, each segment is sequentially taken from the recalled segments (the first-order related search information segments), and then the single segment is spliced ​​with the question text through the template, and input into the knowledge question answering model to determine the relevance of the segment to the question (that is, determine whether the relevant search information segment is related to the question text). An example of a template for determining relevance is shown in the figure: "Please determine whether the provided knowledge is relevant to the question, {text1}, given the above background information, determine whether it is relevant: {query}, please answer yes or no:". Among them, text is text information (that is, relevant search information segments), and query is the question text.

[0090] Exemplarily, the second candidate recalled segment sequence is directly used as the final recalled segment sequence, and the second candidate recalled segment sequence is used as the input of knowledge. Figure 8 , Figure 8 A schematic diagram of the execution flow of a related judgment reordering strategy is shown, such as Figure 8As shown, the relevant retrieval information fragments are reordered according to the results of the large model judgment. First, the results of the large model judgment are taken out according to the rules, and according to the judgment results of yes and no, they are stored in two sorted lists (first sequence or second sequence) in the first order, where the first sequence is a sequence composed of fragments judged by the large model to be related to the question, and the second sequence is a sequence composed of fragments judged by the large model to be irrelevant to the question. At the same time, the fragment order of the two sequences maintains the recall order of the original fragments (first order), that is, in each separate sequence (first sequence or second sequence), the higher the ranking, the higher the vector similarity between the fragment and the question. Finally, the first sequence and the second sequence are spliced, with the first sequence in front and the second sequence in the back, and the spliced ​​sequence is used as a new sorted list (that is, the first spliced ​​sequence). After obtaining the first concatenated sequence, according to the direct interception strategy method, the first N relevant search information fragments in the first concatenated sequence are taken as recall fragments, without changing their arrangement order in the first concatenated sequence, to form a second candidate recall fragment sequence, which is then concatenated with the question text and input into the knowledge question answering model to conduct a large model question answering (such as Figure 8 As shown in the figure, the base strategy is used for question answering) to obtain the final question answering result. The relevance judgment re-ranking strategy uses a re-ranking method, relying on a large model to judge whether the fragment is relevant to the question for re-ranking. This method uses a large model (K+1) times for question answering, where K refers to the total number of relevant retrieval information fragments, K times for judging relevance, and 1 time for the final answer.

[0091] 1.5.2 Bubble reordering strategy.

[0092] This embodiment proposes that the relevant search information fragments may be sorted based on a bubble reordering strategy to obtain a fourth candidate recall fragment sequence.

[0093] In a possible implementation, the relevant search information segments are sorted based on a bubble reordering strategy to obtain a fourth candidate recall segment sequence, including: Step 1, according to the first sequence, take the relevant search information fragments in the window W, input them into the knowledge question and answer big model, and make the knowledge question and answer big model sort them from high to low according to the relevance of each of the relevant search information fragments and the question text, to obtain a third sequence; Step 2, according to the first sequence, taking the relevant search information fragments with a step length of S, and splicing them with the relevant search information fragments with the highest correlation in the third sequence WS, to obtain a first spliced ​​sequence; Step 3, re-inputting the first spliced ​​sequence into the knowledge question and answer big model, so that the knowledge question and answer big model sorts the relevant search information fragments from high to low according to the relevance between the relevant search information fragments and the question text, to obtain a new third sequence; Step 4, repeat steps 2 and 3 until the sorting is completed, and WS relevant search information fragments with the highest relevance are obtained; Step 5: Determine S most relevant relevant retrieval information segments from the remaining multiple relevant retrieval information segments to obtain the fourth candidate recall segment sequence.

[0094] Specifically, refer to Fig. 9 , Fig. 9 A schematic diagram of the execution flow of a bubble reordering strategy is shown, such as Fig. 9 As shown in , the bubble reordering strategy needs to set two hyperparameters, the window size W and the step size S. For the question and the K relevant retrieval information fragments of the given recall, first, take the W relevant retrieval information fragments of the window size in the first order and input them into the knowledge question and answer model to sort the W fragments (that is, sort them from high to low according to the relevance to the question), then take out the WS best fragment results (that is, the WS fragments with the highest relevance), splice the next step size S relevant retrieval information fragments (that is, the first splicing sequence) and input them into the knowledge question and answer model for sorting (that is, step 3), until the sorting is completed, and finally get the WS relevant retrieval information fragments with the highest relevance (that is, step 4). For the sorting template example of W fragments, see Fig. 9 As shown in the example: "Sort the four texts according to their relevance to the question. The texts are as follows: {text1}, {text2}, {text3}, {text4}, and the following is the question: {query}". Where text is the text information (i.e. the relevant search information fragment), and query is the question text. For example, Fig. 9 As shown, the window size is 4, and segments 1-4 are selected for sorting with a step size of 2. According to the sorting results, the two segments with the highest correlation are obtained, namely Fig. 9 The fragments 1 and 3 are shown. Then, two new fragments are taken according to the first order for splicing, that is, fragments 5 and 6 are taken, and re-spliced ​​to obtain a sequence of a window size, which is re-input into the large model for judgment.

[0095] For each ranking result output by the knowledge question and answer big model, it is necessary to use rules to obtain the ranking information, and then save the WS best ranking results each time for the next splicing. When the big model cannot give the ranking information, take WS relevant retrieval information fragments in the original order (ie the first order) to continue the next splicing. After completing the final ranking result, the big model finally takes out the WS best relevant retrieval information fragments, and then takes out the best S relevant retrieval information fragments that do not belong to this sequence from the K relevant retrieval information fragments, and splices them into new W relevant retrieval information fragments (step 5). Specifically, repeat steps 1-4, and each round of steps 1-4 is performed to obtain WS relevant retrieval information fragments until W relevant retrieval information fragments are obtained, and the first N relevant retrieval information fragments are taken as recall fragments to form the fourth recall sequence. According to the direct interception method, it is spliced ​​with the question text and then input into the knowledge question and answer big model to perform a big model question and answer (such as Fig. 9 As shown in the figure, the base strategy is used for question answering) to obtain the final question answering result. The bubble reordering strategy uses a reordering method. When it is divisible, this method uses a large model ((KW) / S+2) times for question answering, of which ((KW) / S+1) times are used for sorting and 1 time is used for the final answer.

[0096] 1.5.3 Keyword re-ranking strategy.

[0097] This embodiment proposes to sort the relevant search information segments based on a keyword re-ranking strategy to obtain a fifth candidate recall segment sequence.

[0098] In a possible implementation, the relevant search information segments are sorted based on a keyword re-ranking strategy to obtain a fifth candidate recall segment sequence, including: Step S401, using the knowledge question and answer model, extract multiple keywords from the question text to obtain a keyword list.

[0099] Step S402: Concatenate each of the relevant search information fragments with the keyword list in the first order to obtain a second concatenated text.

[0100] Step S403: input the second concatenated text into the knowledge question and answer model to determine whether the relevant search information fragment is relevant to all the keywords in the keyword list.

[0101] Step S404, taking out the judgment result through the rule, and storing it in the fourth sequence or the fifth sequence according to the first order, wherein the fourth sequence includes: relevant search information fragments that the large model judges to be relevant to all the keywords in the keyword list, and the fifth sequence includes: relevant search information fragments that the large model judges to be irrelevant to part or all of the keywords in the keyword list, and the arrangement order of the fragments in the two lists is maintained in the first order.

[0102] Step S405, concatenating the fourth sequence and the fifth sequence, with the fourth sequence in front and the fifth sequence in the back, to obtain a second concatenated sequence.

[0103] Step S406: Take the first N relevant search information segments in the second spliced ​​sequence and form the fifth candidate recall segment sequence according to the arrangement order in the second spliced ​​sequence.

[0104] Specifically, refer to Fig.10 , Fig.10 A schematic diagram of the execution flow of a keyword reordering strategy is shown, such as Fig.10 As shown in the figure, first use the knowledge question answering model to generate a keyword list for the question text. The template example is as follows Fig.10 As shown: "Please extract the keyword list of the question. The following is the question: {query}, please extract the keyword list:". After the extraction is completed, the keyword list is taken out according to the rules. Then, for the given K recalled relevant retrieval information fragments, first take each relevant retrieval information fragment from the recalled relevant retrieval information fragments in the first order, and then splice the single relevant retrieval information fragment with the keyword list through the template, and input it into the knowledge question and answer model to determine whether the relevant retrieval information fragment is related to the keyword list. The template example is shown in the figure: "Please determine whether the provided knowledge is related to all the keywords in the keyword list. The following is the knowledge information: {text1}, the following is the keyword list: {list}, please answer yes or no:". Among them, text is the text information (that is, the relevant retrieval information fragment), and list is the keyword list. " Next, the K relevant retrieval information fragments are reordered according to the results of the judgment of the knowledge question and answer big model. First, the results of the big model judgment are taken out according to the rules, and the yes and no are stored in two sorted lists (the fourth sequence and the fifth sequence) in the first order, where the fourth sequence is a sequence composed of fragments that the big model judges to be all relevant to the keyword list, and the fifth sequence is a sequence composed of fragments that the big model judges to be not all relevant to the keyword list. At the same time, the fragment order of the fourth sequence and the fifth sequence maintains the recall order of the original fragments (that is, the first order), that is, in each separate list, the higher the ranking, the higher the vector similarity between the fragment and the question. Finally, the fourth sequence and the fifth sequence are spliced, with the fourth sequence in front and the fifth sequence in the back, and the splicing is used as a new sorted list (that is, the second spliced ​​sequence), and the length of the list is K. After obtaining the second spliced ​​sequence, the first N fragments are selected as recall fragments to form the fifth candidate recall fragment sequence. According to the direct interception method, it is spliced ​​with the question text and then input into the knowledge question and answer big model to conduct a big model question and answer (such as Fig.10 As shown in the figure, the base strategy is used for question answering), and the final question answering result is obtained. The re-ranking strategy for judging relevance uses a re-ranking method, relying on the knowledge question answering big model to judge whether the fragment is related to all the keywords in the keyword list for re-ranking. This method uses the big model (K+2) times for question answering, of which 1 is used to extract the keyword list, K times to judge whether it is related to the keyword list, and 1 time to give the final answer.

[0105] 1.6 Knowledge Question and Answer Model In this embodiment, the knowledge question and answer big model used may be the structure of the Llama big model.

[0106] Large models are generally based on the Transformer framework with some improvements. The original Transformer structure includes multiple layers of encoder and decoder structures. Llama is an open and efficient large language model, and its structure also uses the Transformer decoder-only structure. Llama's structure uses a decoder-only structure, which uses a 32-layer decoder.

[0107] At the same time, Llama made the following improvements: 1) Pre-normalization. In order to improve training stability, Llama normalized the input of each Transformer layer instead of normalizing the output. At the same time, Llama did not use the normalization function of Layer Norm, but used the RMS normalization function. The main difference is that the original mean part is removed. The author of RMS believes that this simplified normalization function can reduce about 7% to 64% of the time. 2) Use the activation function SwiGLU. Llama uses SwiGLU instead of ReLU as the activation function. SwiGLU is a variant of the activation function and a smoothed version of ReLU. It controls the shape of the function by adding a parameter. SwiGLU itself is an attempt to modify various activation functions. From the results, it has a smaller error than ReLU. 3) Rotary Position Embedding (RoPE). RoPE is also a commonly used improvement technology on the Transformer structure, which realizes relative position encoding by absolute position encoding. Unlike the original Transformer, which adds posembedding and token embedding, RoPE multiplies the position encoding and query or key to integrate the position information into the query. Llama3 is an improved model based on Llama, which uses more training data sets and larger context window capabilities. Its structure is the same as Llama. In this embodiment, the Llama3-6b model can be selected as the basic model structure of the knowledge question and answer model used in this application.

[0108] In summary, if Figure 2As shown, the scheme of the large model question-answering enhancement based on the structured knowledge base in the embodiment of the present application includes four main steps, using a vector model to slice the long text for vector representation (described in Section 1.2 above), using a vector database-based method to search for vectors (described in Sections 1.3 and 1.4 above), and then aggregating multiple related retrieval information fragments, and finally using a reordering strategy for knowledge answering (described in Section 1.5 above). The embodiment of the present application proposes to use the vector model method (described in Section 1.2 above) to represent the text as vector information, so it is necessary to segment the text according to the vector model, and then use the vector model to represent the fragment text. The vector model in this step will also serve as the encoding model of the question vector in the next step, and whether the question and the relevant text can be accurately matched depends on the expression ability of the vector model. In addition, the embodiment of the present application adopts the vector database method (described in Sections 1.3 and 1.4 above) to save the text vector information and provide services for vector retrieval. The main functions of this step include compression and accelerated retrieval to reduce time and space overhead. In addition, the model question answering strategy (described in Section 1.5 above) is used to answer questions with knowledge spliced ​​together using the knowledge question answering big model. The knowledge (i.e., the sequence of recall fragments) and questions are spliced ​​together using templates, and then input into the big model to get the answer, so that the big model can answer domain questions by combining the knowledge in the knowledge base and its own knowledge.

[0109] 1.7 Method effectiveness test.

[0110] The present application example also uses an experimental data set to test the above method. The following is the specific test content.

[0111] 1.7.1 Experimental Dataset

[0112] This embodiment uses a large amount of long text information without titles and keywords as the experimental data set. The content of the long text information is not restricted, and the number of original long text information exceeds 1 million. According to the vector model, the segmentation length is set to a maximum length of 512. In order to ensure the semantic coherence of the segmented segments, an overlap length of 20 is set. The specific segmentation method is described in Section 1.2 above. After segmentation, the total number of text segments is more than 3 million.

[0113] 1.7.2 Vector model experiment.

[0114] In this embodiment, multiple vector models are used for comparison in the experiment. Recall calculation is performed without using a vector database or acceleration method. The calculation method is cos similarity calculation. The comparison results of the experiment are as follows: The situations of the 10 data sets are as follows: Experiment No. 0, model Bge, Top-10 recall rate is 30%, Top-20 recall rate is 40%, and the average rank is 3464; Experiment No. 1, model wwn, Top-10 recall rate is 0, Top-20 recall rate is 0, and the average rank is 10w+; Experiment No. 2, model Text2vec, Top-10 recall rate is 30%, Top-20 recall rate is 30%, and the average rank is 4541; Experiment No. 3, model Bge-large, Top-10 recall rate is 90%, Top-20 recall rate is 90%, and the average rank is 6.2.

[0115] The comparative experiments of 20 data are as follows: Experiment No. 0, model Bge, Top-10 recall rate is 40%, Top-20 recall rate is 50%, and the average rank is 3752; Experiment No. 1, model wwm, Top-10 recall rate is 5%, Top-20 recall rate is 5%, and the average rank is 10w+; Experiment No. 2, model Text2vec, Top-10 recall rate is 35%, Top-20 recall rate is 35%, and the average rank is 3971; Experiment No. 3, model Bge-large, Top-10 recall rate is 85%, Top-20 recall rate is 90%, and the average rank is 8.8.

[0116] The comparative experiments of 50 data are as follows: Experiment No. 0, model Bge, Top-10 recall rate is 38%, Top-20 recall rate is 48%, and the average rank is 12114; Experiment No. 1, model wwm, Top-10 recall rate is 2%, Top-20 recall rate is 4%, and the average rank is 10w+; Experiment No. 2, model Text2vec, Top-10 recall rate is 26%, Top-20 recall rate is 28%, and the average rank is 7522; Experiment No. 3, model Bge-large, Top-10 recall rate is 74%, Top-20 recall rate is 76%, and the average rank is 490.9.

[0117] Among them, the Top-10 recall rate is the ratio of correct recall when 10 vectors are recalled. The Top-20 recall rate is the ratio of correct recall when 20 vectors are recalled. Generally, the recall is 10, which can be increased to 20 if there is re-ranking and sufficient time. At the same time, you can check the fault tolerance of the recall when it can be expanded to 20.

[0118] At the beginning, the embodiment of this application prepared 10 and 20 data for experiments. These data are relatively stable and the scope of the questions is relatively limited. Later, 50 data were added. The scope of the questions added later is less limited. There may be multiple articles involving answers or noise, and there may even be situations where the model can answer directly. It is observed that under the fluctuation of these problems, the average rank rate will decrease, especially some problems composed of common words may cause the vector ranking to be very low, but the ranking of each solution is still relatively stable.

[0119] The average rank result is the average of the correct recall rankings in all vector similarities. Taking the bge model result as an example, the rank results of 20 test data recalls are as follows: Test Id1, Rank4; Test Id2, Rank1619; Test Id3, Rank67; Test Id4, Rank703; Test Id5, Rank3629; Test Id6, Rank6368; Test Id7, Rank2; Test Id8, Rank22236; Test Id9, Rank1; Test Id10, Rank11; Test Id11, Rank1; Test ID 12, Rank 1; Test ID 13, Rank 1; Test ID 14, Rank 14; Test ID 15, Rank 38486; Test ID 16, Rank 45; Test ID 17, Rank 1; Test ID 18, Rank 51; Test ID 19, Rank 1795; Test ID 20, Rank 1; Average Rank 3752. It can be seen that in some cases, it is difficult to guarantee that the correct answer can be recalled even if it is expanded to a longer recall number.

[0120] 1.7.3 Vector database experiment.

[0121] This embodiment uses Chroma and LanceDB vector databases to conduct comparative experiments. The results of the comparative experiments of the vector databases are shown below.

[0122] The comparative experiments of 10 data are as follows: Experiment No. 0, model Bge, model recall result 30%, vector database Chroma, acceleration method Hnsw, comprehensive result 10%; Experiment No. 1, model Text2vec, model recall result 30%, vector database Chroma, acceleration method Hnsw, comprehensive result 10%; Experiment No. 2, model Bge-large, model recall result 90%, vector database Chroma, acceleration method Hnsw, comprehensive result 20%; Experiment No. 3, model Bge-large, model recall result 90%, vector database Lancelb, acceleration method Float16, comprehensive result 90%.

[0123] The comparative experiments of 20 data are as follows: Experiment No. 0, model Bge, model recall result 40%, vector database Chroma, acceleration method Hnsw, comprehensive result 10%; Experiment No. 1, model Text2vec, model recall result 35%, vector database Chroma, acceleration method Hnsw, comprehensive result 5%; Experiment No. 2, model Bge-large, model recall result 85%, vector database Chroma, acceleration method Hnsw, comprehensive result 20%; Experiment No. 3, model Bge-large, model recall result 85%, vector database Lancelb, acceleration method Float16, comprehensive result 85%.

[0124] The comparative experiments of 50 data are as follows: Experiment No. 0, model Bge, model recall result 38%, vector database Chroma, acceleration method Husw, comprehensive result 14%; Experiment No. 1, model Text2vec, model recall result 26%, vector database Chroma, acceleration method Husw, comprehensive result 8%; Experiment No. 2, model Bge-large, model recall result 74%, vector database Chroma, acceleration method Husw, comprehensive result 32%; Experiment No. 3, model Bge-large, model recall result 74%, vector database Lancelb, acceleration method Float16, comprehensive result 74%.

[0125] At the same time, this embodiment also conducts a speed experiment, and verifies it by taking 20 pieces of data as an example. The speed experiment results are as follows: Experiment No. 0, model Bge, vector database Chroma, acceleration method Hnsw, average retrieval time 1.01 seconds; Experiment No. 1, model Text2vec, vector database Chroma, acceleration method Hnsw, average retrieval time 1.72 seconds; Experiment No. 2, model Bge-large, vector database Chroma, acceleration method Hnsw, average retrieval time 0.94 seconds; Experiment No. 3, model Bge-large, vector database Lancelb, acceleration method Float16, average retrieval time 2.81 seconds; Experiment No. 4, model Bge-large, no vector database, no acceleration method, average retrieval time 1325 seconds.

[0126] At the same time, this embodiment also conducts a speed experiment, and verifies it by taking 50 data as an example. The speed experiment results are as follows: Experiment No. 0, model Bge, vector database Chroma, acceleration method Hnsw, average retrieval time 0.98 seconds; Experiment No. 1, model Text2vec, vector database Chroma, acceleration method Hnsw, average retrieval time 0.45 seconds; Experiment No. 2, model Bge-large, vector database Chroma, acceleration method Hnsw, average retrieval time 1.68 seconds; Experiment No. 3, model Bge-large, vector database Lancelb, acceleration method Float16, average retrieval time 0.92 seconds; Experiment No. 4, model Bge-large, no vector database used, no acceleration method used, average retrieval time 1290 seconds.

[0127] This embodiment conducts an experiment without using a vector database, as shown in Experiment No. 4. The method of this experiment is to load a vector document stored using jsonline, one line at a time, and then compare the vector with the query to see if they are similar. Because the number of queries is relatively small, all queries will be compared and recorded at one time, without the need to read all vector documents for each query. In this case, this embodiment averages all the time divided by the number of queries to obtain the average retrieval time of the query, which is far longer than the case where the vector database is not used.

[0128] Through experiments, it is found that the recall loss of the Chroma vector database is high, regardless of the vector model. At the same time, LanceDB does not use clustering for acceleration, so there is basically no loss of effect. At the same time, the time cost is similar to Chroma. Therefore, the embodiment of this application finally selected the Lancedb vector database method for subsequent experiments.

[0129] 1.7.4 Knowledge question and answer strategy experiment.

[0130] In terms of knowledge question answering strategies, this embodiment uses the five reordering strategies mentioned above (see the relevant content of the reordering strategies described in Section 1.5 above) and the direct question answering method to conduct a comparative experiment on knowledge question answering. The following experiments are based on the bge-large vector model and the lancedb vector database method.

[0131] The overall experimental results are as follows: Experiment No. 1, question and answer strategy "direct question and answer", no rerank, no knowledge base, effect 12%; Experiment No. 2, question and answer strategy "direct interception strategy", no rerank, use of knowledge base, effect 66%; Experiment No. 3, question and answer strategy "judgment related reranking strategy", rerank, use of knowledge base, effect 72%; Experiment No. 4, question and answer strategy "scoring reranking strategy", rerank, use of knowledge base, effect 68%; Experiment No. 5, question and answer strategy "bubble reranking strategy", rerank, use of knowledge base, effect 68%; Experiment No. 6, question and answer strategy "keyword reranking strategy", rerank, use of knowledge base, effect 66%; Experiment No. 7, question and answer strategy "machine learning ranking strategy", rerank, use of knowledge base, effect 70%.

[0132] It can be seen that experiments 2-6 used a knowledge base to enhance the question-answering ability of the large model, which was more than 50% higher than that of experiment 1, indicating that the overall massive long-text knowledge enhancement strategy has a significant improvement in the question-answering ability of the large model. Experiment 2 did not use the rerank strategy. As the simplest direct interception strategy, it had the lowest accuracy in experiments 2-6, and we also used it as the base strategy for subsequent experiments. Experiments 3-6 all used the rerank method, among which the judgment-related reordering strategy method of experiment 3 had the best effect, with a question-answering accuracy of 72%. It should be noted that the vector recall result is 74%, that is, only 74% of the most concerned related text is input into the large model, and the recall results will be cropped due to length.

[0133] Next, this embodiment conducts separate experimental effect analysis on the five methods of experimental numbers 2-6.

[0134] The effects of the direct interception strategy method are as follows: Category No. 0, Number of categories: 27, vector recall was successful, interception was successful, answer was successful, and the reason for success was: through knowledge; Category No. 1, Number of categories: 3, vector recall was successful, interception was successful; Category No. 2, Number of categories: 4, vector recall was successful; Category No. 3, Number of categories: 3, vector recall was successful, answer was successful, and the reason for success was: some other materials were recalled; Category No. 4, Number of categories: 10, all failed; Category No. 5, Number of categories: 3, answer was successful, and the reason for success was: some materials were recalled or the model itself answered; Total: Number of categories: 50, vector recall was successful 37 times (74%), interception was successful 30 times (60%), and answer was successful 33 times (66%).

[0135] The direct interception strategy method will first truncate the recall, which will cause some recall loss. Among the 37 vector recalls, 4 recall possibilities are lost through interception. Among the remaining 33, recall is lost through interception, but 1 is answered by itself, and part of the information is successfully recalled. Among the 37, 30 are answered correctly, of which 3 lose correct recall through interception, but some materials can still be recalled, so the answer is successful. 7 answers are wrong, of which 4 lose correct recall through interception, 3 recall materials correctly, but the model answers incorrectly. 13 are not recalled, 10 are answered incorrectly, 3 are answered correctly, 1 is answered by the model itself, and 2 are answered by correctly recalling part of the materials. That is, among the 37 recalls, the correct knowledge of the remaining 30 is entered into the model through interception, and the model answers 27 of the 30 successfully. The other 7 did not enter the model, but 3 still have some materials recalled and answered successfully. Of the 13 items that were not recalled, 2 were partially recalled and 1 was the model itself.

[0136] Therefore, most of the vector recall results can be intercepted and successfully answered, while some questions can also be answered correctly through the model itself and some materials.

[0137] The effects of the relevant reordering strategies are judged as follows: Category No. 0, Number of categories 27, vector recall successful, interception successful, Base strategy answer successful, sorting successful (including some materials), answer successful, reason: through knowledge; Category No. 1, Number of categories 3, vector recall successful, interception successful; Category No. 2, Number of categories 2, vector recall successful; Category No. 3, Number of categories 2, vector recall successful, reason: answer through sorting; Category No. 4, Number of categories 3, vector recall successful, Base strategy answer successful, sorting successful (including some materials), answer successful, reason: some other materials were recalled; Category No. 5, Number of categories 8, all failed; Category No. 6, Number of categories 2, sorting successful (including some materials), answer successful, reason: recall some materials through sorting; Category No. 7, Number of categories 2, Base strategy answer successful, sorting successful (including some materials), answer successful, reason: answer through some materials; Category No. 8, Number of categories 1, Base strategy answer successful, reason: exclude some materials through sorting. Total: Number of categories: 50, vector recall was successful 37 times (74%), interception was successful 30 times (60%), Base strategy was successful 33 times (66%), sorting was successful (including partial materials) 36 times (72%), and answering was successful 36 times (72%).

[0138] Among them, the base strategy is a direct interception strategy. Through comparison, it is found that compared with the base strategy, the judgment-related reordering strategy successfully answered two texts through successful sorting and successful sorting of partial materials respectively. At the same time, because the sorting method lost a partial material, the answer failed, which improved the effect by 6% compared with the base strategy.

[0139] The effects of the scoring reordering strategy are as follows: Category number 0, number of categories 25, vector recall success, interception success, Base strategy answer success, sorting success (including some materials), answer success, success reason: through knowledge; Category number 1, number of categories 2, vector recall success, interception success, Base strategy answer success, success reason: excluding correct knowledge through sorting; Category number 2, number of categories 3, vector recall success, interception success; Category number 3, number of categories 2, vector recall success; Category number 4, number of categories 2, vector recall success, sorting success (including some materials), answer success, success reason: answer through sorting; Category number 5, number of categories 3, vector recall success, Base strategy The answer was successful, the sorting was successful (including some materials), the answer was successful, and the reason for success was: some other materials were recalled; Category number 6, number of categories 9, all failed; Category number 7, number of categories 1, sorting was successful (including some materials), the answer was successful, and the reason for success was: answered by sorting some materials; Category number 8, number of categories 3, Base strategy answered successfully, sorting was successful (including some materials), the answer was successful, and the reason for success was: answered by sorting some materials; Total: number of categories 50, vector recall was successful 37 times (74%), interception was successful 30 times (60%), Base strategy answered successfully 33 times (66%), sorting was successful (including some materials) 34 times (68%), and answer was successful 34 times (68%).

[0140] By comparing the score re-ranking strategy with the base strategy, it was found that this method successfully ranked and partially ranked two texts and one text respectively. At the same time, because the ranking method lost two texts, the effect was improved by 2% compared with the base strategy.

[0141] The results of the bubble reordering strategy are as follows: Category No. 0, Number of categories 25, vector recall successful, interception successful, Base strategy answer successful, sorting successful (including some materials), answer successful, reason for success: through knowledge; Category No. 1, Number of categories 2, vector recall successful, interception successful, Base strategy answer successful, reason for success: excluding correct knowledge through sorting; Category No. 2, Number of categories 3, vector recall successful, interception successful; Category No. 3, Number of categories 4, vector recall successful, sorting successful (including some materials), answer successful, reason for success: answer through sorting; Category No. 4, Number of categories 5, vector recall successful, Base strategy answer successful, sorting successful (including some materials), answer successful, reason for success: some other materials were recalled; Category No. 5, Number of categories 10, all failed; Category No. 6, Number of categories 2, Base strategy answer successful, answer successful, reason for success: answer through sorting some materials; Category No. 7, Number of categories 1, Base strategy answer successful, reason for success: excluding some materials through sorting. Total: 50 categories, 37 successful vector recalls (74%), 30 successful interceptions (60%), 33 successful Base strategy answers (66%), 34 successful sorts (68%), and 34 successful answers (68%).

[0142] Through comparison, the knowledge question-answering strategy experiment found that compared with the base strategy, the bubble reordering strategy successfully answered two texts through successful sorting and successful sorting of partial materials. At the same time, it failed to answer because the sorting method lost a part of the material, which was 2% better than the base strategy.

[0143] The results of the keyword re-ranking strategy are as follows: Category No. 0, Number of categories 27, Vector recall successful, interception successful, Base strategy answer successful, sorting successful (including some materials), answer successful, reason for success: through knowledge; Category No. 1, Number of categories 3, Vector recall successful, interception successful, Base strategy unsuccessful, sorting unsuccessful, answer unsuccessful; Category No. 2, Number of categories 3, Vector recall successful, interception unsuccessful, Base strategy unsuccessful, sorting unsuccessful, answer unsuccessful; Category No. 3, Number of categories 1, Vector recall successful, interception unsuccessful, Base strategy unsuccessful, sorting successful (including some materials), answer successful, reason for success: answer through sorting; Category No. 4, Number of categories 3, Vector recall successful, interception unsuccessful, Base strategy answer successful Success, sorting success (including some materials), answering success, success reason: some other materials were recalled; Category number 5, number of categories 9, vector recall failed, interception failed, Base strategy failed, sorting failed, answering failed; Category number 6, number of categories 1, vector recall failed, interception failed, Base strategy failed, sorting success (including some materials), answering success, success reason: answering by sorting some materials; Category number 7, number of categories 1, vector recall failed, interception failed, Base strategy answering success, sorting success (including some materials), answering success, success reason: answering by sorting some materials; Category number 8, number of categories 2, vector recall failed, interception failed, Base strategy answering success, sorting failed, answering failed. Total: number of categories 50, vector recall success 37 times (74%), interception success 30 times (60%), Base strategy answering success 33 times (66%), sorting success 33 times (66%), answering success 33 times (66%).

[0144] Through comparison, it was found that compared with the base strategy, the keyword reordering strategy successfully answered one text through successful sorting and successful sorting of partial materials. At the same time, it failed to answer because two partial materials were lost in the sorting method, which was equivalent to the base strategy.

[0145] The results of the machine learning sorting strategy are as follows: Category No. 0, Number of categories 27, Vector recall successful, interception successful, Base strategy answer successful, sorting successful (including some materials), answer successful, reason for success: through knowledge; Category No. 1, Number of categories 1, Vector recall successful, interception unsuccessful, Base strategy answer successful, sorting unsuccessful, answer successful, reason for success: some materials and the model itself answered successfully; Category No. 2, Number of categories 2, Vector recall successful, interception unsuccessful, Base strategy answer successful, sorting successful (including some materials), answer unsuccessful, reason for success: correct sorting but not correct answer to the question; Category No. 3, Number of categories 2, Vector recall successful, interception successful, Base strategy unsuccessful, sorting unsuccessful, answer unsuccessful; Category No. 4, Number of categories 1, Vector recall successful, interception successful, Base strategy unsuccessful, sorting unsuccessful, answer successful, reason for success: adjusted The internal sequence of interception was answered successfully; Category No. 5, Number of categories 2, Vector recall was successful, interception was unsuccessful, Base strategy was unsuccessful, sorting was successful (including some materials), and the answer was successful. The reason for success was: answering through sorting; Category No. 6, Number of categories 2, Vector recall was successful, interception was unsuccessful, Base strategy was unsuccessful, sorting was unsuccessful, and the answer was unsuccessful; Category No. 7, Number of categories 8, Vector recall was unsuccessful, interception was unsuccessful, Base strategy was unsuccessful, sorting was unsuccessful, and the answer was unsuccessful; Category No. 8, Number of categories 2, Vector recall was unsuccessful, interception was unsuccessful, Base strategy was unsuccessful, sorting was successful (including some materials), and the answer was successful. The reason for success was: recalling some materials through sorting and the model’s own capabilities; Category No. 9, Number of categories 3, Vector recall was unsuccessful, interception was unsuccessful, Base strategy answered successfully, sorting was successful (including some materials), and the answer was successful. The reason for success was: answering through some materials. Total: 50 categories, 37 successful vector recalls (74%), 30 successful interceptions (60%), 33 successful Base strategy answers (66%), 36 successful sorts (72%), and 36 successful answers (72%).

[0146] Through comparison, it was found that compared with the base strategy, the sorting strategy based on the machine learning model successfully answered five texts by adjusting the order, successfully sorting, and successfully sorting some materials. At the same time, it failed to answer the question because it sorted correctly but did not answer the question correctly, which was equivalent to the judgment-related re-sorting strategy.

[0147] 1.7.5 Summary of experimental results.

[0148] In the data set, this embodiment selected more than 1 million long text messages for the experiment. After segmentation, there were more than 3 million text segments for the experiment. At the same time, 50 questions were marked for testing.

[0149] In the vector model experiment, this embodiment uses text2vec, bge, wwm, bge-large and other methods for experiments. The effect of the bge-large method is obviously better than other methods. Since the text is long, a better vector method is needed, and the bge-large method is selected to continue the experiment.

[0150] In the vector database experiment, two vector database methods, Chroma and LanceDB, were used for the experiment. Both methods have good performance improvement, but the Chroma method has a higher recall rate loss due to the use of clustering acceleration, so this embodiment uses the Lancedb method to continue the experiment.

[0151] In the knowledge answer strategy experiment, this embodiment uses a variety of answer strategy methods for experiments, including 5 re-ranking methods and 1 direct interception method. At the same time, the model directly answers questions. The experimental results show that the effect of using knowledge question-answering enhancement is much better than that of direct questions. The effect of using the re-ranking method is better or equal to that of the direct interception strategy. Among them, the re-ranking strategy method based on machine learning and judgment is better.

[0152] In order to solve the problem of difficulty in obtaining accurate factual knowledge in large models, the embodiments of the present application make full use of multiple local private domain knowledge to accurately answer questions in related fields. Local domain knowledge has knowledge of rich information, and local knowledge has the characteristics of visibility, confirmability, and easy transferability. The use of multiple local highly reliable, timely, and professional knowledge information can effectively enhance the question-answering capabilities of large models, and help to break the defect of insufficient value content of a single local information. At the same time, local knowledge is also easier to update, and it is also convenient for users to confirm the content of updated local knowledge, and the update of local knowledge avoids the problem of insufficient timeliness of large models. In addition, by combining multiple reordering strategies, the correlation between the acquired recall fragments and the questions is further improved, and the ability of large models to obtain accurate factual knowledge is improved.

[0153] The second aspect of the embodiment of the present application further provides a large model question-answering device based on multi-party private domain long text information, the device is used to execute the steps in the large model question-answering method described in the first aspect; the device includes: The question acquisition module is used to obtain the question text input by the user; A retrieval module, configured to perform vector retrieval in a plurality of private domain vector databases and a local vector database according to the question text, and obtain a plurality of relevant retrieval information fragments; wherein each of the private domain vector database and the local vector database stores a vector representation obtained by converting long text information through a vector model; the long text information refers to text information whose length exceeds the maximum fragment text length; A sorting module is used to aggregate the multiple related search information fragments, remove duplicates of the related search information fragments with similarity higher than a threshold, and sort them using a re-sorting strategy to obtain a recall fragment sequence composed of multiple recall fragments; The knowledge question-answering model is used to use a template to splice the recall segment sequence and the question text, and to answer the question to obtain an answer text.

[0154] The large model question and answer device based on long text information in the private domain of multiple parties provided in the embodiment of the present application can implement the various processes implemented by the large model question and answer method embodiment based on long text information in the private domain of multiple parties described in the first aspect, and achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0155] An embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the various processes of the above-mentioned large model question and answer method embodiment based on long text information in the private domains of multiple parties are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0156] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, each process of the above-mentioned large model question-answering method embodiment based on multi-party private domain long text information is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. Among them, the processor is the processor in the terminal device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a disk or an optical disk, etc.

[0157] The embodiments of the present application further provide a computer program / program product, which is stored in a storage medium and is executed by at least one processor to implement the various processes of the above-mentioned large-model question-and-answer method embodiment based on long text information in the private domains of multiple parties, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0158] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0159] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A large-model question-answering method based on multi-party private domain long text information, characterized in that: The method comprises: Get the question text entered by the user; According to the question text, vector retrieval is performed in multiple private domain vector databases and local vector databases to obtain multiple relevant retrieval information fragments; wherein each of the private domain vector database and the local vector database stores a vector representation obtained by converting long text information through a vector model; the long text information refers to text information whose length exceeds the maximum fragment text length; Aggregating the multiple related search information fragments, removing duplicates from the related search information fragments whose similarity is higher than a threshold, and sorting them using a reordering strategy to obtain a recall fragment sequence consisting of multiple recall fragments; The recall segment sequence and the question text are spliced ​​together using a template, and are input into the knowledge question-answering model for answering to obtain an answer text.

2. The large model question answering method according to claim 1, characterized in that: The method further comprises: Follow the steps below to convert long text information into vector representation using the vector model: Acquire the text to be stored, and preprocess the text to be stored; Determine whether the length of the text to be stored exceeds the maximum fragment text length; When the length of the text to be stored exceeds the maximum length of the text segment, segmenting the text to be stored to obtain a plurality of segmented text segments; The segmented text segments are converted into vector representations to be stored through a vector model; When the length of the text to be stored does not exceed the maximum length of the text segment, converting the text to be stored into a vector representation to be stored by using a vector model; The vector representation to be stored is stored in a vector database.

3. The large model question answering method according to claim 2, characterized in that: The text to be stored is segmented to obtain a plurality of segmented text segments, including: According to the maximum segment text length and the character overlap length, the text to be stored is segmented to obtain a plurality of segmented segment texts, wherein the maximum segment text length is the maximum length of the vector model, and the character overlap length is the length of the overlapping characters between the segment texts.

4. The large model question answering method according to claim 1, characterized in that: The reordering strategy is used to perform sorting to obtain a recall segment sequence consisting of multiple recall segments, including: Acquire the plurality of related search information fragments arranged in a first order from high to low in similarity after deduplication; According to a plurality of reordering strategies, the relevant search information fragments are respectively ordered to obtain a plurality of candidate recall fragment sequences; The recall segment sequence is determined according to the multiple candidate recall segment sequences.

5. The large model question answering method according to claim 4, characterized in that: According to a plurality of reordering strategies, the relevant search information segments are respectively ordered to obtain a plurality of candidate recall segment sequences, including at least one of the following: Based on the direct interception strategy, the N related search information segments with the highest similarity to the vector represented by the question text are sorted from high to low according to the similarity to obtain a first candidate recall segment sequence; Based on the relevance judgment reordering strategy, the relevant search information segments are ordered to obtain a second candidate recall segment sequence; Based on the score re-ranking strategy, the knowledge question answering model is used to determine the relevance score between each of the relevant search information segments and the question text, and the N relevant search information segments with the highest relevance scores are sorted from large to small according to the relevance scores to obtain a third candidate recall segment sequence; Based on the bubble reordering strategy, the relevant search information fragments are sorted to obtain a fourth candidate recall fragment sequence; Based on the keyword reordering strategy, the relevant search information fragments are ordered to obtain a fifth candidate recall fragment sequence; Based on the machine learning sorting strategy, the relevant retrieval information fragments are reordered using a rearrangement model to obtain a sixth candidate recall fragment sequence.

6. The large model question answering method according to claim 5, characterized in that: Based on the relevance judgment reordering strategy, the relevant search information segments are ordered to obtain a second candidate recall segment sequence, including: According to the first sequence, each of the relevant search information fragments is concatenated with the question text to obtain a first concatenated text; Inputting the first concatenated text into the knowledge question and answer model to determine whether the relevant search information fragment is relevant to the question text; The judgment result is retrieved by the rule and stored in the first sequence or the second sequence according to the first order, wherein the first sequence includes: relevant search information fragments that are judged by the large model to be relevant to the question text, and the second sequence includes: relevant search information fragments that are judged by the large model to be irrelevant to the question text, and the arrangement order of the fragments in the first sequence and the second sequence are both maintained in the first order; splicing the first sequence and the second sequence, with the first sequence in front and the second sequence in the back, to obtain a first spliced ​​sequence; The first N relevant search information segments in the first spliced ​​sequence are taken and, according to the arrangement order in the first spliced ​​sequence, are used to form the second candidate recall segment sequence.

7. The large model question answering method according to claim 5, characterized in that: Based on the bubble reordering strategy, the relevant search information fragments are sorted to obtain a fourth candidate recall fragment sequence, including: Step 1, according to the first sequence, take the relevant search information fragments in the window W, input them into the knowledge question and answer big model, and make the knowledge question and answer big model sort them from high to low according to the relevance of each of the relevant search information fragments and the question text, to obtain a third sequence; Step 2: according to the first sequence, take the relevant search information fragments with a step length of S, and splice them with the relevant search information fragments with the highest correlation in the third sequence WS, to obtain a first spliced ​​sequence; Step 3, re-inputting the first concatenated sequence into the knowledge question and answer big model, so that the knowledge question and answer big model sorts the relevant search information fragments from high to low according to the relevance between the relevant search information fragments and the question text, to obtain a new third sequence; Step 4, repeat steps 2 and 3 until the sorting is completed, and WS relevant search information fragments with the highest relevance are obtained; Step 5: Determine S most relevant relevant retrieval information segments from the remaining multiple relevant retrieval information segments to obtain the fourth candidate recall segment sequence.

8. The large model question answering method according to claim 5, characterized in that: Based on the keyword reordering strategy, the relevant search information fragments are ordered to obtain a fifth candidate recall fragment sequence, including: Utilizing the knowledge question answering model, extracting multiple keywords from the question text to obtain a keyword list; According to the first sequence, each of the relevant search information fragments is concatenated with the keyword list to obtain a second concatenated text; Inputting the second concatenated text into the knowledge question and answer macromodel to determine whether the relevant search information fragment is relevant to all the keywords in the keyword list; The judgment result is retrieved by the rule and stored in the fourth sequence or the fifth sequence according to the first sequence, wherein the fourth sequence includes: relevant search information fragments that are judged by the large model to be relevant to all the keywords in the keyword list, and the fifth sequence includes: relevant search information fragments that are judged by the large model to be irrelevant to some or all of the keywords in the keyword list, and the order of the fragments in the two lists is maintained in the first order; splicing the fourth sequence and the fifth sequence, with the fourth sequence in front and the fifth sequence in the back, to obtain a second spliced ​​sequence; The first N relevant search information segments in the second spliced ​​sequence are taken and, in accordance with the order of arrangement in the second spliced ​​sequence, are used to form the fifth candidate recall segment sequence.

9. The large model question answering method according to claim 5, characterized in that: Determining the recalled segment sequence according to the plurality of candidate recalled segment sequences includes: Based on a preset scoring rule, each recalled segment is scored according to its position in the plurality of candidate recalled segment sequences, and all recalled segments are sorted according to the scores to obtain the recalled segment sequence.

10. A large model question-answering device based on multi-party private domain long text information, characterized in that: The device is used to execute the steps in the large model question answering method according to any one of claims 1 to 9; wherein the device comprises: The question acquisition module is used to obtain the question text input by the user; A retrieval module, configured to perform vector retrieval in a plurality of private domain vector databases and a local vector database according to the question text, and obtain a plurality of relevant retrieval information fragments; wherein each of the private domain vector database and the local vector database stores a vector representation obtained by converting long text information through a vector model; the long text information refers to text information whose length exceeds the maximum fragment text length; A sorting module is used to aggregate the multiple related search information fragments, remove duplicates of the related search information fragments with similarity higher than a threshold, and sort them using a re-sorting strategy to obtain a recall fragment sequence composed of multiple recall fragments; The knowledge question-answering model is used to use a template to splice the recall segment sequence and the question text, and to answer the question to obtain an answer text.

Citation Information

Patent Citations

  • Question and answer method and device and electronic equipment

    CN117194646A

  • Online intelligent question answering method and device based on instruction fine tuning and retrieval enhancement generation

    CN117688163A

  • Search question-answering system and method based on large model and electronic equipment

    CN117708274A

  • Traditional Chinese medicine question and answer method and device based on long document retrieval enhancement generation and medium

    CN117828050A

  • Multi-modal multi-scale multi-recall large language model retrieval enhancement generation method

    CN118296120A

Cited By

  • Individual and enterprise knowledge base collaborative search and backflow method based on privacy protection

    CN121786272A

  • Data retrieval method and device, electronic equipment and readable storage medium

    CN122112042A