Information storage method, retrieval method and question answering method

By building an index tree containing traceability structure, the information loss caused by information slicing in the question-and-answer model is solved, which improves the accuracy and speed of reply and reduces resource usage.

CN120336479APending Publication Date: 2025-07-18DATA SPACE RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510422217.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing question and answer model causes global and local information to be lost during the information slicing process, reducing the reply accuracy.

Method used

Build an index tree containing traceability structure, form an index tree through slice text, retain the global and local information of the document, and use clustering nodes and parent-child relationships to reduce calls to large language models.

Benefits of technology

It greatly improves the reply accuracy and reply speed of the Q&A model, and reduces the use of computing resources and storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336479A_ABST
    Figure CN120336479A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information retrieval, and particularly relates to an information storage method, a retrieval method and a question answering method. The invention discloses an information storage method. The method comprises the following steps: S1, converting newly acquired information into a new document; one piece of new information corresponds to one piece of new document; s2, slicing the new document to form a plurality of sliced texts; s3, constructing an index tree of the current document based on all the slice texts of the new document, and storing the index tree; one new document corresponds to m index trees, and m is a positive integer. According to the information storage method, the index tree containing the traceability structure can be formed, the loss amount of information in the document is greatly reduced, and the reply accuracy and reply speed of a follow-up question and answer model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information retrieval, and particularly relates to an information storage method, a retrieval method, and a question answering method. Background Art

[0002] With the development of technology, more and more question answering models are used to assist human learning and life. No matter what kind of question answering model, the essence of serving humans is to make corresponding answers according to user needs, such as the RAG (Retrieval-Augmented Generation) model.

[0003] At present, all answers made by question answering models are not generated out of thin air. Instead, after the question answering model grabs information related to the demand and records it as basic information, operations such as integrating and reasoning these basic information are performed, and finally an answer can be made. Therefore, no matter how the integration and reasoning process of the question answering model is, the accuracy of the answer it makes is closely related to the basic information.

[0004] In the prior art, the question answering model slices all the collectible information into text documents and then stores the sliced text in a database, so as to conveniently grab the sliced text related to the demand as basic information through the keywords in the demand later. Compared with storing the whole document in the database and grabbing the whole document as basic information, slicing the document not only reduces the overall field range of the grabbed basic information, but also can reduce the introduction of irrelevant information, improve the grabbing speed of the basic information, and enhance the answering efficiency of the question answering model. However, the sliced text will cause the loss of some global information and local information in the document (global information refers to the document overview that can only be obtained based on the whole document, and local information refers to the connection between some sliced texts or the overview that can only be obtained based on some related sliced texts), which results in a decrease in the accuracy of the answer finally made by the question answering model. Summary of the Invention

[0005] The object of the present invention is to overcome the above-mentioned deficiencies of the prior art, and provide an information storage method that can form an index tree containing a traceability structure, greatly reducing the amount of information lost in the document, and improving the answer accuracy and answer speed of the subsequent question answering model.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] An information storage method includes the following steps:

[0008] S1, convert the newly acquired information into a new document; one new piece of information corresponds to one new document;

[0009] S2, slice the new document to form several sliced texts;

[0010] S3. Based on all the sliced texts of a new document, construct and store the index tree of the current document; a new document corresponds to m index trees, where m is a positive integer.

[0011] Preferably, in S2, the following sub-steps are further included:

[0012] S21. Use a segmentation algorithm to segment a new document into n sliced texts. Denote the set of sliced texts of the current new document as D, then D = (D1, D2,..., D i ,..., D n ), where 1 ≤ i ≤ n and both i and n are positive integers, and D i represents the i-th sliced text in the current new document;

[0013] S22. Encode each sliced text into a corresponding sliced vector. Denote the set of sliced vectors corresponding to the set of sliced texts D as A = (A1, A2,..., A i ,..., A n ), where A i represents the sliced vector corresponding to the sliced text D i .

[0014] Preferably, in S3, the following sub-steps are further included:

[0015] S31. When n ≥ 2, take the n sliced vectors in the set of sliced vectors A as the node vectors of n leaf nodes; calculate the distances between each pair of sliced vectors in the set of sliced vectors A. If the distances between a certain sliced vector and k1 other sliced vectors are all within the first radius, then hang the corresponding (k1 + 1) leaf nodes under the same first-layer clustering node; 1 < k1 + 1 ≤ n and k1 is a positive integer;

[0016] If the distance between a certain sliced vector and any sliced vector is greater than the first radius, then the leaf node corresponding to this sliced vector is recorded as a first-layer free node;

[0017] If there is only 1 first-layer clustering node and no first-layer free nodes, then the current first-layer clustering node is the root node, and the index tree construction of the current set of sliced vectors A is completed; otherwise, execute S32;

[0018] When n = 1, that is, there is only 1 sliced vector in the set of sliced vectors A, then the node corresponding to the current sliced vector is the root node, and the index tree construction of the current set of sliced vectors A is completed; otherwise, execute S32;

[0019] S32. After comprehensively summarizing the text content of all leaf nodes subordinate to each first-layer clustering node through a large language model, obtain the text content of the corresponding first-layer clustering node; then encode the text content of the first-layer clustering node into a corresponding node vector;

[0020] S33. Calculate the distances between the node vectors of each pair of the first-layer clustering nodes and the first-layer free nodes. If the distances between the node vectors of k2 first-layer clustering nodes and k3 first-layer free nodes are all within the second radius, then hang these k2 first-layer clustering nodes and k3 first-layer free nodes under the same second-layer clustering node; 1 < k2 + k3 ≤ n, and both k2 and k3 are positive integers. If the distances between the node vectors of a certain first-layer clustering node and any other first-layer clustering node / first-layer free node are all greater than the second radius, then this first-layer clustering node is recorded as a second-layer free node. If the distances between the node vectors of a certain first-layer free node and any other first-layer clustering node / first-layer free node are all greater than the second radius, then this first-layer free node is recorded as a second-layer free node. If there is only 1 second-layer clustering node and no second-layer free nodes, then the current second-layer clustering node is the root node, and the index tree construction of the current slice vector set A is completed; otherwise, execute S34;

[0021] S34. After comprehensively summarizing the text content of all child nodes subordinate to each second-layer clustering node through a large language model, obtain the text content of the corresponding second-layer clustering node; then encode the text content of the second-layer clustering node into a corresponding node vector;

[0022] S35. Calculate the distances between the node vectors of each pair of the second-layer clustering nodes and the second-layer free nodes. If the distances between each pair of the node vectors of k4 second-layer clustering nodes and k5 second-layer free nodes are all within the third radius, then hang these k4 second-layer clustering nodes and k5 second-layer free nodes under the same third-layer clustering node; 1 < k4 + k5 ≤ n, and both k4 and k5 are positive integers. If the distance between a certain second-layer clustering node and the node vectors of any other second-layer clustering nodes / second-layer free nodes is greater than the third radius, then this second-layer clustering node is recorded as a third-layer free node. If the distance between a certain second-layer free node and the node vectors of any other second-layer clustering nodes / second-layer free nodes is greater than the third radius, then this second-layer free node is recorded as a third-layer free node. If there is only 1 third-layer clustering node and no third-layer free nodes, then the current third-layer clustering node is the root node, and the index tree construction of the current slice vector set A is completed. Otherwise, continuously obtain the upper-layer clustering nodes and upper-layer free nodes until the current clustering node is the root node, and the index tree construction of the current slice vector set A is completed. The specific method for obtaining the upper-layer clustering nodes and upper-layer free nodes is the same as that in S34 - S35.

[0023] Preferably, the number of child nodes of any clustering node is less than M: On the premise that the distances between each pair of node vectors are all within the corresponding radius, hang the first M child nodes with the shortest distances between node vectors under the same clustering node. If the number of child nodes is less than M, then hang all child nodes under the same clustering node; M > 1 and M is a positive integer.

[0024] Preferably, the height of the index tree is within the set number of layers L max as follows: If the height of the index tree is greater than the set number of layers L max , then starting from the top layer, cut off the layers greater than L max in the current index tree to obtain m index trees.

[0025] The present invention also provides a retrieval method, including the following steps:

[0026] Step 1, obtain the user's requirement, and generate a requirement vector and a requirement keyword set;

[0027] Step 2, perform rough screening according to the requirement vector and the requirement keyword set: Calculate the first similarity score Score1 between the user's requirement and the root node of the index tree, and retain the index trees with the first similarity score Score1 above the first score value; The index tree is obtained by adopting an information storage method as described above.

[0028] Step 3: Conduct a refined screening on the initially screened index tree. Calculate the second similarity scores between the user requirements and the clustering nodes and leaf nodes in the index tree, and retain the clustering nodes and leaf nodes whose total similarity scores are above the second score value. Output the root nodes screened out initially and the clustering nodes and leaf nodes screened out through refinement as the basic information corresponding to the user requirements.

[0029] Preferably, in Step 2, it includes the following content: The first similarity score Score1 = α1 × P1 + α2 × P2, where α1 represents the first weight parameter, α2 represents the second weight parameter, 0 ≤ α1 ≤ 1 and 0 ≤ α2 ≤ 1, P1 represents the similarity score between the current requirement vector and the node vector of the current root node, and P2 represents the similarity score between the current requirement keyword set and the node text of the current root node. Retain the index trees with the first similarity score Score1 above the first score value, and at the same time, use the node content of the root nodes with the first similarity score Score1 above the first score value as the basic information.

[0030] Preferably, in Step 3, it includes the following content:

[0031] From top to bottom, calculate the second similarity score Score2 = α3 × P3 + α4 × P4 between the user requirements and the clustering nodes or leaf nodes that are not root nodes in the index tree, where α3 represents the third weight parameter, α4 represents the fourth weight parameter, 0 ≤ α3 ≤ 1 and 0 ≤ α4 ≤ 1; P3 represents the similarity score between the current requirement vector and the node vector of the current clustering node or leaf node that is not a root node; P4 represents the similarity score between the current requirement keyword set and the node vector of the current clustering node or leaf node that is not a root node.

[0032] If the second similarity score Score2 between the user requirements and the current clustering node is above the second score value, then use the node content of the current clustering node as the basic information, and then calculate the second similarity score between the user requirements and the child nodes of the current clustering node until the child nodes of the current clustering node are leaf nodes; if the second similarity score Score2 between the user requirements and the current clustering node is less than the second score value, then discard the current clustering node.

[0033] If the second similarity score Score2 between the user requirements and the current leaf node is above the second score value, then use the node content of the current leaf node as the basic information; if the second similarity score Score2 between the user requirements and the current leaf node is less than the second score value, then discard the current leaf node.

[0034] Output the node content of the root nodes screened out initially and the clustering nodes and leaf nodes screened out through refinement as the basic information corresponding to the user requirements.

[0035] Preferably, all leaf nodes with a second similarity score above the second score with respect to the user requirements are sorted in descending order of the second similarity score, and the node contents of the top q leaf nodes are taken as the basic information, where q is a positive integer.

[0036] The present invention also provides a question answering method, including the following: sending the user requirements and the corresponding basic information into a question answering model, and the question answering model outputs a corresponding answer; the basic information is obtained by using a retrieval method as described above.

[0037] The beneficial effects of the present invention are as follows:

[0038] (1) The information storage method of the present invention has good versatility. For text documents converted from newly obtained information, regardless of whether they contain a chapter structure, corresponding index trees can be formed.

[0039] (2) The information storage method of the present invention stores the newly obtained information in the form of a structured index tree. The nodes in the index tree not only retain the information of the sliced text itself, but also retain the global information and local information in a document through a unique traceability structure, that is, clustering nodes and the parent-child relationship between nodes, greatly reducing the loss of information in the document and significantly improving the accuracy of the answers given by the subsequent question answering model.

[0040] (3) The information storage method of the present invention makes full use of the chapter structure contained in the document, directly uses the headings in the current document as the text contents of different clustering nodes, and then sets the parent-child node relationship between the corresponding clustering nodes according to the hierarchical relationship between the headings. The contents of a large number of clustering nodes do not need to be obtained by comprehensive generalization by a large language model. Therefore, while ensuring the accuracy of the index tree, the time for constructing the index tree can be shortened, and the occupation of computing resources can be reduced.

[0041] (4) In the retrieval method of the present invention, since the information has been stored as an index tree with a traceability structure in the early stage, it is basically not necessary to call a large language model, and it further avoids the process of calling the large language model multiple times to extract entities and relationships and construct a knowledge graph; moreover, the method for constructing the index tree has a high degree of generality and does not require special training to extract entities and relationships of information in a specific field, and the call rate of the large language model is also significantly lower than the prior art. Therefore, the retrieval method of this embodiment greatly reduces the occupation of computing resources and storage resources.

[0042] (5) A retrieval method of the present invention can screen out several index trees from hundreds of millions of index trees in the database only by initially coarsely screening the root node, directly narrowing the screening scope of the basic information. Then, among these index trees after coarse screening, the nodes that can serve as basic information in each index tree are further determined through fine screening. Especially for clustering nodes, if the node content of the current clustering node is not sufficient to serve as basic information, there is no need to calculate whether the node content of its child nodes is sufficient to serve as basic information. We only need to further finely screen the child nodes of the clustering nodes that can serve as basic information. Therefore, the retrieval method of this embodiment can quickly, accurately, and efficiently determine the nodes that serve as basic information, improving the accuracy and response speed of the subsequent question-and-answer model's responses based on this basic information.

[0043] (6) A retrieval method of the present invention determines whether the node content of a node is sufficient to serve as basic information by calculating the similarity between the user's requirement and the node. The similarity between the user's requirement and the node is comprehensively judged by the similarity between the requirement vector and the current node vector, as well as the similarity between the set of requirement keywords and the current node text, which is more reasonable and scientific, further improving the accuracy of the nodes that serve as basic information.

[0044] (7) A retrieval method of the present invention outputs the node content of the index tree with a traceability structure as the basic information. Therefore, the output basic information is also traceable and contains global information and local information, avoiding the loss of global information and local information caused by sliced text during the previous information storage process. This also further improves the effective information corresponding to the user's requirement contained in the basic information, and further improves the accuracy rate of the subsequent question-and-answer model's responses.

[0045] (8) A question-answering method of the present invention can improve the accuracy rate and response speed of the question-and-answer model. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart of an information storage method of the present invention;

[0047] Figure 2 is a schematic diagram of the index tree before a document of the present invention is cropped;

[0048] Figure 3 is Figure 2 a schematic diagram of the cropped index tree in

[0049] Figure 4 is a schematic diagram of the index tree corresponding to the new document containing a chapter structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the technical solution of the present invention clearer and more definite, the present invention will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Any equivalent replacement and conventional reasoning of the technical features of the technical solution of the present invention by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment 1

[0052] As Figure 1 shown, it is a flowchart of an information storage method according to this embodiment, including the following steps:

[0053] S1. Convert the newly obtained information into a new document; one new piece of information corresponds to one new document;

[0054] S2. After slicing the new document, form several sliced texts;

[0055] S3. Based on all the sliced texts of one new document, construct and store the index tree of the current document; one new document corresponds to m index trees, where m is a positive integer.

[0056] In S1, the following sub-steps are further included:

[0057] S11. Compare the grabbed information with the information in the first database. If the current information does not exist in the first database, store the current information in the first database and at the same time determine that the current information is recorded as new information;

[0058] S12. Convert the new information into a corresponding new document.

[0059] In S12, if the new information contains pictures, videos or audios, use a multimodal large model to describe the content of the pictures, videos or audios, and convert the pictures, videos or audios into pure text form, so that one new piece of information is converted into one new document; there is no non-text content in the converted new document.

[0060] The multimodal large model used in this embodiment is Qwen2.5-VL-7B.

[0061] In S2, the following sub-steps are further included:

[0062] S21. Use a segmentation algorithm to slice one new document into n sliced texts. Denote the slice set of the current new document as D, then D=(D1, D2,..., D i ,..., D n ), 1≤i≤n and both i and n are positive integers, and D i represents the i-th sliced text in the current new document.

[0063] In this implementation, the segmentation algorithm used is based on SeaModel, combined with BERT encoding and an adaptive sliding window for slicing.

[0064] S22. Encode each slice text into a corresponding slice vector, and denote the set of slice vectors corresponding to the slice set D as A = (A1, A2,..., A i ,..., A n ), A i represents the slice vector corresponding to the slice text D i .

[0065] In this implementation, the BGE encoding model is used to encode the slice text into corresponding vectors.

[0066] Optionally, in S22: Encode the slice set D into a first vector set A′ = (A1′, A2′,..., A i ′,..., A n ′), A i ′ represents the vector corresponding to the slice text D i ; after reducing the dimensions of each vector, obtain the corresponding slice vector set A = (A1, A2,..., A i ,..., A n ), A i represents the reduced-dimensional vector corresponding to the vector A i ′, and is also the slice vector corresponding to the slice text D i .

[0067] In this implementation, the dimensionality reduction algorithm used is the UMAP algorithm, which reduces the dimensions of each slice vector to 30.

[0068] The specific dimension after dimensionality reduction is not a limitation of the present invention. Those skilled in the art can determine it according to actual needs (i.e., the computing power of the processor, the occupancy of computing resources, etc.). Although dimensionality reduction is accompanied by information loss in some dimensions of the vector, dimensionality reduction can greatly reduce the subsequent computational burden. Therefore, those skilled in the art need to control the dimension after dimensionality reduction and, while ensuring that the slice vectors do not lose a large amount of dimensional information, try to reduce the subsequent computational burden as much as possible.

[0069] In S3, the following sub-steps are further included:

[0070] S31. When \(n\geq2\), use the \(n\) slice vectors in the slice vector set \(A\) as the node vectors of \(n\) leaf nodes; calculate the distances between every two slice vectors in the slice vector set \(A\). If the distances between a certain slice vector and \(k1\) other slice vectors are all within the first radius, then hang the corresponding \((k1 + 1)\) leaf nodes under the same first-layer clustering node; \(0\lt k1 + 1\leq n\), \(n\geq2\) and \(k1\) is a positive integer.

[0071] If the distance between a certain slice vector and any other slice vector is greater than the first radius, then the leaf node corresponding to this slice vector is recorded as a first-layer free node.

[0072] If there is only 1 first-layer clustering node and no first-layer free nodes, then the current first-layer clustering node is the root node, and the index tree construction of the current slice vector set \(A\) is completed; otherwise, execute S32.

[0073] When \(n = 1\), that is, there is only 1 slice vector in the slice vector set \(A\), then the node corresponding to the current slice vector is the root node, and the index tree construction of the current slice vector set \(A\) is completed; otherwise, execute S32.

[0074] The Euclidean distance formula is used to calculate the distances between every two node vectors.

[0075] S32. After comprehensively summarizing the text content of all leaf nodes under each first-layer clustering node through a large language model, obtain the text content of the corresponding first-layer clustering node; then encode the text content of the first-layer clustering node into the corresponding node vector.

[0076] In S32, the text content of the leaf node is the corresponding slice text.

[0077] In this embodiment, the large language model used is Qwen2.5 - 72B.

[0078] S33. Calculate the distances between every two node vectors of the first-layer clustering nodes and the first-layer free nodes. If the distances between the node vectors of \(k2\) first-layer clustering nodes and \(k3\) first-layer free nodes are all within the second radius, then hang these \(k2\) first-layer clustering nodes and \(k3\) first-layer free nodes under the same second-layer clustering node; \(1\lt k2 + k3\leq n\), \(n\geq2\) and \(k2\), \(k3\) are both positive integers.

[0079] If the distance between a certain first-layer clustering node and any other first-layer clustering node / first-layer free node is greater than the second radius, then this first-layer clustering node is recorded as a second-layer free node.

[0080] If the distance between the node vectors of a certain first-layer free node and any other first-layer clustering node / first-layer free node is greater than the second radius, then this first-layer free node is recorded as a second-layer free node;

[0081] If there is only one second-layer clustering node and no second-layer free node, then the current second-layer clustering node is the root node, and the index tree construction of the current slice vector set A is completed; otherwise, execute S34.

[0082] S34. After comprehensively summarizing the text content of all child nodes under each second-layer clustering node through a large language model, the text content of the corresponding second-layer clustering node is obtained; then the text content of the second-layer clustering node is encoded into the corresponding node vector.

[0083] S35. Calculate the distance between the node vectors of the second-layer clustering nodes and the second-layer free nodes pairwise. If the distances between the node vectors of k4 second-layer clustering nodes and k5 second-layer free nodes are all within the third radius, then hang these k4 second-layer clustering nodes and k5 second-layer free nodes under the same third-layer clustering node; 1 < k4 + k5 ≤ n, n ≥ 2 and k4, k5 are all positive integers;

[0084] If the distance between the node vectors of a certain second-layer clustering node and any other second-layer clustering node / second-layer free node is greater than the third radius, then this second-layer clustering node is recorded as a third-layer free node;

[0085] If the distance between the node vectors of a certain second-layer free node and any other second-layer clustering node / second-layer free node is greater than the third radius, then this second-layer free node is recorded as a third-layer free node.

[0086] If there is only one third-layer clustering node and no third-layer free node, then the current third-layer clustering node is the root node, and the index tree construction of the current slice vector set A is completed; otherwise, continuously obtain the upper-layer clustering nodes and upper-layer free nodes until the current clustering node is the root node, and the index tree construction of the current slice vector set A is completed.

[0087] The specific method for obtaining the upper-layer clustering nodes and upper-layer free nodes is similar to S34 - S35, and will not be elaborated here.

[0088] Optionally, during the process of constructing the index tree, the number of child nodes of any clustering node is less than M: on the premise that the distances between the node vectors pairwise are all within the corresponding radius, hang the first M child nodes with the shortest pairwise node vector distances under the same clustering node. If the number of child nodes is less than M, then hang all child nodes under the same clustering node; M > 1 and M is a positive integer.

[0089] During the process of constructing the index tree, the number of child nodes under each clustering node is restricted, further improving the subsequent retrieval efficiency.

[0090] It should be emphasized that an index tree has only one root node. A child node is a node directly connected to its parent node. The root node has no parent node, and the leaf node has no child node. In this embodiment, the parent node of the leaf node must be a clustering node.

[0091] Optionally, the height of the index tree is within the set number of layers L max below.

[0092] If the height of the index tree is greater than the set number of layers L max , then starting from the top layer, the layers greater than L in the current index tree are trimmed downwards to obtain m new index trees. max

[0093] For ease of understanding, the following is combined with Figures 2 to 3 illustrative examples:

[0094] Such as Figure 2Shown is the index tree before document cropping. The slice vector set A contains 8 slice vectors, so the 8 leaf nodes are denoted as V0 to V7. The number of child nodes of any clustering node does not exceed 3. Then, during the process of constructing the current index tree: The distances between the slice vectors corresponding to the leaf nodes V1 to V2 are all within the first radius. Therefore, the leaf nodes V1 to V2 are hung under the same first-layer clustering node, and this clustering node is denoted as U3. Similarly, the leaf nodes V4 to V5 are hung under the same first-layer clustering node, and this clustering node is denoted as U4. The leaf nodes V6 to V7 are hung under the same first-layer clustering node, and this clustering node is denoted as U5. And the distance between the slice vector corresponding to the leaf node V1 and the other 7 slice vectors is greater than the first radius. So, the leaf node V1 is also a first-layer free node. Calculate the distances between the first-layer free node V1 and the first-layer clustering nodes U3 to U5 pairwise. Since the distance between the node vector corresponding to the first-layer free node V1 and the first-layer clustering node U3 is within the second radius, the first-layer free node V1 and the first-layer clustering node U3 are hung under the same second-layer clustering node, and this clustering node is denoted as U1. Similarly, the first-layer clustering node U4 and the first-layer clustering node U5 are hung under the same second-layer clustering node, and this clustering node is denoted as U2. And there is no second-layer free node. Since the distance between the node vectors corresponding to the second-layer clustering node U1 and the second-layer clustering node U2 is within the third radius, the second-layer clustering node U1 and the second-layer clustering node U2 are hung under the same third-layer clustering node, and this clustering node is denoted as U0. And there is no third-layer free node. Since there is only 1 third-layer clustering node and no third-layer free node, the clustering node U0 is the root node, and the index tree construction of the current slice vector set A is completed.

[0095] Figure 2 The index tree in contains 4 layers, which are the 0th layer, the 1st layer, the 2nd layer, and the 3rd layer from bottom to top. The 0th layer contains the leaf nodes V1 to V7, the 1st layer contains the leaf node V0 and the clustering nodes U3 to U5, the 2nd layer contains the clustering nodes U1 to U2, and the 3rd layer contains the clustering node U0.

[0096] The child node points to its parent node.

[0097] Set the number of layers L max = 3, then as Figure 3 shown, for the Figure 2 index tree in after cropping, two new index trees are obtained: We crop the top layer of the index tree in, that is, after cropping the clustering node U0, m = 2 new index trees are obtained. Figure 2

[0098] ​In summary, a document needs to be trimmed due to the excessive height of the index tree, resulting in more than one index tree corresponding to the document finally.

[0099] Optionally, when the new document contains a chapter structure, after hierarchically splitting the new document according to the chapter structure, construct the corresponding index tree and store it, where the text content of several clustering nodes in the index tree is the chapter title.

[0100] Specifically, it includes the following content:

[0101] A new document containing a chapter structure contains f levels of headings, where f is a positive integer, and each heading at the last level contains a sub-document without a chapter structure.

[0102] Take each heading in the current document as the text content of different clustering nodes, and then set the parent-child node relationship between the corresponding clustering nodes according to the hierarchical relationship between the headings; denote the clustering node corresponding to a certain last-level heading as X, and construct the sub-index tree of the clustering node X for the sub-document included under the last-level heading according to S2~S3 and its sub-steps. The root node of the corresponding sub-index tree is the clustering node X. At this time, the index tree corresponding to the current new document containing a chapter structure is constructed, and the current index tree is stored.

[0103] For the convenience of understanding, the following combines Figure 4 Take an example to illustrate:

[0104] A new document containing a chapter structure has two levels of headings. There is only 1 first-level heading, denoted as T01; there are two second-level headings under the first-level heading T01, denoted as T02 and T03 respectively. The sub-document N02 is included under the second-level heading T02, and the sub-document N03 is included under the second-level heading T03. Denote the three clustering nodes corresponding to the headings T01~T03 as X1~X3 respectively. Then the clustering node X2 and the clustering node X3 are the sub-nodes of the clustering node X1.

[0105] After slicing sub-document N02, four sliced texts are obtained. These four sliced texts are respectively used as the text content of leaf nodes V8 to V11, and a sub-index tree of clustering node X2 is constructed. The sub-index tree of clustering node X2 is as shown within the red dashed box. During the process of constructing the sub-index tree of clustering node X2, the distances between the sliced vectors corresponding to leaf nodes V8 to V11 are all within the first radius. However, in this embodiment, M = 3, so not all of leaf nodes V8 to V11 can be used as the child nodes of clustering node X2. The first 3 leaf nodes V9 to V11 with the shortest distances between the node vectors of leaf nodes V8 to V11 are hung under the same clustering node U6, that is, leaf nodes V9 to V11 are located at the 0th layer in the current sub-index tree, and leaf node V8 is recorded as a free node at the first layer; then leaf node V8 and clustering node U6 are used as the child nodes of clustering node X2, that is, leaf node V8 and clustering node U6 are located at the 1st layer in the current sub-index tree, and clustering node X2 is located at the 2nd layer in the current sub-index tree, which is also the root node of the current sub-index tree.

[0106] After slicing sub-document N03, four sliced texts are obtained. These four sliced texts are respectively used as the text content of leaf nodes V12 to V15, and a sub-index tree of clustering node X3 is constructed. The sub-index tree of clustering node X3 is as shown within the green dashed box. During the process of constructing the sub-index tree of clustering node X3, the distances between the sliced vectors corresponding to leaf nodes V12 to V13 are all within the first radius, and the distances between the sliced vectors corresponding to leaf nodes V14 to V15 are all within the first radius. Therefore, leaf nodes V12 and V13 are hung under the same clustering node U7, leaf nodes V14 and V15 are hung under the same clustering node U8, and clustering nodes U7 and U8 are hung under clustering node X3; that is, in the current sub-index tree, leaf nodes V12 to V15 are all located at the 0th layer in the current sub-index tree, clustering nodes U7 and U8 are located at the 1st layer in the current sub-index tree, and clustering node X3 is located at the 2nd layer in the current sub-index tree, which is also the root node of the current sub-index tree.

[0107] For the current new document containing chapter structures, clustering nodes X2 and X3 are located at the 2nd layer of the entire index tree, clustering node X1 is located at the 3rd layer of the entire index tree, and the entire index tree contains four layers in total.

[0108] The specific content of pruning the index tree corresponding to the new document containing chapter structures is the same as above, and will not be elaborated here.

[0109] When the index tree construction is completed, each node generates corresponding node content, which includes: the text content of the node, the node vector, the parent node of the current node, the level of the current node in the index tree, and the child nodes of the current node. Among them, the corresponding node vector is obtained by encoding the text content of the node; if the current node is the root node, the parent node of the current node is a null value; if the current node is a leaf node, the child nodes of the current node are null values.

[0110] In the prior art, the collected information is converted into a text document and then sliced, and then the obtained several sliced texts are directly stored in the database, which will not only cause the loss of some global information and local information in the document, but also reduce the accuracy of the final reply made by the subsequent question-and-answer model; compared with the prior art, the information storage method of this embodiment stores the newly obtained information in the form of a structured index tree. The nodes in the index tree not only retain the information of the sliced text itself, but also retain the global information and local information in a document through a unique traceability structure, that is, clustering nodes and the parent-child relationship between nodes, greatly reducing the loss of information in the document and significantly improving the accuracy of the reply made by the subsequent question-and-answer model.

[0111] An information storage method of this embodiment has good versatility. For a text document converted from newly obtained information, whether it contains a chapter structure or not, a corresponding index tree can be formed.

[0112] Titles at different levels are highly accurate summaries of the content they contain, so the chapter structure contained in the document is the natural local information and global information of the document. An information storage method of this embodiment makes full use of the chapter structure contained in the document, directly uses the titles in the current document as the text content of different clustering nodes, and then sets the parent-child node relationship between the corresponding clustering nodes according to the hierarchical relationship between the titles. A large amount of clustering node content does not need to be obtained after being comprehensively summarized by a large language model. Therefore, while ensuring the accuracy of the index tree, it can also shorten the time for constructing the index tree and reduce the occupation of computing resources (using a large language model will occupy a large amount of computing resources).

[0113] Embodiment 2

[0114] This embodiment also provides a retrieval method, including the following steps:

[0115] Step 1, obtain user requirements, and generate a requirement vector and a set of requirement keywords;

[0116] Step 2, according to the requirement vector and the set of requirement keywords, perform rough screening: calculate the first similarity score Score1 between the user requirement and the root node of the index tree, and retain the index trees with the first similarity score Score1 above the first score value;

[0117] Step 3: Perform fine screening on the roughly screened index tree: Calculate the second similarity score between the user requirements and the clustering nodes and leaf nodes in the index tree, and retain the clustering nodes and leaf nodes whose total similarity score is above the second score value; Output the root nodes screened out roughly and the clustering nodes and leaf nodes screened out finely as the basic information corresponding to the user requirements.

[0118] The index tree is obtained by adopting an information storage method described in Embodiment 1.

[0119] In Step 1, use a text embedding model to convert the text requirements into vector form, i.e., requirement vectors; Extract keywords from the text requirements to obtain one or more keywords, and form the requirement keyword set of the current user.

[0120] In this embodiment, the TextRank algorithm is adopted for keyword extraction.

[0121] In Step 2, the first similarity score Score1 = α1×P1 + α2×P2, where α1 represents the first weight parameter, α2 represents the second weight parameter, 0≤α1≤1 and 0≤α2≤1, P1 represents the similarity score between the current requirement vector and the node vector of the current root node, and P2 represents the similarity score between the current requirement keyword set and the node text of the current root node; Retain the index tree whose first similarity score Score1 is above the first score value, and at the same time, take the node content of the root node whose first similarity score Score1 is above the first score value as the basic information.

[0122] In this embodiment, the BM25 retrieval algorithm is used to calculate the similarity score between the requirement keyword set node texts, and the cosine similarity algorithm is used to calculate the similarity score between the requirement vector and the node vector.

[0123] In Step 3, the following content is also included:

[0124] From the upper layer to the lower layer, calculate the second similarity score Score2 = α3×P3 + α4×P4 between the user requirements and the non-root clustering nodes or leaf nodes in the index tree, where α3 represents the third weight parameter, α4 represents the fourth weight parameter, 0≤α3≤1 and 0≤α4≤1; P3 represents the similarity score between the current requirement vector and the node vector of the current non-root clustering node or leaf node; P4 represents the similarity score between the current requirement keyword set and the node vector of the current non-root clustering node or leaf node;

[0125] If the second similarity score Score2 between the user requirement and the current clustering node is above the second score value, the node content of the current clustering node is used as the basic information. Then, the second similarity score between the user requirement and the child nodes of the current clustering node is calculated until the child nodes of the current clustering node are leaf nodes. If the second similarity score Score2 between the user requirement and the current clustering node is less than the second score value, the current clustering node is discarded.

[0126] If the second similarity score Score2 between the user requirement and the current leaf node is above the second score value, the node content of the current leaf node is used as the basic information. If the second similarity score Score2 between the user requirement and the current leaf node is less than the second score value, the current leaf node is discarded.

[0127] The node content of the root nodes roughly screened out, the clustering nodes and leaf nodes finely screened out is output as the basic information corresponding to the user requirement.

[0128] Optionally, during the fine screening process, all leaf nodes with a second similarity score above the second score value with respect to the user requirement are sorted in descending order according to their corresponding second similarity scores, and the node content of the top q leaf nodes is taken as the basic information, where q is a positive integer.

[0129] In the prior art retrieval method, the collected piece of information is converted into a text document and then sliced, and the obtained several sliced texts are directly stored in the database. When retrieving in the sliced texts in the database, it is necessary to call the large language model multiple times to extract entities and relationships from the sliced texts, and construct a knowledge graph according to the entities and relationships (a large language model also needs to be called a lot during the process of constructing the knowledge graph), and then traverse the knowledge graph to retrieve the corresponding entities and relationships as the basic information for output. And in order to ensure the accuracy of the large language model in extracting entities and relationships, it is also necessary to train the large language model according to the type and field of the sliced texts. Therefore, the retrieval method of the prior art not only takes a long time, but also has too high an occupancy rate of storage computing resources and storage resources.

[0130] In the retrieval method of this embodiment, since the information has been stored as an index tree with a traceability structure in the early stage, it is basically unnecessary to call the large language model, thus avoiding the process of calling the large language model multiple times to extract entities and relationships and construct a knowledge graph. Moreover, the method of constructing the index tree has a high degree of generality, does not require special training to extract entities and relationships in specific fields of information, and the call rate of the large language model is also significantly lower than that of the prior art. Therefore, the retrieval method of this embodiment greatly reduces the occupancy of computing resources and storage resources.

[0131] The retrieval method of this embodiment only needs to first perform a rough screening of the root node, and then several index trees can be screened out from hundreds of millions of index trees in the database, directly narrowing the screening range of the basic information. Then, among these index trees after rough screening, the nodes that can be used as basic information in each index tree are further determined through fine screening. Especially for clustering nodes, if the node content of the current clustering node is not sufficient to be used as basic information, there is no need to calculate whether the node content of its child nodes is sufficient to be used as basic information. We only need to further fine-screen the child nodes of the clustering nodes that can be used as basic information. Therefore, the retrieval method of this embodiment can quickly, accurately and efficiently determine the nodes that can be used as basic information, improving the accuracy and response speed of the subsequent Q&A model's responses based on this basic information.

[0132] The retrieval method of this embodiment determines whether the node content of a node is sufficient to be used as basic information by calculating the similarity between the user's demand and the node. The similarity between the user's demand and the node is comprehensively judged by the similarity between the demand vector and the current node vector, as well as the similarity between the set of demand keywords and the current node text, which is more reasonable and scientific, and further improves the accuracy of the nodes used as basic information.

[0133] Since the node content of the index tree with a traceability structure is used as the basic information for output in the retrieval method of this embodiment, the output basic information is also traceable and contains global information and local information, avoiding the loss of global information and local information caused by sliced text during the previous information storage process. This also further improves the effective information corresponding to the user's demand contained in the basic information, and further improves the accuracy rate of the subsequent Q&A model's responses.

[0134] Embodiment 3

[0135] This embodiment is a question answering method, including the following steps: sending the user's demand and the corresponding basic information into the Q&A model, and the Q&A model outputs the corresponding response.

[0136] The basic information is obtained by using a retrieval method as described in Embodiment 2.

[0137] In this embodiment, the Q&A model is the RAG model.

[0138] After verification by technical personnel, compared with directly using the RAG model to answer based on the sliced text of the initial information (the initial information refers to the information directly grabbed from various platforms) in the prior art, for the same 10 user demands based on the same 10,000 initial information, the response speed is increased by an average of 27%, and the accuracy rate of the response is increased by an average of 15%.

[0139] A problem answering method according to this embodiment can improve the answering accuracy and speed of the question answering model.

[0140] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies. It should also be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Each component or step in the embodiments of the present invention can be decomposed and / or recombined, and these decompositions and / or recombinations should be regarded as equivalent solutions of the present application and should all fall within the protection scope of the present invention.

Claims

1. An information storage method, characterized in that, Including the following steps: S1. Convert the newly obtained information into a new document; one new piece of information corresponds to one new document. S2. Slice the new document to form a number of sliced texts. S3. Based on all the sliced texts of one new document, construct an index tree for the current document and store it; one new document corresponds to m index trees, where m is a positive integer.

2. The information storage method according to claim 1, wherein In S2, the following sub-steps are further included: S21. Use the segmentation algorithm to split a new document into n sliced texts. Denote the set of slices of the current new document as D, then D = (D1, D2,..., D i ,..., D n ), where 1 ≤ i ≤ n and both i and n are positive integers, and D i represents the i-th sliced text in the current new document; S22, encode each slice text into a corresponding slice vector, and denote the slice vector set corresponding to the slice set D as A = (A1, A2,..., A i ,..., A n ), where A i represents the slice vector corresponding to the slice text D i .

3. The information storage method according to claim 2, wherein In S3, the following sub-steps are further included: S31. When n≥2, use the n sliced vectors in the sliced vector set A as the node vectors of n leaf nodes; calculate the distances between each pair of sliced vectors in the sliced vector set A. If the distances between a certain sliced vector and k1 other sliced vectors are all within the first radius, then hang the corresponding (k1 + 1) leaf nodes under the same first-layer clustering node; 1<k1 + 1≤n and k1 is a positive integer. If the distance between a certain sliced vector and any other sliced vector is greater than the first radius, then the leaf node corresponding to this sliced vector is marked as a first-layer free node. If there is only 1 first-layer clustering node and no first-layer free nodes, then the current first-layer clustering node is the root node, and the index tree for the current sliced vector set A is constructed; otherwise, execute S32. When n = 1, that is, there is only 1 sliced vector in the sliced vector set A, then the node corresponding to the current sliced vector is the root node, and the index tree for the current sliced vector set A is constructed; otherwise, execute S32. S32. After comprehensively summarizing the text contents of all the leaf nodes under each first-layer clustering node through a large language model, obtain the text content corresponding to the first-layer clustering node; then encode the text content of the first-layer clustering node into a corresponding node vector. S33. Calculate the distances between the node vectors of each pair of first-layer clustering nodes and first-layer free nodes. If the distances between the node vectors of k2 first-layer clustering nodes and k3 first-layer free nodes are all within the second radius, then hang these k2 first-layer clustering nodes and k3 first-layer free nodes under the same second-layer clustering node; 1<k2 + k3≤n, and both k2 and k3 are positive integers. If the distance between a certain first-layer clustering node and the node vectors of any other first-layer clustering node / first-layer free node is greater than the second radius, then this first-layer clustering node is marked as a second-layer free node. If the distance between a certain first-layer free node and the node vectors of any other first-layer clustering node / first-layer free node is greater than the second radius, then this first-layer free node is marked as a second-layer free node. If there is only 1 second-layer clustering node and no second-layer free nodes, then the current second-layer clustering node is the root node, and the index tree for the current sliced vector set A is constructed; otherwise, execute S34. S34. After comprehensively summarizing the text contents of all the child nodes under each second-layer clustering node through a large language model, obtain the text content corresponding to the second-layer clustering node; then encode the text content of the second-layer clustering node into a corresponding node vector. S35. Calculate the distances between the node vectors of each pair of the second-layer clustering nodes and the second-layer free nodes. If the distances between the node vectors of each pair of the k4 second-layer clustering nodes and the k5 second-layer free nodes are all within the third radius, then hang these k4 second-layer clustering nodes and k5 second-layer free nodes under the same third-layer clustering node; 1 < k4 + k5 ≤ n, and both k4 and k5 are positive integers. If the distances between the node vectors of a certain second-layer clustering node and any other second-layer clustering node / second-layer free node are all greater than the third radius, then this second-layer clustering node is recorded as a third-layer free node. If the distances between the node vectors of a certain second-layer free node and any other second-layer clustering node / second-layer free node are all greater than the third radius, then this second-layer free node is recorded as a third-layer free node. If there is only 1 third-layer clustering node and no third-layer free nodes, then the current third-layer clustering node is the root node, and the index tree for the current slice vector set A is constructed; otherwise, continuously obtain the upper-layer clustering nodes and upper-layer free nodes until the current clustering node is the root node, at which time the index tree for the current slice vector set A is constructed; the specific method for obtaining the upper-layer clustering nodes and upper-layer free nodes is the same as that in S34 - S35.

4. The information storage method according to claim 3, characterized in that, The number of child nodes of any clustering node is less than M: On the premise that the distances between the node vectors of each pair are all within the corresponding radius, hang the first M child nodes with the shortest distances between the node vectors under the same clustering node. If the number of child nodes is less than M, then hang all child nodes under the same clustering node; M > 1 and M is a positive integer.

5. The information storage method according to claim 3, wherein The height of the index tree is at the set number of layers L max The following: If the height of the index tree is greater than the set number of layers L max , then starting from the top layer, the layers greater than L in the current index tree are trimmed downwards to obtain m index trees. max ​ 6. A retrieval method, characterized in that, It includes the following steps: Step 1. Obtain the user requirements, and generate a requirement vector and a requirement keyword set. Step 2. According to the requirement vector and the requirement keyword set, conduct a rough screening: Calculate the first similarity score Score1 between the user requirements and the root node of the index tree, and retain the index trees with the first similarity score Score1 above the first score value; the index tree is obtained by adopting an information storage method described in any one of claims 3 - 5. Step 3. Conduct a fine screening on the index trees obtained from the rough screening: Calculate the second similarity score between the user requirements and the clustering nodes and leaf nodes in the index tree, and retain the clustering nodes and leaf nodes with the total similarity score above the second score value. Output the root node obtained from the rough screening and the clustering nodes and leaf nodes obtained from the fine screening as the basic information corresponding to the user requirements.

7. A retrieval method according to claim 6, wherein In step 2, it includes the following content: The first similarity score Score1 = α1 × P1 + α2 × P2, where α1 represents the first weight parameter, α2 represents the second weight parameter, 0 ≤ α1 ≤ 1 and 0 ≤ α2 ≤ 1, P1 represents the similarity score between the current requirement vector and the node vector of the current root node, and P2 represents the similarity score between the current requirement keyword set and the node text of the current root node; retain the index trees with the first similarity score Score1 above the first score value, and at the same time use the node content of the root node with the first similarity score Score1 above the first score value as the basic information.

8. A retrieval method according to claim 6, characterized in that In step 3, the following is included: From top to bottom, calculate the second similarity score Score2 = α3×P3 + α4×P4 between the user requirement and the clustering node or leaf node of the non-root node in the index tree, where α3 represents the third weight parameter, α4 represents the fourth weight parameter, 0 ≤ α3 ≤ 1 and 0 ≤ α4 ≤ 1; P3 represents the similarity score between the current requirement vector and the node vector of the current clustering node or leaf node of the non-root node; P4 represents the similarity score between the current requirement keyword set and the node vector of the current clustering node or leaf node of the non-root node. If the second similarity score Score2 between the user requirement and the current clustering node is above the second score value, then take the node content of the current clustering node as the basic information, and then calculate the second similarity score between the user requirement and the child nodes of the current clustering node until the child nodes of the current clustering node are leaf nodes; if the second similarity score Score2 between the user requirement and the current clustering node is less than the second score value, then discard the current clustering node. If the second similarity score Score2 between the user requirement and the current leaf node is above the second score value, then take the node content of the current leaf node as the basic information; if the second similarity score Score2 between the user requirement and the current leaf node is less than the second score value, then discard the current leaf node. Take the node content of the root node roughly screened and the clustering nodes and leaf nodes finely screened as the basic information corresponding to the user requirement for output.

9. The retrieval method according to claim 8, wherein: Arrange all the leaf nodes with the second similarity score above the second score value with respect to the user requirement in descending order of the second similarity score, and take the node content of the top q leaf nodes as the basic information, where q is a positive integer.

10. A question answering method, characterized in that, The following is included: send the user requirement and the corresponding basic information into the Q&A model, and the Q&A model outputs the corresponding reply; the basic information is obtained by using a retrieval method described in any one of claims 6-9.

Citation Information

Patent Citations

  • Concept hierarchy establishing method based on product review document set

    CN103761264A

  • Face image clustering method and device, electronic equipment and storage medium

    CN115359541A

  • All-weather RAG intelligent agent automatic newspaper design method and device and electronic equipment

    CN119476247A

  • Text processing method, device, equipment, storage medium and program product

    CN119760127A

  • Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search

    US20240386015A1