Structured retrieval method and device based on rhetorical structure analysis and large language model tree

Through rhetorical structure analysis and large language model tree structured retrieval methods, a grammatical binary tree index structure is constructed and the BERT embedding model is optimized, which solves the problem of insufficient professional knowledge of large language models, improves retrieval accuracy and generation effects, and reduces retrieval costs.

CN120353803BActive Publication Date: 2025-09-09HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510855165.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-09
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In existing technologies, large language models lack the ability to store knowledge in specialized fields, leading to misunderstandings and hallucinations. Fixed-length block segmentation methods ignore document semantics and logical structures, resulting in high retrieval costs and an inability to effectively utilize knowledge from long documents.

Method used

It adopts rhetorical structure parsing and large language model tree structured retrieval methods, optimizes the BERT embedding model through grammatical binary tree construction and comparative learning, constructs a grammatical tree index structure, and performs semantic encoding and retrieval.

Benefits of technology

It improves the adaptability and generation effect of large language models in professional fields, respects the semantics and logical structure of documents, reduces retrieval costs, and is suitable for large-scale data applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353803B_ABST
    Figure CN120353803B_ABST
Patent Text Reader

Abstract

The present invention provides a structured retrieval method and device based on rhetorical structure parsing and a large language model tree, relating to the field of natural language processing technology. The method comprises: using a rhetorical structure parsing model to perform structural parsing on an original long document and construct a grammatical binary tree; using a large language model to semantically summarize the text of the nodes of the grammatical binary tree, and obtaining a summary text of the grammatical binary tree through an iterative loop mechanism; constructing a contrastive learning training sample, and optimizing a pre-trained BERT embedding model using a contrastive learning framework to obtain an optimized BERT embedding model; constructing a retrieval index structure using FAISS-based vector retrieval technology; inputting a user's query into the optimized BERT embedding model for encoding to obtain a query vector representation; calculating cosine similarity based on the query vector representation and the retrieval index structure, and obtaining retrieval results based on the similarity score. The present invention can improve the effectiveness of tree-structured retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a structured retrieval method and device based on rhetorical structure analysis and a large language model tree. Background Art

[0002] With the rapid development of large language model technology in recent years, large models have been widely used due to their powerful natural language understanding and generation capabilities, such as knowledge question answering and copywriting. However, when faced with different professional fields, the knowledge stored in large model parameters is often insufficient, which can lead to misunderstandings and generation illusions. To address this problem, retrieval-enhanced generation technology extracts knowledge from relevant professional documents based on the given query and provides it as context to the large model, significantly improving its domain adaptability and generation performance at a low cost. Due to the limited length of the large model's context window and its susceptibility to interference from irrelevant information, the accuracy of the retrieval results directly affects the effectiveness of retrieval-enhanced generation.

[0003] Current retrieval techniques for professional documents typically first divide the document into several text blocks. Different retrieval methods are then used to identify the blocks most relevant to the generated content as contextual knowledge. Block segmentation methods typically pre-set a suitable length, based on which the document is divided into consecutive blocks. For block retrieval, dense vector retrieval methods are commonly used, converting the text blocks and the query into vector representations and determining relevance based on vector similarity. Furthermore, methods using large models to read documents and directly generate key evidence have also gained some application.

[0004] Existing methods primarily include dense vector retrieval (DV) and generative model-based methods. DV retrieval involves two steps: chunking and retrieval. The chunking step involves segmenting long documents into chunks of equal length. In the retrieval step, semantic representation vectors for the chunks are calculated and stored using a pre-trained embedding model, such as the BERT embedding model. The BERT embedding model is then used to calculate the query vector representation. Vector similarity calculation methods, such as cosine similarity, measure the similarity between the query semantic representation vector and the representation vectors of different chunks, and select the one with the highest similarity as the relevant text. However, this method is highly dependent on the effectiveness of text chunking, and the optimal chunk length varies for different documents. Fixed-length chunking methods ignore the semantics and logical structure of the document itself, mechanically truncating semantically related paragraphs and hindering language model understanding. Furthermore, small chunks provide only local knowledge and fail to provide a comprehensive understanding and grasp of document knowledge. Generative model-based methods, on the other hand, do not require chunking. Instead, they receive the entire or partial document and, after understanding it, directly provide the query-related text. This approach is unaffected by chunking, and the rich context facilitates a more comprehensive understanding and utilization of document content, often resulting in better retrieval results. However, this approach is also limited by the length of the context and is extremely expensive to use. Each new query requires the generative model to reread the document, making it unsuitable for large quantities of long documents. Summary of the Invention

[0005] To address the technical issues of existing technologies, such as context length limitations, high costs, and fixed-length chunking methods that ignore the semantics and logical structure of the document itself, the present invention provides a structured retrieval method and device based on rhetorical structure analysis and a large language model tree. The technical solution is as follows:

[0006] On the one hand, a structured retrieval method based on rhetorical structure parsing and a large language model tree is provided. The method is implemented by a structured retrieval device based on rhetorical structure parsing and a large language model tree, and the method includes:

[0007] S1. Obtain the original long document; use the rhetorical structure parsing model to perform structural analysis on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at the paragraph level; use the pre-trained large language model to perform semantic induction on the paragraph-level subtrees layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; use the root node text splicing of each paragraph-level subtree as the upper-level abstract sub-document structure, and perform rhetorical structure analysis on the sub-documents to obtain the grammatical binary tree between paragraphs, and use the pre-trained large language model to perform semantic induction on the grammatical binary tree between paragraphs layer by layer; through the iterative loop mechanism, continuously iterate until the grammatical binary tree structure at the original document level is obtained; merge the grammatical binary trees at all levels to obtain the complete grammatical binary tree of the original long document;

[0008] S2. Based on the tree structure of the grammar binary tree and the pre-annotated evidence information, locate the positive sample nodes related to the pre-annotated query; sample negative samples from the remaining nodes in the grammar tree, and obtain negative sample nodes by calculating the sampling probability of negative sample nodes and positive sample nodes; use the pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample nodes and the text corresponding to the positive sample nodes, and obtain the semantic vector representation corresponding to the negative sample nodes and the semantic vector representation corresponding to the positive sample nodes; construct contrastive learning training samples based on the positive sample nodes, the semantic vector representation corresponding to the positive sample nodes, the negative sample nodes, and the semantic vector representation corresponding to the negative sample nodes;

[0009] S3. Based on the contrastive learning training samples, the contrastive learning framework is used to optimize the pre-trained BERT embedding model through the constructed contrastive learning loss function to obtain the optimized BERT embedding model;

[0010] S4. Input the text corresponding to all nodes in the grammatical binary tree into the optimized BERT embedding model for semantic encoding to obtain text vector representations corresponding to all nodes; construct a retrieval index structure based on the FAISS vector retrieval technology according to the text vector representation; input the obtained user query into the optimized BERT embedding model for semantic encoding to obtain a query vector representation; calculate the cosine similarity between the query vector representation and the node representation in the retrieval index structure to obtain a similarity score; sort the similarity scores from high to low, and select the tree node texts corresponding to the k vectors with the highest scores as the retrieval results.

[0011] Optionally, the process of using the pre-trained large language model in S1 to perform semantic induction on each paragraph-level subtree layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree is expressed by the following formula (1):

[0012] (1)

[0013] in, is the left subtree text; is the right subtree text; R is the logical relationship prediction, Summarize text for semantics; Indicates summarization using a large language model.

[0014] Optionally, the process of using the pre-trained BERT embedding model in S2 to perform semantic vector mapping on the text corresponding to the negative sample node and the text corresponding to the positive sample node is expressed by the following formula (2):

[0015] (2)

[0016] in, represents the encoder of the BERT embedding model; Indicates the added semantic start tag, Represents semantic vector representation, using The encoding output of the corresponding position is marked as a node representation; Represents the text in a node.

[0017] Optionally, the process of obtaining negative sample nodes by calculating the sampling probabilities of negative sample nodes and positive sample nodes in S2 is expressed by the following formula (3):

[0018] (3)

[0019] Among them, min() means taking the minimum value of positive and negative distances, and exp means exponential operation. The weight to control the scale of the distance value; represents the distance between all positive nodes and all negative nodes; Represents a negative node ; represents the set of all positive nodes; Represents a negative node The probability of being sampled as a negative example; represents the set of all negative nodes; Represents the distance between all positive nodes and negative nodes.

[0020] Optionally, the contrastive learning loss function is expressed by the following formula (4):

[0021] (4)

[0022] in, represents contrastive learning loss; Embedding vector representation for query; is the embedding vector representation of the positive node; is the embedding vector representation of the negative node; is the dot product similarity calculation, is the temperature hyperparameter, set to 0.01.

[0023] Optionally, the step S4 constructs a search index structure based on the text vector representation using a FAISS vector search technology, including:

[0024] Based on the text vector representations corresponding to all nodes, a structured index is performed on the nodes of each document syntax tree, and a vector index table is constructed, including the node embedding vector, the document to which the node belongs, the logical relationship label, and the hierarchical position.

[0025] According to the constructed vector index table, a retrieval index structure is constructed based on the FAISS vector retrieval technology.

[0026] Optionally, the process of calculating the cosine similarity between the query vector representation and the node representation in the retrieval index structure is represented by the following formula (5):

[0027] (5)

[0028] in, Represents the similarity score between the query vector representation and the node representation in the retrieval index structure; represents the query vector representation; Represents a vector norm; Represents the dot product operation; Represents a node representation in a retrieved index structure.

[0029] On the other hand, a structured retrieval device based on rhetorical structure parsing and a large language model tree is provided. The device is applied to a structured retrieval method based on rhetorical structure parsing and a large language model tree. The device includes:

[0030] The structural parsing unit is used to obtain the original long document; the rhetorical structure parsing model is used to perform structural parsing on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at the paragraph level; the pre-trained large language model is used to perform semantic induction on the paragraph-level subtrees layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; the root node text of each paragraph-level subtree is concatenated as the upper-level abstract sub-document structure, and the rhetorical structure of the sub-document is performed on the sub-document to obtain the grammatical binary tree between paragraphs; the pre-trained large language model is used to perform semantic induction on the grammatical binary tree between paragraphs layer by layer; through the iterative loop mechanism, it is continuously iterated until the grammatical binary tree structure at the original document level is obtained; the grammatical binary trees at all levels are merged to obtain the complete grammatical binary tree of the original long document;

[0031] A construction unit is used to locate positive sample nodes related to the pre-annotated query based on the tree structure of the grammar binary tree and pre-annotated evidence information; sample negative samples from the remaining nodes in the grammar tree, and obtain negative sample nodes by calculating the sampling probabilities of negative sample nodes and positive sample nodes; use a pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample node and the text corresponding to the positive sample node to obtain the semantic vector representation corresponding to the negative sample node and the semantic vector representation corresponding to the positive sample node; construct comparative learning training samples based on the positive sample node, the semantic vector representation corresponding to the positive sample node, the negative sample node, and the semantic vector representation corresponding to the negative sample node;

[0032] An optimization unit is used to optimize the pre-trained BERT embedding model based on the contrastive learning training samples, using a contrastive learning framework and a constructed contrastive learning loss function to obtain an optimized BERT embedding model;

[0033] The acquisition unit is used to input the text corresponding to all nodes in the grammatical binary tree into the optimized BERT embedding model for semantic encoding to obtain the text vector representations corresponding to all nodes; based on the text vector representations, a retrieval index structure is constructed using the FAISS vector retrieval technology; the obtained user query is input into the optimized BERT embedding model for semantic encoding to obtain a query vector representation; the query vector representation and the node representation in the retrieval index structure are subjected to cosine similarity calculation to obtain a similarity score; the similarity scores are sorted from high to low, and the tree node texts corresponding to the k vectors with the highest scores are selected as the retrieval results.

[0034] Optionally, a pre-trained large language model is used to perform semantic induction on each paragraph-level subtree layer by layer, and the process of obtaining the summary text of the intermediate nodes and root nodes of each subtree is expressed by the following formula (1):

[0035] (1)

[0036] in, is the left subtree text; is the right subtree text; R is the logical relationship prediction, Summarize text for semantics; Indicates summarization using a large language model.

[0037] Optionally, the process of using the pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample node and the text corresponding to the positive sample node is expressed by the following formula (2):

[0038] (2)

[0039] in, represents the encoder of the BERT embedding model; Indicates the added semantic start tag, Represents semantic vector representation, using The encoding output of the corresponding position is marked as a node representation; Represents the text in a node.

[0040] Optionally, the process of obtaining negative sample nodes by calculating the sampling probabilities of negative sample nodes and positive sample nodes is expressed by the following formula (3):

[0041] (3)

[0042] Among them, min() means taking the minimum value of positive and negative distances, and exp means exponential operation. The weight to control the scale of the distance value; represents the distance between all positive nodes and all negative nodes; Represents a negative node ; represents the set of all positive nodes; Represents a negative node The probability of being sampled as a negative example; represents the set of all negative nodes; Represents the distance between all positive nodes and negative nodes.

[0043] Optionally, the contrastive learning loss function is expressed by the following formula (4):

[0044] (4)

[0045] in, represents contrastive learning loss; Embedding vector representation for query; is the embedding vector representation of the positive node; is the embedding vector representation of the negative node; is the dot product similarity calculation, is the temperature hyperparameter, set to 0.01.

[0046] Optionally, constructing a search index structure based on the text vector representation by using a FAISS vector search technology includes:

[0047] Based on the text vector representations corresponding to all nodes, a structured index is performed on the nodes of each document syntax tree, and a vector index table is constructed, including the node embedding vector, the document to which the node belongs, the logical relationship label, and the hierarchical position.

[0048] According to the constructed vector index table, a retrieval index structure is constructed based on the FAISS vector retrieval technology.

[0049] Optionally, the process of calculating the cosine similarity between the query vector representation and the node representation in the retrieval index structure is represented by the following formula (5):

[0050] (5)

[0051] in, Represents the similarity score between the query vector representation and the node representation in the retrieval index structure; represents the query vector representation; Represents a vector norm; Represents the dot product operation; Represents a node representation in a retrieved index structure.

[0052] On the other hand, a structured retrieval device based on rhetorical structure parsing and a large language model tree is provided, and the structured retrieval device based on rhetorical structure parsing and a large language model tree includes: a processor; a memory, wherein computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned structured retrieval methods based on rhetorical structure parsing and a large language model tree is implemented.

[0053] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned methods based on rhetorical structure parsing and large language model tree structured retrieval methods.

[0054] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0055] The embodiment of the present invention first obtains the original long document; uses the rhetorical structure parsing model to perform structural analysis on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at each paragraph level; uses the pre-trained large language model to perform semantic induction on each paragraph level subtree layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; uses the root node text of each paragraph level subtree as the upper layer abstract sub-document structure, and performs rhetorical structure analysis on the sub-document to obtain the grammatical binary tree between paragraphs; uses the pre-trained large language model to perform semantic induction on the grammatical binary tree between paragraphs layer by layer The grammar binary tree structure of the original document level is gradually iterated through the iterative loop mechanism; the grammar binary trees at all levels are merged to obtain the complete grammar binary tree of the original long document; secondly, according to the tree structure of the grammar binary tree and the pre-annotated evidence information, the positive sample nodes related to the pre-annotated query are located; negative samples are sampled from the remaining nodes in the grammar tree, and the negative sample nodes are obtained by calculating the sampling probability of the negative sample and the positive sample nodes; the pre-trained BERT embedding model is used to embed the text corresponding to the negative sample node and the text corresponding to the positive sample node. The text is mapped to semantic vectors to obtain the semantic vector representations corresponding to the negative sample nodes and the semantic vector representations corresponding to the positive sample nodes; a contrastive learning training sample is constructed based on the positive sample nodes, the semantic vector representations corresponding to the positive sample nodes, the negative sample nodes and the semantic vector representations corresponding to the negative sample nodes; based on the contrastive learning training sample, the contrastive learning framework is adopted to optimize the pre-trained BERT embedding model through the constructed contrastive learning loss function to obtain the optimized BERT embedding model; the text corresponding to all nodes in the grammatical binary tree is input into the optimized BERT embedding model for semantic encoding to obtain the text vector representations corresponding to all nodes; based on the text vector representation, a retrieval index structure is constructed using the FAISS vector retrieval technology; finally, the obtained user query is input into the optimized BERT embedding model for semantic encoding to obtain the query vector representation; the cosine similarity between the query vector representation and the node representation in the retrieval index structure is calculated to obtain the similarity score; the similarity scores are sorted from high to low, and the tree node texts corresponding to the k vectors with the highest scores are selected as the retrieval results.

[0056] The embodiment of the present invention adopts document syntax and logical structure analysis to obtain a syntax tree, replacing the text blocks obtained by fixed-length segmentation, so that the segmentation and parsing process of long documents can fully respect the semantics and logical structure of the document. Each subtree of the syntax tree is a whole obtained by logical aggregation of text units rather than mechanical division, which is more conducive to the understanding and use of retrieval and generation models. The embodiment of the present invention adopts a large language model to perform hierarchical summary of the syntax tree, thereby introducing the understanding and summarization capabilities of the large language model, so that the upper nodes of the syntax tree can obtain a broader field of view and summarize the text containing the macroscopic understanding of the document. The embodiment of the present invention designs tree structure-aware node sampling and comparative learning, so that the embedding model optimizes the semantic representation of the syntax tree structure and improves the effect of tree structure retrieval. The syntax tree construction, large model summary and embedding representation of the embodiment of the present invention only need to be performed once, and can be reused for any subsequent retrieval task. The retrieval cost is at the same order of magnitude as traditional dense vector retrieval, which is suitable for application on large-scale data. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0058] Figure 1 This is a flow chart of a structured retrieval method based on rhetorical structure analysis and a large language model tree provided by an embodiment of the present invention;

[0059] Figure 2 This is a flowchart of an implementation process based on rhetorical structure analysis provided by an embodiment of the present invention;

[0060] Figure 3 This is a general flow chart of a structured retrieval method based on rhetorical structure analysis and a large language model tree provided by an embodiment of the present invention;

[0061] Figure 4 This is a schematic diagram of a simple example of a document syntax tree obtained by parsing a rhetorical structure provided by an embodiment of the present invention;

[0062] Figure 5 This is a block diagram of a structured retrieval device based on rhetorical structure analysis and large language model tree provided by an embodiment of the present invention;

[0063] Figure 6 This is a schematic diagram of the structure of a structured retrieval device based on rhetorical structure analysis and a large language model tree provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0065] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0066] In the embodiments of the present invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same. The terms "of," "corresponding," and "corresponding" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same.

[0067] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0068] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0069] The embodiment of the present invention provides a method for structured retrieval based on rhetorical structure analysis and a large language model tree. The method can be implemented by a device for structured retrieval based on rhetorical structure analysis and a large language model tree. The device for structured retrieval based on rhetorical structure analysis and a large language model tree can be a terminal or a server. Figure 1 The flowchart of the structured retrieval method based on rhetorical structure analysis and large language model tree is shown. The processing flow of the method may include the following steps:

[0070] S1. Obtain the original long document; use the rhetorical structure parsing model to perform structural analysis on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at the paragraph level; use the pre-trained large language model to perform semantic induction on the paragraph-level subtrees layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; use the root node text splicing of each paragraph-level subtree as the upper-level abstract sub-document structure, and perform rhetorical structure analysis on the sub-documents to obtain the grammatical binary trees between paragraphs, and use the pre-trained large language model to perform semantic induction on the grammatical binary trees between paragraphs layer by layer; through the iterative loop mechanism, continuously iterate until the grammatical binary tree structure at the original document level is obtained; merge the grammatical binary trees at all levels to obtain the complete grammatical binary tree of the original long document.

[0071] Among them, such as Figure 2 The figure shows a flow chart of an implementation process based on rhetorical structure analysis provided by an embodiment of the present invention.

[0072] The syntax tree includes leaf nodes and intermediate nodes; the leaf nodes are sentence-level semantic units, and the intermediate nodes represent the logical relationships between corresponding substructures, including: causal relationships, parallel relationships, concession relationships, and contrast relationships.

[0073] Among them, such as Figure 3 The figure shows a flow chart of a structured retrieval method based on rhetorical structure parsing and a large language model tree provided by an embodiment of the present invention; in a feasible implementation manner, obtaining the content of the original document includes: "Current large language models show strong generalization capabilities when processing general tasks. However, they still have obvious shortcomings in acquiring professional domain knowledge. To make up for this deficiency, researchers have begun to explore methods of introducing domain knowledge." The original document is subjected to rhetorical structure parsing, and a grammatical binary tree is constructed; the grammatical binary tree is divided into two layers including transitions and reasons; the content of the transition is: "Current large language models show strong generalization capabilities when processing general tasks"; the content of the reasons includes: "However, they still have obvious shortcomings in acquiring professional domain knowledge; to make up for this deficiency, researchers have begun to explore methods of introducing domain knowledge."

[0074] Optionally, S1 uses a pre-trained large language model to perform semantic induction on each paragraph-level subtree layer by layer, and the process of obtaining the summary text of the intermediate nodes and root nodes of each subtree is expressed by the following formula (1):

[0075] (1)

[0076] in, is the left subtree text; is the right subtree text; R is the logical relationship prediction, Summarize text for semantics; Indicates summarization using a large language model.

[0077] S2. Based on the tree structure of the grammar binary tree and the pre-annotated evidence information, locate the positive sample nodes related to the pre-annotated query; sample negative samples from the remaining nodes in the grammar tree, and obtain the negative sample nodes by calculating the sampling probability of negative and positive sample nodes; use the pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample nodes and the text corresponding to the positive sample nodes, and obtain the semantic vector representation corresponding to the negative sample nodes and the semantic vector representation corresponding to the positive sample nodes; construct comparative learning training samples based on the positive sample nodes, the semantic vector representation corresponding to the positive sample nodes, the negative sample nodes, and the semantic vector representation corresponding to the negative sample nodes.

[0078] Optionally, S2 uses a pre-trained BERT embedding model to map the semantic vectors of the text corresponding to the negative sample node and the text corresponding to the positive sample node using the following formula (2):

[0079] (2)

[0080] in, represents the encoder of the BERT embedding model; Indicates the added semantic start tag, Represents semantic vector representation, using The encoding output of the corresponding position is marked as a node representation; Represents the text in a node.

[0081] Optionally, the process of obtaining negative sample nodes by calculating the sampling probabilities of negative sample nodes and positive sample nodes is expressed by the following formula (3):

[0082] (3)

[0083] Among them, min() means taking the minimum value of positive and negative distances, and exp means exponential operation. The weight to control the scale of the distance value; represents the distance between all positive nodes and all negative nodes; Represents a negative node ; represents the set of all positive nodes; Represents a negative node The probability of being sampled as a negative example; represents the set of all negative nodes; Represents the distance between all positive nodes and negative nodes.

[0084] Among them, the sampling probability can ensure that nodes with farther semantics and structures are more likely to become negative sample nodes, so that the embedding distance of the representation model is related to the distance of the syntax tree nodes, making it more suitable for tree structure retrieval.

[0085] S3. Based on the contrastive learning training samples, the contrastive learning framework is used to optimize the pre-trained BERT embedding model through the constructed contrastive learning loss function to obtain the optimized BERT embedding model;

[0086] Among them, the contrastive learning framework is used to fine-tune the pre-trained BERT embedding model, and the optimization goal is to make the query The distance between the positive sample node and the negative sample node is minimized, while the distance between the negative sample node and the negative sample node is increased. During the optimization, a number of negative nodes are randomly sampled according to the negative sample sampling probability. Triplet training optimizes representations.

[0087] Optionally, the contrast loss function is expressed by the following formula (4):

[0088] (4)

[0089] in, represents contrastive learning loss; Embedding vector representation for query; is the embedding vector representation of the positive node; is the embedding vector representation of the negative node; is the dot product similarity calculation, is the temperature hyperparameter, set to 0.01.

[0090] S4. Input the text corresponding to all nodes in the grammatical binary tree into the optimized BERT embedding model for semantic encoding to obtain the text vector representation corresponding to all nodes; based on the text vector representation, use the FAISS vector retrieval technology to build a retrieval index structure; input the obtained user query into the optimized BERT embedding model for semantic encoding to obtain the query vector representation; calculate the cosine similarity between the query vector representation and the node representation in the retrieval index structure to obtain the similarity score; sort the similarity scores from high to low, and select the tree node texts corresponding to the k vectors with the highest scores as the retrieval results.

[0091] Optionally, S4 constructs a search index structure based on the text vector representation using FAISS vector search technology, including:

[0092] Based on the text vector representations corresponding to all nodes, we perform a structured index on the nodes of each lesson document's syntax tree, and construct a vector index table that includes the node embedding vector, the document to which the node belongs, the logical relationship label, and the hierarchical position.

[0093] According to the constructed vector index table, a retrieval index structure is constructed based on the FAISS vector retrieval technology.

[0094] The FAISS vector retrieval technology is a technology known to those skilled in the art and will not be further elaborated in the embodiments of the present invention.

[0095] Optionally, the process of calculating the cosine similarity between the query vector representation and the node representation of the retrieval index structure value is expressed by the following formula (5):

[0096] (5)

[0097] in, Represents the similarity score between the query vector representation and the node representation in the retrieval index structure; represents the query vector representation; Represents a vector norm; Represents the dot product operation; Represents a node representation in a retrieved index structure.

[0098] Among them, the retrieval results can not only return the most relevant sentence-level semantic units to users, but also trace back the subtree where the most relevant sentence-level semantic units are located and obtain the mid-level semantic summary nodes, realizing the result display of two granularities: the underlying original text details and the high-level context summary.

[0099] in, Figure 4 It is a general flow chart of the tree-structured retrieval method based on rhetorical structure analysis and large language model provided by an embodiment of the present invention; in a feasible implementation manner, the embodiment of the present invention constructs three modules including: a document recursive parsing and syntax tree construction module, a tree structure-aware contrastive learning representation model optimization module and a multi-granularity tree-structured retrieval module; wherein, the document recursive parsing and syntax tree construction module is used to parse the logic and syntax structure of a long document to form a document syntax tree; the tree structure-aware contrastive learning representation model optimization module is used to introduce tree structure information into the representation model to optimize it in a direction that is more suitable for tree retrieval; the multi-granularity tree-structured retrieval module is used to construct a completed document syntax analysis binary tree and an optimized embedded representation model, and to construct a structured information retrieval mechanism for semantic understanding. This module supports multi-granularity and hierarchical query response capabilities, and can quickly locate the substructure content that best matches the query semantics in the tree structure according to the query semantics, thereby improving the accuracy and interpretability of the retrieval.

[0100] In one feasible implementation, the rhetorical structure parsing method is based on a transformation-based system that constructs a grammatical binary tree through a series of well-defined actions. The system can capture local and global logical structure relationships while maintaining computational efficiency. The transformation system maintains two main data structures: a stack , used to store partially built trees; a queue , used to store unprocessed sentences and build them step by step from bottom to top.

[0101] The conversion system gradually builds a structure tree through the following three basic operations:

[0102] (1) “Shift” operation: When new content is needed for processing, a sentence is moved from the queue to the stack.

[0103] (2) “Reduce” operation: combine the two adjacent subtrees at the top of the stack into a new subtree and identify the rhetorical structural relationship between them.

[0104] (3) “Pop root” operation: When the complete grammar binary tree is successfully constructed, the whole process ends.

[0105] Each state of the system is represented by , from the initial Start, where is the initial state, S is the set of all sentences. At this time, all sentences are in the queue and eventually end at ,in is the terminal state, and T is the final grammar binary tree.

[0106] In one possible implementation, the conversion system follows a deterministic process guided by a neural scoring model including:

[0107] (1) Initialization: ;

[0108] (2) When Not empty or , do the following:

[0109] (a) If and If it is not empty, perform the "shift" operation and move the next sentence from Move to ;

[0110] (b) If If it is empty, the "reduce" operation is performed to merge the two subtrees at the top of the stack;

[0111] (c) Otherwise, use the neural scoring model based on the current and The state of determines whether to perform a "shift" or "reduce" operation.

[0112] (3) Return stack The only tree T left in .

[0113] Among them, the neural scoring model can consider the top three subtrees in the stack and the next sentence in the queue Perform action scoring and next action selection. The design takes the following factors into consideration:

[0114] (a) and is an immediate candidate for the next potential "reduce" operation;

[0115] (b) Provides important contextual information about recently constructed structures;

[0116] (c) Useful for determining whether new content should be introduced via a "shift" operation.

[0117] Among them, when the scoring model calculates the next action, it first calculates the three subtrees and sentences Calculate its representation vector: For each leaf node in the subtree , or a sentence , and calculate its representation vector using the following formula (6):

[0118] (6)

[0119] in is the leaf node text, is the node representation vector, For sentences The representation vector, Pre-trained XLNET model; for any intermediate node of the three subtrees , and its representation vector is calculated using the following formula (7):

[0120] (7)

[0121] in Intermediate node The representation vector of is the set of child nodes of the intermediate node t, is the representation vector of each child node.

[0122] The scoring model calculates the action score based on the above representation vector, which is expressed by the following formula (8):

[0123] (8)

[0124] Where a is a specific action, To calculate the score of action a, represents the first learnable parameter matrix; is the second learnable parameter matrix, Represents a subtree The root node of represents a vector; Represents a subtree The root node of represents a vector; Represents a subtree The root node represents the vector, For sentences The representation vector, is a vector concatenation operation. According to the action score, the selection probability of each action is calculated using the following formula (9):

[0125] (9)

[0126] in is the predicted probability of selecting action a in state c, is the natural exponential operation, is the score of action a, A is the set of all actions, Any optional action in action set A The scoring model calculates the predicted probability of all actions and uses the action with the highest probability as the next action in rhetorical structure analysis.

[0127] The embodiment of the present invention first obtains the original long document; uses the rhetorical structure parsing model to perform structural analysis on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at each paragraph level; uses the pre-trained large language model to perform semantic induction on each paragraph level subtree layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; uses the root node text of each paragraph level subtree as the upper layer abstract sub-document structure, and performs rhetorical structure analysis on the sub-document to obtain the grammatical binary tree between paragraphs; uses the pre-trained large language model to perform semantic induction on the grammatical binary tree between paragraphs layer by layer The method is to perform semantic induction; through the iterative loop mechanism, it is continuously iterated until the grammatical binary tree structure of the original document level is obtained; the grammatical binary trees at all levels are merged to obtain the complete grammatical binary tree of the original long document; secondly, according to the tree structure of the grammatical binary tree and the pre-annotated evidence information, the positive sample nodes related to the pre-annotated query are located; negative samples are sampled from the remaining nodes in the grammatical tree, and the negative sample nodes are obtained by calculating the sampling probability of the negative sample and the positive sample nodes; the pre-trained BERT embedding model is used to embed the text corresponding to the negative sample node and the text corresponding to the positive sample node. The text is mapped to a semantic vector to obtain the semantic vector representation corresponding to the negative sample node and the semantic vector representation corresponding to the positive sample node; a contrastive learning training sample is constructed based on the positive sample node, the semantic vector representation corresponding to the positive sample node, the negative sample node and the semantic vector representation corresponding to the negative sample node; based on the contrastive learning training sample, the contrastive learning framework is adopted to optimize the pre-trained BERT embedding model through the constructed contrastive learning loss function to obtain the optimized BERT embedding model; the text corresponding to all nodes in the grammatical binary tree is input into the optimized BERT embedding model for semantic encoding to obtain the text vector representation corresponding to all nodes; based on the text vector representation, a retrieval index structure is constructed using the FAISS vector retrieval technology; finally, the obtained user query is input into the optimized BERT embedding model for semantic encoding to obtain the query vector representation; the cosine similarity between the query vector representation and the node representation in the retrieval index structure is calculated to obtain the similarity score; the similarity scores are sorted from high to low, and the tree node texts corresponding to the k vectors with the highest scores are selected as the retrieval results.

[0128] The embodiment of the present invention adopts document syntax and logical structure analysis to obtain a syntax tree, replacing the text blocks obtained by fixed-length segmentation, so that the segmentation and parsing process of long documents can fully respect the semantics and logical structure of the document. Each subtree of the syntax tree is a whole obtained by logical aggregation of text units rather than mechanical division, which is more conducive to the understanding and use of retrieval and generation models. The embodiment of the present invention adopts a large language model to perform hierarchical summary of the syntax tree, thereby introducing the understanding and summarization capabilities of the large language model, so that the upper nodes of the syntax tree can obtain a broader field of view and summarize the text containing the macroscopic understanding of the document. The embodiment of the present invention designs tree structure-aware node sampling and comparative learning, so that the embedding model optimizes the semantic representation of the syntax tree structure and improves the effect of tree structure retrieval. The syntax tree construction, large model summary and embedding representation of the embodiment of the present invention only need to be performed once, and can be reused for any subsequent retrieval task. The retrieval cost is at the same order of magnitude as traditional dense vector retrieval, which is suitable for application on large-scale data.

[0129] Figure 5 This is a block diagram of a structured search device based on rhetorical structure analysis and a large language model tree according to an exemplary embodiment. The device is used for a structured search method based on rhetorical structure analysis and a large language model tree. Figure 5 The device includes a structure analysis unit 510, a construction unit 520, an optimization unit 530, and an acquisition unit 540.

[0130] The structural parsing unit 510 is used to obtain the original long document; use the rhetorical structure parsing model to perform structural parsing on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at each paragraph level; use the pre-trained large language model to perform semantic induction on each paragraph level subtree layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; use the root node text of each paragraph level subtree to splice as the upper-level abstract sub-document structure, and perform rhetorical structure parsing on the sub-document to obtain the grammatical binary tree between paragraphs; use the pre-trained large language model to perform semantic induction on the grammatical binary tree between paragraphs layer by layer; through an iterative loop mechanism, continuously iterate until the grammatical binary tree structure at the original document level is obtained; merge the grammatical binary trees at all levels to obtain the complete grammatical binary tree of the original long document;

[0131] A construction unit 520 is configured to locate positive example nodes related to a pre-annotated query based on the tree structure of the grammar binary tree and pre-annotated evidence information; sample negative examples from the remaining nodes in the grammar tree, and obtain negative example nodes by calculating the sampling probabilities of negative examples and positive example nodes; use a pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative example nodes and the text corresponding to the positive example nodes to obtain semantic vector representations corresponding to the negative example nodes and semantic vector representations corresponding to the positive example nodes; and construct comparative learning training samples based on the positive example nodes, the semantic vector representations corresponding to the positive example nodes, the negative example nodes, and the semantic vector representations corresponding to the negative example nodes.

[0132] An optimization unit 530 is configured to optimize the pre-trained BERT embedding model using a contrastive learning framework and a constructed contrastive learning loss function based on the contrastive learning training sample to obtain an optimized BERT embedding model;

[0133] The acquisition unit 540 is used to input the text corresponding to all nodes in the grammatical binary tree into the optimized BERT embedding model for semantic encoding to obtain text vector representations corresponding to all nodes; based on the text vector representation, a retrieval index structure is constructed using FAISS vector retrieval technology; the obtained user query is input into the optimized BERT embedding model for semantic encoding to obtain a query vector representation; cosine similarity is calculated between the query vector representation and the node representation in the retrieval index structure to obtain a similarity score; the similarity scores are sorted from high to low, and the tree node texts corresponding to the k vectors with the highest scores are selected as retrieval results.

[0134] Optionally, a pre-trained large language model is used to perform semantic induction on each paragraph-level subtree layer by layer, and the process of obtaining the summary text of the intermediate nodes and root nodes of each subtree is expressed by the following formula (1):

[0135] (1)

[0136] in, is the left subtree text; is the right subtree text; R is the logical relationship prediction, Summarize text for semantics; Indicates summarization using a large language model.

[0137] Optionally, the process of using the pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample node and the text corresponding to the positive sample node is expressed by the following formula (2):

[0138] (2)

[0139] in, represents the encoder of the BERT embedding model; Indicates the added semantic start tag, Represents semantic vector representation, using The encoding output of the corresponding position is marked as a node representation; Represents the text in a node.

[0140] Optionally, the process of obtaining negative sample nodes by calculating the sampling probabilities of negative sample nodes and positive sample nodes is expressed by the following formula (3):

[0141] (3)

[0142] Among them, min() means taking the minimum value of positive and negative distances, and exp means exponential operation. The weight to control the scale of the distance value; represents the distance between all positive nodes and all negative nodes; Represents a negative node ; represents the set of all positive nodes; Represents a negative node The probability of being sampled as a negative example; represents the set of all negative nodes; Represents the distance between all positive nodes and negative nodes.

[0143] Optionally, the contrastive learning loss function is expressed by the following formula (4):

[0144] (4)

[0145] in, represents contrastive learning loss; Embedding vector representation for query; is the embedding vector representation of the positive node; is the embedding vector representation of the negative node; is the dot product similarity calculation, is the temperature hyperparameter, set to 0.01.

[0146] Optionally, constructing a search index structure based on the text vector representation by using a FAISS vector search technology includes:

[0147] Based on the text vector representations corresponding to all nodes, a structured index is performed on the nodes of each document syntax tree, and a vector index table is constructed, including the node embedding vector, the document to which the node belongs, the logical relationship label, and the hierarchical position.

[0148] According to the constructed vector index table, a retrieval index structure is constructed based on the FAISS vector retrieval technology.

[0149] Optionally, the process of calculating the cosine similarity between the query vector representation and the node representation in the retrieval index structure is represented by the following formula (5):

[0150] (5)

[0151] in, Represents the similarity score between the query vector representation and the node representation in the retrieval index structure; represents the query vector representation; Represents a vector norm; Represents the dot product operation; Represents a node representation in a retrieved index structure.

[0152] The embodiment of the present invention first obtains the original long document; uses the rhetorical structure parsing model to perform structural analysis on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at each paragraph level; uses the pre-trained large language model to perform semantic induction on each paragraph level subtree layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; uses the root node text of each paragraph level subtree as the upper layer abstract sub-document structure, and performs rhetorical structure analysis on the sub-document to obtain the grammatical binary tree between paragraphs; uses the pre-trained large language model to perform semantic induction on the grammatical binary tree between paragraphs layer by layer The method is to perform semantic induction; through the iterative loop mechanism, it is continuously iterated until the grammatical binary tree structure of the original document level is obtained; the grammatical binary trees at all levels are merged to obtain the complete grammatical binary tree of the original long document; secondly, according to the tree structure of the grammatical binary tree and the pre-annotated evidence information, the positive sample nodes related to the pre-annotated query are located; negative samples are sampled from the remaining nodes in the grammatical tree, and the negative sample nodes are obtained by calculating the sampling probability of the negative sample and the positive sample nodes; the pre-trained BERT embedding model is used to embed the text corresponding to the negative sample node and the text corresponding to the positive sample node. The text is mapped to a semantic vector to obtain the semantic vector representation corresponding to the negative sample node and the semantic vector representation corresponding to the positive sample node; a contrastive learning training sample is constructed based on the positive sample node, the semantic vector representation corresponding to the positive sample node, the negative sample node and the semantic vector representation corresponding to the negative sample node; based on the contrastive learning training sample, the contrastive learning framework is adopted to optimize the pre-trained BERT embedding model through the constructed contrastive learning loss function to obtain the optimized BERT embedding model; the text corresponding to all nodes in the grammatical binary tree is input into the optimized BERT embedding model for semantic encoding to obtain the text vector representation corresponding to all nodes; based on the text vector representation, a retrieval index structure is constructed using the FAISS vector retrieval technology; finally, the obtained user query is input into the optimized BERT embedding model for semantic encoding to obtain the query vector representation; the cosine similarity between the query vector representation and the node representation in the retrieval index structure is calculated to obtain the similarity score; the similarity scores are sorted from high to low, and the tree node texts corresponding to the k vectors with the highest scores are selected as the retrieval results.

[0153] The embodiment of the present invention adopts document syntax and logical structure analysis to obtain a syntax tree, replacing the text blocks obtained by fixed-length segmentation, so that the segmentation and parsing process of long documents can fully respect the semantics and logical structure of the document. Each subtree of the syntax tree is a whole obtained by logical aggregation of text units rather than mechanical division, which is more conducive to the understanding and use of retrieval and generation models. The embodiment of the present invention adopts a large language model to perform hierarchical summary of the syntax tree, thereby introducing the understanding and summarization capabilities of the large language model, so that the upper nodes of the syntax tree can obtain a broader field of view and summarize the text containing the macroscopic understanding of the document. The embodiment of the present invention designs tree structure-aware node sampling and comparative learning, so that the embedding model optimizes the semantic representation of the syntax tree structure and improves the effect of tree structure retrieval. The syntax tree construction, large model summary and embedding representation of the embodiment of the present invention only need to be performed once, and can be reused for any subsequent retrieval task. The retrieval cost is at the same order of magnitude as traditional dense vector retrieval, which is suitable for application on large-scale data.

[0154] Figure 6 This is a schematic diagram of a structured retrieval device based on rhetorical structure analysis and large language model tree provided by an embodiment of the present invention. Figure 6 As shown, the structured retrieval device based on rhetorical structure analysis and large language model tree can include the above Figure 5 The illustrated apparatus for retrieval based on rhetorical structure analysis and large language model tree structure. Optionally, the apparatus for retrieval based on rhetorical structure analysis and large language model tree structure 610 may include a first processor 2001 .

[0155] Optionally, the structured retrieval device 610 based on rhetorical structure parsing and large language model tree may further include a memory 2002 and a transceiver 2003 .

[0156] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0157] The following combination Figure 6 The components of the structured retrieval device 610 based on rhetorical structure analysis and large language model tree are introduced in detail:

[0158] The first processor 2001 is the control center of the rhetorical structure parsing and large language model tree structured retrieval device 610, and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), or application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs) or one or more field programmable gate arrays (FPGAs).

[0159] Optionally, the first processor 2001 can perform various functions of the rhetorical structure parsing and large language model tree structured retrieval device 610 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0160] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 are shown in FIG.

[0161] In a specific implementation, as an embodiment, the structured retrieval device 610 based on rhetorical structure analysis and large language model tree may also include multiple processors, such as Figure 6 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0162] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0163] Alternatively, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and access the first processor 2001 through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0164] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0165] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0166] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or can exist independently and can be retrieved through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0167] It should be noted that Figure 6 The structure of the structured retrieval device 610 based on rhetorical structure parsing and large language model tree shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0168] In addition, the technical effects of the structured retrieval device 610 based on rhetorical structure analysis and large language model tree can refer to the technical effects of the structured retrieval method based on rhetorical structure analysis and large language model tree described in the above method embodiment, and will not be repeated here.

[0169] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0170] It should also be understood that the memory in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0171] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0172] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0173] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0174] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0175] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0176] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0177] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0178] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0179] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0180] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.

[0181] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A structured retrieval method based on rhetorical structure analysis and large language model tree, characterized in that: The method comprises: S1. Obtain the original long document; use the rhetorical structure parsing model to perform structural analysis on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at the paragraph level; use the pre-trained large language model to perform semantic induction on the paragraph-level subtrees layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; use the root node text splicing of each paragraph-level subtree as the upper-level abstract sub-document structure, perform rhetorical structure analysis on the sub-documents, and obtain the grammatical binary tree between paragraphs; use the pre-trained large language model to perform semantic induction on the grammatical binary tree between paragraphs layer by layer; through the iterative loop mechanism, continuously iterate until the grammatical binary tree structure at the original document level is obtained; merge the grammatical binary trees at all levels to obtain the complete grammatical binary tree of the original long document; S2. Based on the tree structure of the grammar binary tree and the pre-annotated evidence information, locate the positive sample nodes related to the pre-annotated query; sample negative samples from the remaining nodes in the grammar tree, and obtain negative sample nodes by calculating the sampling probability of negative sample nodes and positive sample nodes; use the pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample nodes and the text corresponding to the positive sample nodes, and obtain the semantic vector representation corresponding to the negative sample nodes and the semantic vector representation corresponding to the positive sample nodes; construct contrastive learning training samples based on the positive sample nodes, the semantic vector representation corresponding to the positive sample nodes, the negative sample nodes, and the semantic vector representation corresponding to the negative sample nodes; S3. Based on the contrastive learning training samples, the contrastive learning framework is used to optimize the pre-trained BERT embedding model through the constructed contrastive learning loss function to obtain the optimized BERT embedding model; S4. Input the text corresponding to all nodes in the grammatical binary tree into the optimized BERT embedding model for semantic encoding to obtain text vector representations corresponding to all nodes; construct a retrieval index structure based on the FAISS vector retrieval technology according to the text vector representation; input the obtained user query into the optimized BERT embedding model for semantic encoding to obtain a query vector representation; calculate the cosine similarity between the query vector representation and the node representation in the retrieval index structure to obtain a similarity score; sort the similarity scores from high to low, and select the tree node texts corresponding to the k vectors with the highest scores as the retrieval results.

2. The structured retrieval method based on rhetorical structure analysis and large language model tree according to claim 1 is characterized in that: The process of using the pre-trained large language model in S1 to perform semantic induction on each paragraph-level subtree layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree is expressed by the following formula (1): (1) in, is the left subtree text; is the right subtree text; R is the logical relationship prediction, Summarize text for semantics; Indicates summarization using a large language model.

3. The structured retrieval method based on rhetorical structure analysis and large language model tree according to claim 1 is characterized in that: The process of using the pre-trained BERT embedding model in S2 to map the semantic vectors of the text corresponding to the negative sample node and the text corresponding to the positive sample node is expressed by the following formula (2): (2) in, represents the encoder of the BERT embedding model; Indicates the added semantic start tag, Represents semantic vector representation, using The encoding output of the corresponding position is marked as a node representation; Represents the text in a node.

4. The structured retrieval method based on rhetorical structure analysis and large language model tree according to claim 1 is characterized in that: The process of obtaining negative sample nodes by calculating the sampling probabilities of negative sample nodes and positive sample nodes in S2 is expressed by the following formula (3): (3) Among them, min() means taking the minimum value of positive and negative distances, and exp means exponential operation. The weight to control the scale of the distance value; represents the distance between all positive nodes and all negative nodes; Represents a negative node ; represents the set of all positive nodes; Represents a negative node The probability of being sampled as a negative example; represents the set of all negative nodes; Represents the distance between all positive nodes and negative nodes.

5. The structured retrieval method based on rhetorical structure analysis and large language model tree according to claim 1 is characterized in that: The contrastive learning loss function is expressed by the following formula (4): (4) in, represents contrastive learning loss; Embedding vector representation for query; is the embedding vector representation of the positive node; is the embedding vector representation of the negative node; is the dot product similarity calculation, is the temperature hyperparameter, set to 0.

01.

6. The structured retrieval method based on rhetorical structure analysis and large language model tree according to claim 1 is characterized in that: The step S4 constructs a search index structure based on the text vector representation using the FAISS vector search technology, including: Based on the text vector representations corresponding to all nodes, a structured index is performed on the nodes of each document syntax tree, and a vector index table is constructed, including the node embedding vector, the document to which the node belongs, the logical relationship label, and the hierarchical position. According to the constructed vector index table, a retrieval index structure is constructed based on the FAISS vector retrieval technology.

7. The structured retrieval method based on rhetorical structure analysis and large language model tree according to claim 1 is characterized in that: The process of calculating the cosine similarity between the query vector representation and the node representation in the retrieval index structure is expressed by the following formula (5): (5) in, Represents the similarity score between the query vector representation and the node representation in the retrieval index structure; represents the query vector representation; Represents a vector norm; Represents the dot product operation; Represents a node representation in a retrieval index structure.

8. A structured search device based on rhetorical structure analysis and a large language model tree, wherein the structured search device based on rhetorical structure analysis and a large language model tree is used to implement the structured search method based on rhetorical structure analysis and a large language model tree as described in any one of claims 1 to 7, characterized in that: The device comprises: The structural parsing unit is used to obtain the original long document; the rhetorical structure parsing model is used to perform structural parsing on the paragraphs of the original long document to obtain the grammatical binary tree subtrees at the paragraph level; the pre-trained large language model is used to perform semantic induction on the paragraph-level subtrees layer by layer to obtain the summary text of the intermediate nodes and root nodes of each subtree; the root node text of each paragraph-level subtree is concatenated as the upper-level abstract sub-document structure, and the rhetorical structure of the sub-document is performed on the sub-document to obtain the grammatical binary tree between paragraphs; the pre-trained large language model is used to perform semantic induction on the grammatical binary tree between paragraphs layer by layer; through the iterative loop mechanism, it is continuously iterated until the grammatical binary tree structure at the original document level is obtained; the grammatical binary trees at all levels are merged to obtain the complete grammatical binary tree of the original long document; A construction unit is used to locate positive sample nodes related to the pre-annotated query based on the tree structure of the grammar binary tree and pre-annotated evidence information; sample negative samples from the remaining nodes in the grammar tree, and obtain negative sample nodes by calculating the sampling probability of negative sample nodes and positive sample nodes; use the pre-trained BERT embedding model to perform semantic vector mapping on the text corresponding to the negative sample node and the text corresponding to the positive sample node to obtain the semantic vector representation corresponding to the negative sample node and the semantic vector representation corresponding to the positive sample node; construct contrastive learning training samples based on the positive sample node, the semantic vector representation corresponding to the positive sample node, the negative sample node, and the semantic vector representation corresponding to the negative sample node; An optimization unit is used to optimize the pre-trained BERT embedding model based on the contrastive learning training samples, using a contrastive learning framework and a constructed contrastive learning loss function to obtain an optimized BERT embedding model; The acquisition unit is used to input the text corresponding to all nodes in the grammatical binary tree into the optimized BERT embedding model for semantic encoding to obtain the text vector representations corresponding to all nodes; based on the text vector representations, a retrieval index structure is constructed using the FAISS vector retrieval technology; the obtained user query is input into the optimized BERT embedding model for semantic encoding to obtain a query vector representation; the query vector representation and the node representation in the retrieval index structure are subjected to cosine similarity calculation to obtain a similarity score; the similarity scores are sorted from high to low, and the tree node texts corresponding to the k vectors with the highest scores are selected as the retrieval results.

9. A structured retrieval device based on rhetorical structure analysis and large language model tree, characterized in that: The structured retrieval device based on rhetorical structure analysis and large language model tree includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for generating text data and computer equipment

    CN112328734A

  • Text retrieval enhancement generation method and system, terminal equipment and readable storage medium

    CN120011512A