Intelligent literature retrieval system and method, electronic equipment and storage medium
By introducing cited literature extraction, text blocking and vectorization processing, hierarchical clustering and search tree construction and search adjustment modules into the literature search system, the existing system's insufficient clustering in reference digestion, cited knowledge utilization and dimensionality reduction clustering is solved, and more efficient and accurate literature search is achieved.
Patent Information
- Application Number
- CN202411941760.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
The existing literature search system deals with the problem of insufficient reference digestion, utilization of cited knowledge and dimensionality reduction clustering in the text, resulting in insufficient search results.
By introducing cited literature extraction module, text chunking module, text vectorization and dimension reduction module, hierarchical clustering and search tree construction module and search adjustment module, we realize the extraction and utilization of cited sentences in the literature, text chunking and vectorization processing, hierarchical clustering and search tree construction, and adjust the height of the search tree according to user needs.
It improves the accuracy and efficiency of literature search, can understand semantic relationships in text more accurately, make full use of the condensed knowledge in cited sentences, and enhances the relevance and accuracy of the search results.
Smart Images

Figure CN119938807A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to an intelligent document retrieval system, method, electronic equipment and storage medium. Background Art
[0002] With the rapid development of information technology, a large amount of academic literature and technical data are constantly being generated. Existing literature retrieval systems mainly rely on keyword matching and retrieval methods based on vector space models (such as TF-IDF, Word2Vec, etc.). However, these traditional methods cannot handle the semantic relationships in the text well, resulting in inaccurate retrieval results, especially when dealing with literature with complex citation relationships.
[0003] At present, some document retrieval systems have begun to introduce deep learning technology, using neural network models to generate vector representations of documents to improve retrieval results, but the following problems still exist:
[0004] 1. Traditional text segmentation and vectorization processing methods cannot effectively handle the problem of reference resolution in text, especially the resolution of pronouns (such as "it", "he", "she", etc.) and demonstrative pronouns (such as "this", "that"), resulting in the inability to accurately identify and associate the core information in the text during the retrieval process.
[0005] 2. Most of the current retrieval systems fail to effectively utilize the cited sentences in the literature as external knowledge, do not fully utilize the citation information in the literature, and ignore the scientific research results and ideas condensed in the cited literature.
[0006] 3. The accuracy of dimensionality reduction and clustering strategies after document vectorization is low, resulting in inaccurate classification of retrieval results.
[0007] A Chinese patent with publication number CN118861088A discloses a short text query expansion and enhanced retrieval method based on a knowledge base hierarchical tree structure. The invention captures the macro and detail level information of the knowledge base document by constructing a knowledge base hierarchical tree structure, and implements different levels of retrieval strategies in the two-stage expansion of short text queries and result retrieval according to contextual requirements. At the macro level, the short text query is retrieved on the basis of a generalized expansion of the knowledge base hierarchical tree to implement a secondary specific expansion of the query, so as to achieve the alignment of the user query with the semantic space of the knowledge base; at the detail level, the original text fragments in the underlying leaf nodes are accurately retrieved through the summary layer located during the secondary expansion, thereby generating highly relevant query responses. This technical solution focuses on expanding queries and retrieval by constructing a hierarchical structure, without considering how to effectively use the reference information in the document as an external knowledge source, especially how to extract and use the reference sentences in the document. Moreover, the retrieval tree constructed by the invention is relatively static, and does not involve how to flexibly adjust the height and complexity of the tree structure under different retrieval requirements. Summary of the invention
[0008] The present invention provides an intelligent document retrieval system, method, electronic device and storage medium, aiming to solve the problems of insufficient reference resolution, utilization of reference knowledge and dimensionality reduction clustering in existing document retrieval systems.
[0009] The technical solution of the present invention is as follows:
[0010] On the one hand, the present invention provides an intelligent document retrieval system, including: a reference document extraction module, a text segmentation module, a text vectorization module, a hierarchical clustering and retrieval tree construction module and a retrieval adjustment module.
[0011] The cited literature extraction module is used to obtain cited articles as external knowledge through crawler tools for subsequent retrieval.
[0012] The text segmentation module is used to perform a coreference resolution operation on external knowledge through a large language model, and to segment the text of the external knowledge that has undergone coreference resolution into sub-units.
[0013] The text vectorization and dimensionality reduction module is used to vectorize the text of each sub-unit and reduce the dimensionality of the obtained text vector.
[0014] The hierarchical clustering and retrieval tree construction module is used to perform hierarchical clustering on the text vector after dimensionality reduction, build a retrieval tree based on the hierarchical clustering results, and merge the information of the nodes in the retrieval tree through a large language model.
[0015] The retrieval adjustment module is used to adjust the height of the retrieval tree according to the different requirements of the user input query during the retrieval process, perform matching retrieval based on the adjusted retrieval tree, and output the retrieval results.
[0016] Preferably, the cited articles obtained as external knowledge by crawler tools in the cited document extraction module are specifically:
[0017] By identifying the database to which each existing document belongs, constructing a URL to access the detailed page of each document, parsing the HTML structure in the page, and extracting the document title information.
[0018] Using the document title information, construct a URL to download the PDF file of the article that cites the document, and download the PDF file using a crawler tool.
[0019] Use the document parsing tool to parse the PDF file into an XML file. The parsed content of the XML file includes the main text, references, and document structure information.
[0020] Locate the reference part of the XML file to extract the title and serial number of the document, and extract the corresponding reference sentence and the reference paragraph in the body of the XML file according to the title and serial number of the document, and use the reference sentence and reference paragraph as external knowledge for subsequent retrieval.
[0021] Preferably, the text segmentation module performs a coreference resolution operation on the external knowledge through a large language model, and performs text segmentation on the external knowledge that has undergone coreference resolution, and decomposes the text into sub-units, specifically:
[0022] Taking external knowledge as the initial text block, the initial text block is semantically analyzed through a large language model to determine the specific entity or concept referred to by the pronoun in the initial text block, and the pronoun is replaced with the corresponding entity to complete the reference resolution process.
[0023] By performing syntactic analysis on the text block after coreference resolution, the sentences containing reference marks in the text block are identified and extracted, and the sentences containing reference marks are saved as subunits.
[0024] For complex sentences or long sentences containing quotation marks in the text block, the complete sentences are extracted as subunits according to the grammatical structure to ensure semantic integrity.
[0025] Preferably, the text vectorization and dimension reduction module performs dimension reduction on the obtained text vector as follows:
[0026] A similarity graph is constructed based on the text vectors by calculating the similarity of each pair of data.
[0027] Heat diffusion modeling is performed by solving the Laplacian matrix of the similarity graph, and the local similarity is extended to the entire similarity graph through the diffusion matrix.
[0028] The similarity graph after heat diffusion is subjected to feature dimensionality reduction through spectral embedding, and a low-dimensional representation of the text vector is output.
[0029] Preferably, the hierarchical clustering and knowledge tree construction module performs hierarchical clustering on the text vector after dimension reduction, and constructs a retrieval tree according to the hierarchical clustering result as follows:
[0030] Initialize the text vectors, treating each text vector as a cluster.
[0031] Calculate the similarity between all cluster pairs and select the cluster pairs with the highest similarity to merge.
[0032] Cluster merging is continued and the similarity matrix is updated until all text vectors are merged, and a retrieval tree is constructed based on the hierarchical clustering results.
[0033] The root node of the retrieval tree represents the set of all text vectors, the internal nodes represent clusters, and the leaf nodes correspond to specific text vectors.
[0034] Preferably, the retrieval adjustment module uses the large language model to merge the information of nodes in the retrieval tree in detail as follows:
[0035] The search tree is merged from bottom to top in pairs, with the pair of nodes closest to each other being the merged objects, until the merge is completed, and finally the entire search tree is formed.
[0036] For each pair of nodes to be merged, the merged node information is generated according to the Prompt preset by the large language model.
[0037] In another aspect, the present invention provides an intelligent document retrieval method, comprising the following steps:
[0038] The cited articles are obtained through crawler tools as external knowledge for subsequent retrieval.
[0039] The external knowledge is subjected to a coreference resolution operation through a large language model, and the external knowledge that has undergone coreference resolution is segmented into text blocks and decomposed into sub-units.
[0040] The text of each sub-unit is vectorized and the dimension of the obtained text vector is reduced.
[0041] Hierarchical clustering is performed on the text vector after dimensionality reduction, a retrieval tree is constructed based on the hierarchical clustering results, and the information of the nodes in the retrieval tree is merged through a large language model.
[0042] During the retrieval process, the height of the retrieval tree is adjusted according to the different requirements of the user input query, and a matching search is performed based on the adjusted retrieval tree to output the retrieval results.
[0043] On the other hand, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the intelligent document retrieval method as described in any embodiment of the present invention is implemented.
[0044] In yet another aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the intelligent document retrieval method as described in any embodiment of the present invention is implemented.
[0045] Compared with the prior art, the present invention has the following technical effects:
[0046] 1. The present invention can more accurately understand the semantic relationship in the text through refined text segmentation and vectorization processing, especially when processing complex texts, and can effectively improve the retrieval accuracy.
[0047] 2. By adjusting the height of the retrieval tree, the present invention can flexibly adjust the depth and accuracy of the retrieval according to different needs, thereby improving the efficiency and flexibility of document retrieval and optimizing the retrieval performance.
[0048] 3. The present invention makes full use of the condensed knowledge in the quoted sentences to improve the utilization rate of external knowledge and enhance the relevance between the retrieved documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is an overall flow chart of the intelligent document retrieval method described in the present invention. DETAILED DESCRIPTION
[0050] In order to make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in combination with specific embodiments of the present application and with reference to the accompanying drawings.
[0051] Embodiment 1
[0052] This embodiment provides an intelligent document retrieval system, including: a reference document extraction module, a text segmentation module, a text vectorization module, a hierarchical clustering and retrieval tree construction module and a retrieval adjustment module.
[0053] The cited literature extraction module is used to obtain cited articles as external knowledge through crawler tools for subsequent retrieval.
[0054] The text segmentation module is used to perform a coreference resolution operation on external knowledge through a large language model, and to segment the external knowledge that has undergone coreference resolution into sub-units. Specifically, the choice of large language models is not limited here, and may include but is not limited to the GPT series, BERT, T5, RoBERTa, XLNet, ERNIE, etc.
[0055] The text vectorization and dimensionality reduction module is used to perform text vectorization processing on each subunit and reduce the dimensionality of the obtained text vector. Specifically, the vectorization processing method is not limited here. In this embodiment, the OPENAI model is preferably used for text vectorization processing.
[0056] The hierarchical clustering and retrieval tree construction module is used to perform hierarchical clustering on the text vector after dimensionality reduction, build a retrieval tree based on the hierarchical clustering results, and merge the information of the nodes in the retrieval tree through a large language model.
[0057] The search adjustment module is used to adjust the height of the search tree according to the different requirements of the user input query during the search process, perform the search according to the adjusted search tree, and output the search results. Specifically, during the search process, the user's search description can be input according to the user's needs, the height of the search tree is defined according to different user needs, and the search results are obtained by calculating the TOPN nodes through the cosine similarity of the vector.
[0058] Furthermore, users can also use big language to perform comprehension queries for search results. After the user enters a query, the big language model generates the answer the user needs by semantically matching and contextually analyzing the TOP N nodes with the user query, so that the user can understand the searched documents.
[0059] As a preferred implementation of this embodiment, the cited articles obtained as external knowledge by the crawler tool in the cited document extraction module are specifically:
[0060] By identifying the database (such as PubMed) to which each existing document belongs, constructing a URL to access the detailed page of each document, using crawler tools (Requests library and BeautifulSoup library) to parse the HTML structure in the page and extract the document title information, which may include the title and document identifier (such as PMC ID, which is a unique identifier used to identify a document in the PubMed database, and through which citation articles can be further obtained);
[0061] Use the document title information to construct a URL to download the PDF file of the article that cites the document, and use a crawler tool to download the PDF file;
[0062] Use a document parsing tool (such as Grobid) to parse the PDF file into an XML file in TEI format. The parsed content of the XML file includes the main text, reference part and document structure information.
[0063] Locate the reference part of the XML file to extract the title and serial number of the document, and extract the corresponding reference sentence and the reference paragraph in the body of the XML file according to the title and serial number of the document, and use the reference sentence and reference paragraph as external knowledge for subsequent retrieval.
[0064] Furthermore, the extraction of quoted sentences can use string matching methods to locate the quoted mark (such as (Smith et al., 2020) or [1]) and extract the complete sentence containing the quoted sentence. The context information (such as paragraphs) of the quoted sentence will also be extracted to ensure semantic integrity.
[0065] As a preferred implementation of this embodiment, the text segmentation module performs a coreference resolution operation on the external knowledge through a large language model, and performs text segmentation on the external knowledge that has undergone coreference resolution, and decomposes it into sub-units, specifically:
[0066] Taking external knowledge as the initial text block, the initial text block is semantically analyzed through a large language model to determine the specific entity or concept referred to by the pronoun in the initial text block, and the pronoun is replaced with the corresponding entity to complete the reference resolution process.
[0067] The pronoun refers to a word that replaces the previously mentioned noun or concept in the text, such as "he", "she", "it" in English, "he", "she", "it" in Chinese.
[0068] The text segmentation module determines the specific entity or concept referred to by the pronoun through context analysis. For example, in "Smith (2020) proposed a new method, and it was widely adopted.", "it" refers to "method". Through the semantic analysis of the large language model, this reference relationship is automatically identified and the pronoun is replaced with its corresponding entity.
[0069] In the text segmentation process, pronoun identification will affect the decomposition of sub-units. In a sentence with a reference, if there is a pronoun, the text segmentation module will first perform pronoun resolution and use the resolved text as a new sub-unit for subsequent processing. This ensures that the reference relationship in the text is correctly processed and does not affect the integrity and semantics of the reference sentence.
[0070] By performing syntactic analysis on the text block after coreference resolution, we identify and extract sentences containing reference markers in the text block. Reference markers usually appear in a specific format, such as "(Smith et al., 2020)", "[1]", etc. Sentences containing reference markers are saved as subunits, and each sentence containing reference information is regarded as an independent subunit.
[0071] For complex sentences or long sentences containing quotation marks in the text block, the complete sentence is extracted as a subunit according to the grammatical structure to ensure semantic integrity. Ensure that the entire sentence with quotation information is taken as a subunit.
[0072] In complex sentences involving references, the principal clause and the subordinate clauses are processed together, and the entire sentence (including the principal clause and the related subordinate clauses) is regarded as a subunit to ensure semantic integrity.
[0073] As a preferred implementation of this embodiment, the intelligent document retrieval system described in this embodiment supports multilingual text input, such as English, Chinese, French, German, Spanish, etc. When processing multilingual text, the subunit decomposition rules are adjusted according to the grammar and structure of different languages, but the basic principle of division is still to divide the blocks according to the sentences with citations. Further, the differences in the decomposition rules are as follows:
[0074] For English: Quoted sentences in English texts are usually clearly marked, such as "Smith (2020)", "[1]", etc. The text chunking module identifies sentences with quotations based on these tags.
[0075] For Chinese: Citation information in Chinese sentences may be represented by "(author, year)" or similar formats. The text chunking module identifies sentences containing citations through word segmentation and syntactic analysis and treats them as subunits.
[0076] For other languages: For languages such as French and German, the text chunking module also identifies sentences with citation information based on the grammatical structure of the language and processes them as subunits.
[0077] As a preferred implementation of this embodiment, the text vectorization and dimensionality reduction module performs dimensionality reduction on the obtained text vector, and the dimensionality reduction method is not limited. This embodiment preferably uses the PHATE algorithm for dimensionality reduction, specifically:
[0078] A similarity graph is constructed based on the text vectors by calculating the similarity of each pair of data.
[0079] Heat diffusion modeling is performed by solving the Laplacian matrix of the similarity graph, and the local similarity is extended to the entire similarity graph through the diffusion matrix.
[0080] The similarity graph after heat diffusion is subjected to feature dimensionality reduction through spectral embedding, and a low-dimensional representation of the text vector is output.
[0081] The above dimensionality reduction operation not only retains the local and global structure of text data, but also improves the efficiency and accuracy of retrieval through dimensionality reduction, providing effective support for subsequent hierarchical clustering and retrieval tree construction.
[0082] The purpose of hierarchical clustering is to group texts according to their similarities, thereby providing a more efficient structure for retrieval tasks. Hierarchical clustering algorithms such as hierarchical clustering and agglomerative clustering can be used. The algorithm starts with each text vector as a separate cluster, calculates the similarities between clusters, and gradually merges the most similar clusters until a complete tree structure is formed. As a preferred implementation of this embodiment, the hierarchical clustering and knowledge tree construction module hierarchically clusters the text vectors after dimensionality reduction, and constructs a retrieval tree based on the hierarchical clustering results as follows:
[0083] Initialize the text vectors, treating each text vector as a cluster (the size of each cluster is 1).
[0084] Calculate the similarity between all cluster pairs, and select the cluster pairs with the highest similarity for merging. Specifically, by calculating the similarity between the text vectors after dimensionality reduction (such as Euclidean distance or cosine similarity), a similarity matrix of the text vectors is constructed, which reflects the similarity between each text vector and other texts.
[0085] Clusters are continuously merged and the similarity matrix is updated until all text vectors are merged, and a retrieval tree is constructed based on the hierarchical clustering results. Different similarity measures can be selected when merging clusters, including but not limited to single linkage, complete linkage or average linkage.
[0086] The structure of the retrieval tree is as follows: Root node: The root node represents the set of all text vectors and is the topmost node of the tree, containing the text information of all documents. Internal nodes: Internal nodes represent clusters, and the text vectors in the clusters have high semantic similarity. Each internal node will contain multiple child nodes, representing multiple clusters at the same level. Leaf nodes: Leaf nodes correspond to specific text vectors. Leaf nodes are the bottom layer of the retrieval tree, and each leaf node points to the specific content of a document.
[0087] The process of building a retrieval tree:
[0088] Initial structure: According to the results of hierarchical clustering, the root node is initialized to a set containing all text vectors.
[0089] Stepwise splitting: Create child nodes and group similar text vectors together based on the structure of the clusters. Each cluster is treated as an internal node.
[0090] Controlling the depth of the tree: The depth of the tree is controlled by a threshold that can be set during the query process. For example, some large document collections may require a deeper tree structure to refine different topics and subtopics.
[0091] In order to improve the structure and retrieval efficiency of the retrieval tree after hierarchical clustering, node merging is an important optimization step. In particular, when using a large language model for node merging, a bottom-up two-by-two merging method is adopted to avoid exceeding the window limit of the large language model.
[0092] As a preferred implementation of this embodiment, the retrieval adjustment module uses a large language model to merge the information of nodes in the retrieval tree as follows:
[0093] The search tree is merged from bottom to top in pairs, with the pair of nodes closest to each other being the merged objects, until the merge is completed, and finally the entire search tree is formed.
[0094] For each pair of clusters to be merged, the merged node information is generated according to the Prompt preset by the large language model. The Prompt here is not used to determine whether to merge, but as part of the merging process, it provides the identification or label of the new node after the merger.
[0095] Furthermore, the purpose of the prompt design is to ensure that no useful information is lost during the merge process. The prompt format can be set as follows:
[0096] Based on the quotations from the cited papers, summarize their core themes and research directions and combine them into a coherent summary.
[0097] Requirements: 1. Generate less than two sentences. 2. Remember as much information as possible. 3. Use professional terms and do not make up without permission.
[0098] During the merging process, the large language model will automatically generate merged node information based on the preset prompt to avoid exceeding the input limit of the model.
[0099] In the bottom-up two-by-two merging, the number of nodes merged each time is relatively small, and a smaller set of nodes is processed each time to avoid inputting too much information at one time, thereby ensuring that each input does not exceed the window limit of the large language model. At the same time, each merge only involves a small amount of text data, and the operations during each merge are more focused on the content of the current cluster, which improves the efficiency of the merge. The merged node information (such as summary or label) will continue to merge upward as a new node until the entire retrieval tree is finally formed.
[0100] Through the bottom-up merging operation, the final retrieval tree has the following structure:
[0101] Root node: The root node represents all text collections and contains all clusters after the final merger.
[0102] Internal nodes: Each internal node represents a cluster, which contains text vectors with similar semantics. Each internal node carries representative information of the cluster (such as topic tags, abstracts, or other identifiers).
[0103] Leaf nodes: Each leaf node is still the original text vector, which is gradually classified into the corresponding cluster clusters through hierarchical clustering and node merging operations.
[0104] As a preferred implementation of this embodiment, the height of the search tree can also be adjusted according to different requirements through the search adjustment module, specifically:
[0105] Seeking the best default parameters: Through comparative experiments, find the best height as the initial parameter.
[0106] Dynamic adjustment: Dynamically adjust the height of the search tree according to the needs of the user during the query. For example, the user can specify that a more precise or broader search result is required, and the search adjustment module adjusts the tree level according to the needs to query more detailed results.
[0107] Adaptive algorithm: Based on the feedback information of the search results, the height of the search tree is adaptively adjusted. If the search results do not meet the expected accuracy or are too broad, you can choose to increase the depth of the tree; if the search results are too redundant or the depth is too large, resulting in reduced efficiency, you can reduce the depth of the tree.
[0108] Embodiment 2
[0109] Accordingly, this embodiment provides an intelligent document retrieval method, which is implemented based on the intelligent document retrieval system described in Embodiment 1. Figure 1 As shown, the following steps are included:
[0110] The cited articles are obtained through crawler tools as external knowledge for subsequent retrieval.
[0111] The external knowledge is subjected to a coreference resolution operation through a large language model, and the external knowledge that has undergone coreference resolution is segmented into text blocks and decomposed into sub-units.
[0112] The text of each sub-unit is vectorized and the dimension of the obtained text vector is reduced.
[0113] Hierarchical clustering is performed on the text vector after dimensionality reduction, a retrieval tree is constructed based on the hierarchical clustering results, and the information of the nodes in the retrieval tree is merged through a large language model.
[0114] During the retrieval process, the height of the retrieval tree is adjusted according to the different requirements of the user input query, and a matching search is performed based on the adjusted retrieval tree to output the retrieval results.
[0115] Embodiment 3
[0116] This embodiment provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the intelligent document retrieval method as described in any embodiment of the present invention is implemented.
[0117] Embodiment 4
[0118] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the intelligent document retrieval method as described in any embodiment of the present invention is implemented.
[0119] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.
[0120] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0121] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0122] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.
[0123] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. An intelligent document retrieval system, characterized in that: include: Citation document extraction module, text segmentation module, text vectorization module, hierarchical clustering and retrieval tree construction module and retrieval adjustment module; The citation literature extraction module is used to obtain the cited articles as external knowledge through crawler tools for subsequent retrieval; The text segmentation module is used to perform a coreference resolution operation on external knowledge through a large language model, and to segment the text of the external knowledge that has undergone coreference resolution into sub-units. The text vectorization and dimension reduction module is used to vectorize the text of each sub-unit and reduce the dimension of the obtained text vector; Hierarchical clustering and retrieval tree construction module, which is used to perform hierarchical clustering on the text vector after dimensionality reduction, build a retrieval tree based on the hierarchical clustering results, and merge the information of nodes in the retrieval tree through a large language model; The retrieval adjustment module is used to adjust the height of the retrieval tree according to the different requirements of the user input query during the retrieval process, perform matching retrieval based on the adjusted retrieval tree, and output the retrieval results.
2. The intelligent document retrieval system according to claim 1, characterized in that: The cited articles obtained as external knowledge by the crawler tool in the cited document extraction module are specifically: By identifying the database to which each existing document belongs, constructing a URL to access the detailed page of each document, parsing the HTML structure in the page, and extracting the document title information; Use the document title information to construct a URL to download the PDF file of the article that cites the document, and use a crawler tool to download the PDF file; Use the document parsing tool to parse the PDF file into an XML file. The parsed content of the XML file includes the main text, references, and document structure information. Locate the reference part of the XML file to extract the title and serial number of the document, and extract the corresponding reference sentence and the reference paragraph in the body of the XML file according to the title and serial number of the document, and use the reference sentence and reference paragraph as external knowledge for subsequent retrieval.
3. The intelligent document retrieval system according to claim 1, characterized in that: In the text segmentation module, external knowledge is subjected to a coreference resolution operation through a large language model, and the external knowledge subjected to the coreference resolution processing is subjected to text segmentation, which is decomposed into sub-units, specifically: Take the external knowledge as the initial text block, perform semantic analysis on the initial text block through the large language model, determine the specific entity or concept referred to by the pronoun in the initial text block, and replace the pronoun with the corresponding entity to complete the reference resolution process; By performing syntactic analysis on the text block after reference resolution, the sentences containing reference marks in the text block are identified and extracted, and the sentences containing reference marks are saved as subunits; For complex sentences or long sentences containing quotation marks in the text block, the complete sentences are extracted as subunits according to the grammatical structure to ensure semantic integrity.
4. The intelligent document retrieval system according to claim 1, characterized in that: The text vectorization and dimensionality reduction module performs dimensionality reduction on the obtained text vector as follows: Construct a similarity graph based on the text vectors by calculating the similarity of each pair of data; Heat diffusion modeling is performed by solving the Laplacian matrix of the similarity graph, and the local similarity is extended to the entire similarity graph through the diffusion matrix; The similarity graph after heat diffusion is subjected to feature dimensionality reduction through spectral embedding, and a low-dimensional representation of the text vector is output.
5. The intelligent document retrieval system according to claim 1, characterized in that: In the hierarchical clustering and knowledge tree construction module, hierarchical clustering is performed on the text vector after dimension reduction, and a retrieval tree is constructed according to the hierarchical clustering result as follows: Initialize the text vectors and treat each text vector as a cluster; Calculate the similarity between all cluster pairs and select the cluster pairs with the highest similarity to merge; Continuously merge clusters and update the similarity matrix until all text vectors are merged, and then build a retrieval tree based on the hierarchical clustering results; The root node of the retrieval tree represents the set of all text vectors, the internal nodes represent clusters, and the leaf nodes correspond to specific text vectors.
6. The intelligent document retrieval system according to claim 5, characterized in that: The retrieval adjustment module uses the large language model to merge the information of nodes in the retrieval tree specifically as follows: The search tree is merged from bottom to top, with the two merged objects being the closest pair of nodes, until the merge is completed, and finally the entire search tree is formed; For each pair of nodes to be merged, the merged node information is generated according to the Prompt preset by the large language model.
7. An intelligent document retrieval method, characterized in that: The method is implemented based on the intelligent document retrieval system according to any one of claims 1 to 6, and comprises the following steps: Use crawler tools to obtain cited articles as external knowledge for subsequent retrieval; The external knowledge is subjected to a coreference resolution operation through a large language model, and the coreference resolution external knowledge is segmented into text blocks and decomposed into sub-units; Perform text vectorization on each sub-unit and reduce the dimension of the obtained text vector; Perform hierarchical clustering on the text vector after dimensionality reduction, build a retrieval tree based on the hierarchical clustering results, and merge the information of the nodes in the retrieval tree through a large language model; During the retrieval process, the height of the retrieval tree is adjusted according to the different requirements of the user input query, and a matching search is performed based on the adjusted retrieval tree to output the retrieval results.
8. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the intelligent document retrieval method according to claim 7 when executing the computer program.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the intelligent document retrieval method according to claim 7 is implemented.
Citation Information
Patent Citations
A local citation recommendation method and system based on neural machine translation technology
CN109145190A
Literature subject term aggregation method and device, computer equipment and readable storage medium
CN111898366A
Literature retrieval method and device based on multi-dimensional semantic index and medium
CN118093770A
Retrieval enhancement method and device, equipment and storage medium
CN118394793A
Short text query expansion enhancement retrieval method based on knowledge base hierarchical tree structure
CN118861088A