Literature retrieval method and system based on large language model

By introducing large language models into the literature search system, the problems of low efficiency and insufficient accuracy in dealing with large-scale and diverse literatures are solved, and more efficient and accurate literature search is achieved.

CN120067296APending Publication Date: 2025-05-30NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510168761.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When traditional literature search methods deal with large-scale and diverse literature, they lack semantic understanding ability, resulting in too many search results and low correlation, making it difficult to meet the needs of users.

Method used

Using a literature search method based on large language models, the introduction of large language models gives the search process stronger semantic understanding and analysis capabilities, build a search formula and search in selected databases, perform data cleaning and key literature screening.

Benefits of technology

It improves the efficiency and accuracy of searches, can more accurately judge the correlation between the literature and the search topic, avoids the limitations and ambiguity of traditional keyword searches, and provides a more efficient and convenient literature search experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067296A_ABST
    Figure CN120067296A_ABST
Patent Text Reader

Abstract

The invention discloses a literature retrieval method and system based on a large language model. The literature retrieval method comprises the steps that a retrieval subject is determined; selecting a proper literature source; constructing a final retrieval formula; utilizing a large language model to select a database related to the retrieval subject; performing retrieval in a selected database by using the constructed retrieval formula, namely performing first screening to obtain a list of related literatures and information features of each literature; performing data cleaning on the obtained literature list; and understanding the literature content by using a large language model, and judging whether the literature conforms to a retrieval subject, namely screening key literatures for the second time, and storing the literature with the content conforming to the requirement in a local file. The method has the core advantage that the large language model is introduced, so that the retrieval process is endowed with stronger semantic comprehension and analysis capability, and the retrieval efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of document retrieval, and in particular relates to a document retrieval method and system based on a large language model. Background Art

[0002] With the explosive growth of document information, users have an increasing demand for accurate retrieval. Traditional retrieval methods are unable to cope with large-scale and diverse documents, resulting in too many retrieval results with low relevance, making it difficult to meet user needs. Specifically, traditional retrieval methods face the following problems: lack of semantic understanding capabilities, relying only on surface matching of keywords, and being unable to handle the complexity of natural language. For example, when searching for "apple", it may include both documents related to the fruit and documents related to Apple. How to give the document retrieval system the ability to understand natural language in order to efficiently and accurately locate the target document in a large amount of documents has become a key issue that needs to be urgently addressed in the current field of document retrieval technology. Summary of the invention

[0003] In order to solve the problems of low retrieval efficiency and insufficient accuracy in the prior art, the purpose of the present invention is to provide a document retrieval method and system based on a large language model. By introducing a large language model, the retrieval process is endowed with stronger semantic understanding and analysis capabilities, thereby improving the efficiency of retrieval.

[0004] The purpose of the present invention is achieved through the following technical solutions: A document retrieval method based on a large language model comprises the following steps: Step S1. Determine the search topic: The user first identifies the topic to be searched, inputs the question into a large language model, and obtains a search formula corresponding to the topic.

[0005] Step S2. Select the source of retrieval literature: Select appropriate academic journals or academic conferences as the source of literature.

[0006] Step S3. Constructing a search formula: constructing a final search formula based on steps S1 and S2.

[0007] Step S4. Select database: given a search topic and several candidate databases, let the large language model output the database that best matches the search topic as the selected database.

[0008] Step S5. Retrieve documents: Use the search formula constructed in step S3 to search in the selected database, i.e., the first screening, to obtain a list of relevant documents and the information characteristics of each document.

[0009] Step S6. Data cleaning: Perform data cleaning on the document list obtained in step S5 and delete irregular data items for subsequent processing.

[0010] Step S7. Screening of key literature: For the literature list obtained by cleaning in Step S6, a second screening is performed based on the literature content, that is, a large language model is used to parse the literature title and abstract to determine whether it meets the retrieval theme. Save the literature whose content meets the requirements to a local file.

[0011] In the above-mentioned Step S2, the appropriate academic journals or academic conferences are automatically obtained from a preset list of journal conferences. The preset list of journal conferences consists of high-quality journals and conferences in this field, such as academic journals in the first and second districts of the Chinese Academy of Sciences or Class A conferences recommended by CCF, etc. Users can determine appropriate academic journals or academic conferences as the literature source according to this list.

[0012] The construction of the retrieval formula in the above-mentioned Step S3 includes defining the retrieval theme, retrieval scope, and literature source. By selecting subject terms, the retrieval scope is limited to the literature directly related to the research theme to avoid overly broad retrieval results. Select a specific retrieval scope, such as "subject / title / abstract", to ensure a high correlation between the retrieval results and the keywords. Select authoritative academic conferences and journals as the literature source to ensure the academic quality and reliability of the retrieval results.

[0013] The database in the above-mentioned Step S4 is determined based on the judgment of a large language model. Given the retrieval theme and several candidate databases, the large language model is made to output the database that best meets the retrieval theme as the selected database.

[0014] In the above-mentioned Step S7, the abstract and title of the literature are used as the input of the large language model to make the large language model judge whether the literature meets the retrieval theme.

[0015] Use a large language model to judge whether the literature meets the retrieval theme based on the title and abstract of the literature, and process multiple literatures simultaneously in a parallel execution manner.

[0016] In the above-mentioned Steps S5 and S7, Python is used to construct a script to automatically save the screened literature to the local.

[0017] A literature retrieval system based on a large language model, characterized by including: A literature source retrieval formula construction module: used to select appropriate academic journals or academic conferences according to a preset list of journal conferences when a retrieval theme is given, and users determine appropriate academic journals or academic conferences as the literature source according to this list; A database selection module: selects an appropriate database according to the retrieval theme; Retrieval module: It is used to perform retrieval in the selected database according to the retrieval formula constructed by the user, that is, the first screening, obtain a list of relevant documents and the information characteristics of each document, and save them locally. Understanding module: It is used to understand the content of the documents and screen out the documents that meet the requirements according to the content of the documents.

[0018] The present invention has the following beneficial effects: The literature retrieval method based on a large language model provided by the present invention, compared with the traditional keyword retrieval method, its core advantage lies in the introduction of the large language model, which endows the retrieval process with stronger semantic understanding and analysis capabilities. The large language model can deeply analyze the titles and abstracts of the documents, understand the semantics and intentions behind them, so as to more accurately judge the relevance of the documents to the retrieval topic, avoiding the limitations and ambiguities of the traditional keyword retrieval. This makes the retrieval results more accurate, can effectively eliminate the unwanted retrieval results, improve the retrieval efficiency, and provide a more efficient and convenient literature retrieval experience for users. Description of the Drawings

[0019] Figure 1 It is the flowchart of the method in the embodiments of this specification Detailed Embodiments

[0020] The following will further elaborate on the content of the present invention in combination with the embodiments, but it is not a limitation to the present invention.

[0021] Embodiment: Taking the search for English literature related to "recommendation system" as an example, the technical solution of the present invention will be further described in detail.

[0022] A literature retrieval method based on a large language model includes the following steps: Step S1. Determine the retrieval topic: The user first clarifies that the topic to be retrieved is "recommendation system", and inputs the prompt "Retrieve papers related to the'recommendation system' topic and provide relevant retrieval formulas" into a large language model (such as ChatGPT, DeepSeek, etc.), so that it outputs the retrieval formulas related to the "recommendation system" topic.

[0023] Step S2. Select the source of the retrieved literature: Select a suitable academic journal or academic conference as the source of the literature. Automatically obtain a suitable academic journal or academic conference as the source of the literature from the preset list of journals and conferences. The preset list of journals and conferences consists of high-quality journals and conferences in this field. In this embodiment, Class A and Class B conferences recommended by CCF are selected, and the sub-fields are selected as "Artificial Intelligence", "Database / Data Mining / Content Retrieval", "Cross / Comprehensive / Emerging" to obtain the retrieval formula for the source of the literature.

[0024] Step S3. Construct a search query: The search query constructed based on the topic "recommendation system" determined in the first step is "("recommendation system" OR "recommender system") AND ("collaborative filtering" OR "content-based" OR "hybrid" OR "matrix factorization" OR "deep learning" OR "neural networks" OR "personalization" OR "user preferences" OR "evaluation" OR "metrics")", and the search source is "Subject / Title / Abstract". Also, the search query for the literature source obtained in Step S2 is "AAAI or CICAI or UAI or IJCAI or KDD or SIGIR or WWW or RecSys or CIKM or WSDM or COLT or ICLR or ICML or NeurIPS or ACL or EMNLP", and the search source is "Conference information".

[0025] Step S4. Select a database: Given the search topic and a number of candidate databases, let the large language model output the database that best matches the search topic as the selected database. In this embodiment, based on the input topic "recommendation system", the large language model outputs the database most similar to this topic as engineering village, a comprehensive engineering science database that provides users with comprehensive engineering information retrieval services.

[0026] Step S5. Retrieve literature: Use the constructed search query to retrieve in engineering village, that is, the first screening, to obtain a list of relevant literature and the information characteristics of each literature, and save this information locally.

[0027] Step S6. Data cleaning: Clean the literature list obtained in Step S5, delete non-standard data items for subsequent further reading of this literature list. The information to be deleted includes but is not limited to "Open Access", "Database", etc.

[0028] Step S7. Screening of key documents: For the list of documents obtained after cleaning in Step S6, a second screening is performed based on the content of the documents. Specifically, Python code is used to read this list of documents. For each piece of document information therein, the abstract and title of the document are input into a large language model, and the large language model is used to determine whether the document meets the retrieval topic. The documents that meet the requirements are saved to a local file.

[0029] A document retrieval system based on a large language model, which implements the above method, includes: Document source retrieval formula construction module: When a given retrieval topic is provided, it is used to select appropriate academic journals or academic conferences according to a preset list of journals and conferences, and the user determines an appropriate academic journal or academic conference as the document source based on this list; Database selection module: Selects an appropriate database according to the retrieval topic; Retrieval module: It is used to perform a retrieval in the selected database according to the retrieval formula constructed by the user, that is, the first screening, obtain a list of relevant documents and the information characteristics of each document, and save them locally; Understanding module: It is used to understand the content of the documents and screen out the documents that meet the requirements according to the content of the documents.

[0030] In this way, all the documents related to the retrieval topic are obtained. By introducing the text understanding function of the large model, the ambiguity of traditional keyword retrieval is avoided, making the retrieval more efficient and the retrieval results more accurate.

Claims

1. A document retrieval method based on a large language model, characterized in that: The following steps are involved: Step S1. Determine the search topic: The user first determines the topic to be searched, inputs the question into the large language model, and obtains the search formula corresponding to the topic; Step S2. Select the source of retrieval literature: select appropriate academic journals or academic conferences as the source of literature; Step S3. Constructing a search formula: constructing a final search formula based on steps S1 and S2; Step S4. Select database: select an appropriate database for searching literature given the search topic; Step S5. Retrieve documents: Use the constructed search formula to search in the selected database, i.e., the first screening, to obtain a list of relevant documents and the information characteristics of each document; Step S6. Data cleaning: Perform data cleaning on the document list obtained in step S5 and delete irregular data items for subsequent processing; Step S7. Based on the large language model, the key documents are screened for the second time according to the document content, and the documents whose content meets the requirements are saved to a local file.

2. The document retrieval method based on a large language model according to claim 1, characterized in that: In step S2, suitable academic journals or academic conferences are automatically obtained from a preset journal and conference list, and the preset journal and conference list consists of high-quality journals and conferences in the field.

3. The document retrieval method based on a large language model according to claim 1, characterized in that: The construction of the search formula in step S3 includes limiting the search subject, search scope, and document source; by selecting subject words, the search scope is limited to documents directly related to the research subject to avoid the search results being too broad.

4. The document retrieval method based on a large language model according to claim 1, characterized in that: The database in step S4 is determined based on a large language model. Given a search topic and several candidate databases, the large language model is made to output a database that best matches the search topic as the selected database.

5. The document retrieval method based on a large language model according to claim 1, characterized in that: In step S7, the abstract and title of the document are used as inputs of the large language model, so that the large language model determines whether the document meets the search topic.

6. The document retrieval method based on a large language model according to claim 5, characterized in that: A large language model is used to determine whether a document meets the search topic based on its title and abstract, and multiple documents are processed simultaneously through parallel execution.

7. The document retrieval method based on a large language model according to claim 1, characterized in that: In steps S5 and S7, Python is used to construct a script to automatically save the screened documents locally.

8. A document retrieval system based on a large language model, characterized in that: include: Literature source search construction module: used to select appropriate academic journals or academic conferences according to the preset journal and conference list when a search topic is given. Users can determine appropriate academic journals or academic conferences as literature sources based on the list; Database selection module: select the appropriate database according to the search topic; Retrieval module: used to search the selected database according to the search formula constructed by the user, that is, the first screening, obtain the list of relevant documents and the information characteristics of each document, and save them locally; Understanding module: used to understand the content of the document and filter out the documents that meet the requirements based on the content of the document.