Literature search and academic report generation method, system and device and storage medium
User problems are obtained through natural language dialogue, rewritten and expanded, and literature screening and reordering is performed by combining rank sum ratio method and similarity calculation. Large language models are used to assist users in understanding papers and generating academic reports, solving the problems of incomplete paper searches and poor article quality in the existing technology, and achieving more efficient and comprehensive paper searches and academic report generation.
Patent Information
- Application Number
- CN202411941756.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
AI Technical Summary
The existing paper search tools are based on keyword matching, making it difficult to search for professional terms required by beginners, and are unable to effectively capture the deep semantic relationships between the literature, resulting in incomplete searches and poor article quality.
Obtain user problems through natural language dialogue, rewrite and expand the problem, combine rank sum ratio method and similarity calculation, conduct initial screening and reordering of literature, use large language models to assist users in understanding papers, and generate academic reports.
It improves the accuracy and comprehensiveness of paper searches, simplifies the difficulty of users to read literature, lowers the threshold, and accelerates the process of beginners to understand the corresponding fields.
Smart Images

Figure CN120011497A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent retrieval technology, and in particular relates to a document search and academic report generation method, system, device and storage medium. Background Art
[0002] Common paper search tools currently available on the market, such as PubMed, Google Scholar, and CNKI, are all based on keyword matching search technology. For beginners in a certain field, it is difficult to find the most representative or professional keywords in this field at the beginning, which makes it difficult to search for relevant papers. At the same time, many words, whether in Chinese or English, have many synonyms or antonyms. Therefore, the search technology based on keyword matching will miss many articles that are highly relevant to the keywords entered by the user, resulting in incomplete search articles.
[0003] Currently, some developers use knowledge graphs or similarity to build paper search frameworks, which simplifies the difficulty of users searching for papers to a certain extent, and also improves the accuracy and comprehensiveness of recommending related papers. However, there are still some shortcomings:
[0004] 1. The paper search framework built using the knowledge graph has improved its comprehensiveness to a certain extent, but it is still limited by keyword matching technology, which leads to some shortcomings.
[0005] 2. For the paper search framework using similarity calculation, although this method has overcome the limitations brought by keyword matching technology, thereby improving the accuracy and comprehensiveness of paper recommendations, it is unable to conduct a more in-depth search and the quality of the articles is mixed. It is difficult to cover relatively new documents or old but classic documents.
[0006] At the same time, these paper search tools are unable to provide a comprehensive report on the current development of the field based on questions or keywords entered by users, nor can they provide auxiliary reading for a particular article.
[0007] A Chinese patent with publication number WO2024078141A1 discloses a method for predicting subject document retrieval, including: constructing a subject resource database, constructing a subject retrieval vocabulary for the digital resources of the document; storing the retrieval data in a file each time a retrieval is performed, searching and matching the retrieval data in the file to determine the retrieval frequency, constructing a retrieval information knowledge graph, and associating other users' retrieval history data to improve the knowledge graph; finally, the knowledge graph determines the retrieval strategy, predicts the retrieval documents and sorts them; when the user browses the retrieval results, the document data browsed and downloaded by the user is recorded and associated, the data is analyzed, and the degree of association between the retrieval strategy and the predicted documents is established; the retrieval results of other users are predicted according to the degree of association between the retrieval strategy and the predicted documents, and the retrieval documents are predicted and sorted and output. Although the invention predicts and sorts documents through word frequency statistics, knowledge graphs and user behavior analysis, there are problems of excessive data dependence and historical data deviation that may affect the accuracy of prediction. At the same time, excessive reliance on keyword matching may not effectively capture the deep semantic associations between documents. Summary of the invention
[0008] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a method, system, device and storage medium for document search and academic report generation to improve the accuracy of paper recommendation and the comprehensiveness of paper search, while simplifying the difficulty of users reading documents.
[0009] The technical solution of the present invention is as follows:
[0010] In one aspect, the present invention provides a method for document search and academic report generation, comprising the following steps:
[0011] S10: Obtaining user questions input by the user in the form of natural language dialogue.
[0012] S20: Rewrite the question according to the user's question.
[0013] S30: Obtain relevant literature from the search engine according to the rewritten user question, and calculate the similarity between the literature abstract and the user question to perform preliminary screening of the literature.
[0014] S40: Obtain detailed information on the literature after the initial screening, and re-sort the literature obtained by the initial screening by combining the rank sum ratio method and similarity to obtain recommended related papers.
[0015] S50: Based on the re-ranked related papers, select the TOP-N related papers for in-depth search and re-recommendation.
[0016] S51: The user selects a paper from all the recommended papers to read, and the user is assisted in understanding the paper by asking questions based on a large language model.
[0017] S52: Obtain the TOP-N related papers in step S50, extract relevant and useful information from the papers and input it into the large language model, and generate an academic report for user reference.
[0018] Preferably, the step S20 is specifically as follows:
[0019] S201: According to the user's question, relevant information of the user's question is searched through a search engine, and the relevant information is input into a large language model in combination with the user's question to expand the question.
[0020] S202: Use the knowledge base of the large language model to answer the user's question, and document the expanded user question and the answer of the large language model to obtain the rewritten user question for subsequent similarity calculation.
[0021] Preferably, the step S30 is specifically as follows:
[0022] S301: Obtain relevant literature from a search engine according to the rewritten user question and obtain a summary of the literature.
[0023] S302: Use the pre-trained word embedding model to generate embedding vectors for two types of text: user questions and document summaries.
[0024] S303: The language similarity between the two types of text is obtained by calculating the cosine similarity of the embedding vectors between the two types of text, specifically:
[0025]
[0026] Where CosineSimilarity() is the cosine similarity; A and B are the embedding vectors of the two texts respectively; · represents the dot product; ‖A‖ and ‖B‖ represent the modulus of the two embedding vectors respectively.
[0027] S304: Perform a preliminary screening of documents according to a set semantic similarity threshold.
[0028] Preferably, the step S40 is specifically as follows:
[0029] S401: Obtain detailed information on the literature after initial screening, including the number of citations of the paper, the impact factor of the journal in which the paper is published, and the similarity between the references of the paper.
[0030] S402: Setting the weight of each information factor in the detailed information of the document by using the rank sum ratio method, and calculating the weighted rank sum ratio of each information element. The calculation formula of the weighted rank sum ratio is specifically:
[0031]
[0032] Where, WRSR iis the weighted rank sum ratio of the i-th paper; n is the total number of papers, i∈[1,n]; p is the total number of information factors, j∈[1,p]; W j is the weight of the jth evaluation index; R ij is the rank of the information factor in the jth column of the ith paper.
[0033] S403: Re-rank the documents obtained by the initial screening by combining the weighted rank sum ratio and the score obtained by adding the similarities, so as to obtain recommended related papers.
[0034] Preferably, the step S50 is specifically as follows:
[0035] According to the re-ranked related papers, select the TOP-N related papers and obtain all the references and citations of the TOP-N related papers. Select the omitted papers and re-rank them according to the weighted rank sum ratio and similarity, and then screen the top M papers for re-recommendation.
[0036] Preferably, the step S51 is specifically as follows:
[0037] S511: Divide the document content into several text blocks based on the semantic information and completeness of the paper.
[0038] S512: Calculate the similarity between the text block and the user's question based on the question raised by the user.
[0039] S513: According to the calculated similarity, TOP-N text blocks are taken out as background information of the user question and input into the large language model, and different keywords are set to require the large language model to answer the user question.
[0040] S514: Summarize the answers to the questions obtained by the large language model based on different text blocks, so as to obtain the final answer to the question, and assist the user in understanding the paper being consulted.
[0041] Preferably, the step S52 is specifically as follows:
[0042] S521: Obtain the TOP-N relevant papers in step S50, and extract information from the papers in blocks according to the parts generated according to the requirements of the academic report template. The information includes corresponding topic background information, research progress in the current field, and outstanding contributions of the experiment.
[0043] S522: Based on the set keywords, the obtained information is further purified through a large language model to prevent harmful information from affecting the final output result.
[0044] S523: The purified information is merged with the user's question, and combined with the set academic report keywords, the information is passed into the large language model to generate an academic report for user reference.
[0045] On the other hand, the present invention provides a document search and academic report generation system, including a user problem processing module, a document retrieval and preliminary screening module, a deep retrieval module, a paper understanding auxiliary module and an academic report generation module.
[0046] The user question processing module is used to obtain user questions input by users in the form of natural language dialogue and rewrite questions according to the user questions.
[0047] The literature retrieval and preliminary screening module is used to obtain relevant literature from the search engine based on the rewritten user questions, calculate the similarity between the literature abstract and the user questions for preliminary screening of the literature, obtain detailed information on the literature after the preliminary screening, and re-sort the literature obtained from the preliminary screening by combining the rank sum ratio method and similarity to obtain recommended related papers.
[0048] The deep retrieval module is used to select the TOP-N related papers for deep search and re-recommendation based on the re-ranked related papers.
[0049] The paper comprehension assistance module is used for users to select papers from all recommended papers for review, and assists users in understanding the reviewed papers by asking questions based on a large language model.
[0050] The academic report generation module is used to obtain the TOP-N relevant papers selected in the deep search module, extract relevant and useful information from the papers and input it into the large language model to generate academic reports for user reference.
[0051] On the other hand, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for document search and academic report generation as described in any embodiment of the present invention is implemented.
[0052] On the other hand, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for document search and academic report generation as described in any embodiment of the present invention is implemented.
[0053] Compared with the prior art, the present invention has the following technical effects:
[0054] The present invention realizes the user's natural language conversational search, question rewriting, and initial screening and refined sorting of papers, thereby improving the accuracy and comprehensiveness of paper search; and by developing multiple functions including academic report generation, paper assisted reading, and innovative idea generation, it simplifies the difficulty of users reading literature, lowers the threshold, and accelerates beginners to understand the corresponding field. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is an overall flow chart of the document search and academic report generation method described in the present invention. DETAILED DESCRIPTION
[0056] In order to make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in combination with specific embodiments of the present application and with reference to the accompanying drawings.
[0057] Embodiment 1
[0058] This embodiment provides a method for document search and academic report generation. Figure 1 As shown, the following steps are included:
[0059] S10: Obtain user questions input by the user in the form of natural language dialogue. The input user questions are the scientific research fields or subject directions that the user is interested in, for example: Which research articles utilize RNA-sequencing (RNA-seq) or RNA in situ hybridization (RNA-scope) technologies for experimental validation or as part of the experimental process? (Which research articles utilize RNA sequence (RNA-seq) or RNA in situ hybridization (RNA-scope) technologies for experimental validation or as part of the experimental process?).
[0060] S20: Rewrite the question according to the user's question.
[0061] As a preferred implementation of this embodiment, step S20 is specifically as follows:
[0062] S201: Based on the user's question, search for relevant information about the user's question through a search engine, input the relevant information into a large language model in combination with the user's question to expand the question, enrich the details of the user's question, and thus improve the accuracy of paper search and recommendation. The selection of the large language model is not limited in practical applications, and may include but is not limited to the GPT series, BERT, RoBERTa, T5, XLNet, BART, etc.
[0063] S202: Use the knowledge base of the large language model to answer the user's question, and document the expanded user question and the answer of the large language model to obtain the rewritten user question for subsequent similarity calculation.
[0064] S30: Obtain relevant documents from the search engine according to the rewritten user question, and calculate the similarity between the document abstract and the user question to perform preliminary screening of the documents. The method of obtaining relevant documents is not limited here, and this embodiment preferably uses crawler technology to collect relevant documents.
[0065] As a preferred implementation of this embodiment, step S30 is specifically as follows:
[0066] S301: Obtain relevant literature from a search engine according to the rewritten user question and obtain a summary of the literature.
[0067] S302: Generate embedding vectors for two types of text, user questions and document abstracts, using a pre-trained word embedding model. The word embedding model can convert text into fixed-dimensional embedding vectors, which are used to capture the semantic information of the text. The choice of word embedding model is not limited here, and may include but is not limited to BERT, Word2Vec, GloVe, etc.
[0068] S303: The language similarity between the two types of texts is obtained by calculating the cosine similarity of the embedding vectors between the two types of texts. The cosine similarity is to evaluate the similarity between the two vectors by calculating the cosine value of the angle between the two vectors. Specifically,
[0069]
[0070] Where CosineSimilarity() is the cosine similarity; A and B are the embedding vectors of the two texts respectively; · represents the dot product; ‖A‖ and ‖B‖ represent the modulus of the two embedding vectors respectively.
[0071] S304: Perform a preliminary screening of documents according to a set semantic similarity threshold.
[0072] S40: Obtain detailed information on the literature after the initial screening, and re-sort the literature obtained by the initial screening by combining the rank sum ratio method and similarity to obtain recommended related papers.
[0073] As a preferred implementation of this embodiment, step S40 is specifically as follows:
[0074] S401: Obtain detailed information on the literature after initial screening, including the number of citations of the paper, the impact factor of the journal in which the paper is published, and the similarity between the references of the paper.
[0075] S402: Setting the weight of each information factor in the detailed information of the document by using the rank sum ratio method, and calculating the weighted rank sum ratio of each information element. The calculation formula of the weighted rank sum ratio is specifically:
[0076]
[0077] Where, WRSRi is the weighted rank sum ratio of the i-th paper; n is the total number of papers, i∈[1,n]; p is the total number of information factors, j∈[1,p]; W j is the weight of the jth evaluation index; R ij is the rank of the information factor in the jth column of the ith paper.
[0078] Furthermore, the weights of each information factor in the detailed information of the document are set by the rank sum ratio method as follows:
[0079] List the original data: one row represents an evaluation object, one column represents an evaluation indicator, and construct a data matrix (n×m).
[0080] Sort each evaluation index by its value to get the rank R. 3. Calculate the RSR value:
[0081] The RSR value and RSR value ranking are obtained by using the rank value calculation.
[0082] List the distribution table of RSR: including the frequency of each group, cumulative frequency, average rank, etc., and calculate the downward cumulative frequency.
[0083] Linear regression was performed using RSR distribution values as independent variables and Probit values as dependent variables.
[0084] The evaluation objects are ranked according to the RSR estimated values corresponding to the regression equation.
[0085] S403: Re-rank the documents obtained by the initial screening by combining the weighted rank sum ratio and the score obtained by adding the similarities, so as to obtain recommended related papers.
[0086] S50: Based on the re-ranked related papers, select the TOP-N related papers for in-depth search and re-recommendation.
[0087] As a preferred implementation of this embodiment, step S50 is specifically as follows:
[0088] According to the re-ranked related papers, select the TOP-N related papers and obtain all the references and citations of the TOP-N related papers, select the omitted documents, re-rank them according to the weighted rank sum ratio and similarity, and then screen the top M documents for re-recommendation. The re-ranking is consistent with the re-ranking method in step S40. M is the preset number of recommended documents. The method of obtaining references and citations is not limited here. In this embodiment, crawler technology is preferred for collection.
[0089] S51: The user selects a paper from all the recommended papers to read, and the user is assisted in understanding the paper by asking questions based on a large language model.
[0090] As a preferred implementation of this embodiment, step S51 is specifically as follows:
[0091] S511: According to the semantic information and integrity of the paper, the document content is divided into several text blocks. The text block method is not limited here. In this embodiment, the large language model based on the llamaindex development framework is preferably used to divide the text into blocks.
[0092] S512: Calculate the similarity between the text block and the user's question based on the question raised by the user.
[0093] S513: According to the calculated similarity, TOP-N text blocks are taken out as background information of the user question and input into the large language model, and different keywords are set to require the large language model to answer the user question.
[0094] S514: Summarize the answers to the questions obtained by the large language model based on different text blocks, so as to obtain the final answer to the question, and assist the user in understanding the paper being consulted.
[0095] S52: Obtain the TOP-N related papers in step S50, extract relevant and useful information from the papers and input it into the large language model, and generate an academic report for user reference.
[0096] As a preferred implementation of this embodiment, step S52 is specifically as follows:
[0097] S521: Obtain the TOP-N relevant papers in step S50, and extract information from the papers in blocks according to the parts generated according to the requirements of the academic report template. The information includes corresponding topic background information, research progress in the current field, and outstanding contributions of the experiment.
[0098] S522: Based on the set keywords, the obtained information is further purified through a large language model to prevent harmful information from affecting the final output result.
[0099] Furthermore, the purification method includes a refining operation and a filtering operation. The refining operation is specifically as follows: combining the set keywords, requiring the large language model to combine the question and information, refining useful information, and removing irrelevant information. The core idea is to generate a summary that integrates information from multiple retrieved documents to support the generation of the target output. The generated summary may contain expressions that are not in the original document, but these expressions are semantically equivalent to the original information, providing greater flexibility. The filtering operation is specifically as follows: selecting the most relevant parts (sentences or phrases) from the retrieved documents, which directly support the generation of the target output; evaluating the relevance between each retrieved sentence and the target output or question; and selecting the most relevant one or several sentences as the summary based on the relevance score. These sentences directly contain the information required to generate the target output. Since the filtering operation only selects the most relevant content, the generated summary is usually shorter, reducing the amount of context processed by the language model. At the same time, the extracted content comes directly from the original document, so the summary remains faithful to the original information.
[0100] S523: The purified information is combined with the user's question, and combined with the set academic report keywords, and input into the large language model to generate an academic report for the user's reference. The academic report may include an introduction, a literature review, an abstract, and the like.
[0101] As a preferred implementation of this embodiment, the method described in this embodiment also includes:
[0102] S53: Obtain the TOP-N related papers in step S50, extract the experimental settings in the papers and input them into the large language model, and generate other feasible solutions for researchers' reference.
[0103] The step S53 is specifically as follows:
[0104] S531: Obtain the TOP-N relevant papers in step S50, and extract the experimental settings involved in the papers, including experimental methods, experimental purposes, and experimental results, etc., based on the obtained relevant papers.
[0105] S532: The experimental settings of relevant papers are passed into the large language model, requiring it to improve the experimental settings based on the existing experimental settings, from the perspective of making up for the shortcomings of the existing experiments and better achieving the experimental objectives, and to propose new innovative solutions.
[0106] S533: Let the large language model determine whether the innovative solution is similar to the previous experimental settings, and verify whether the solution is reasonable and feasible, so as to implement a feedback mechanism, gradually update and optimize the innovative solution proposed by the large language model;
[0107] S534: If the requirements are met, output the final innovative solution; otherwise, repeat steps S531-S533 from TOP-N+1 to TOP-2N papers.
[0108] Embodiment 2
[0109] Accordingly, this embodiment provides a document search and academic report generation system, including a user problem processing module, a document retrieval and preliminary screening module, a deep retrieval module, a paper comprehension assistance module, an academic report generation module and an innovative solution generation module.
[0110] The user question processing module is used to obtain the user questions input by the user in the form of natural language dialogue and rewrite the questions according to the user questions. This module is used to implement the functions of steps S10 and S20 in the first embodiment and will not be repeated here.
[0111] The document retrieval and preliminary screening module is used to obtain relevant documents from the search engine according to the rewritten user question, and calculate the similarity between the document abstract and the user question to perform preliminary screening of the document, obtain detailed information of the document after the preliminary screening, and re-sort the documents obtained by the preliminary screening by combining the rank sum ratio method and the similarity to obtain recommended related papers. This module is used to implement the functions of steps S30 and S40 in Example 1, which will not be repeated here.
[0112] The deep search module is used to select TOP-N related papers for deep search and re-recommendation based on the re-ranked related papers. This module is used to implement the function of step S50 in the first embodiment and will not be described in detail here.
[0113] The paper understanding assistance module is used for users to select papers from all recommended papers for review, and assists users in understanding the reviewed papers based on a large language model in the form of questions. This module is used to implement the function of step S51 in Example 1 and will not be described in detail here.
[0114] The academic report generation module is used to obtain the TOP-N relevant papers selected in the deep search module, extract relevant and useful information from the papers and input it into the large language model, and generate academic reports for user reference. This module is used to implement the function of step S52 in Example 1 and will not be repeated here.
[0115] The innovative solution generation module is used to obtain the TOP-N related papers in step S50, extract the experimental settings in the papers and input them into the large language model, and generate other feasible solutions for researchers to refer to. This module is used to implement the function of step S53 in embodiment 1 and will not be repeated here.
[0116] Embodiment 3
[0117] This embodiment provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the document search and academic report generation method as described in any embodiment of the present invention.
[0118] Embodiment 4
[0119] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for document search and academic report generation as described in any embodiment of the present invention is implemented.
[0120] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.
[0121] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0122] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0123] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.
[0124] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for document search and academic report generation, characterized in that: The following steps are involved: S10: obtaining a user question input by the user in the form of a natural language dialogue; S20: Rewrite the question according to the user's question; S30: Obtain relevant literature from the search engine according to the rewritten user question, and calculate the similarity between the literature abstract and the user question to perform preliminary screening of the literature; S40: Obtain detailed information of the literature after the initial screening, and re-rank the literature obtained by the initial screening by combining the rank sum ratio method and similarity, so as to obtain recommended related papers; S50: Based on the re-ranked related papers, select the TOP-N related papers for in-depth search and re-recommendation; S51: The user selects a paper from all the recommended papers to read, and the user is assisted in understanding the paper by asking questions based on the large language model; S52: Obtain the TOP-N related papers in step S50, extract relevant and useful information from the papers and input it into the large language model, and generate an academic report for user reference.
2. The method for document search and academic report generation according to claim 1, characterized in that: The step S20 is specifically as follows: S201: searching for relevant information about the user's question through a search engine according to the user's question, and inputting the relevant information into a large language model in combination with the user's question to expand the question; S202: Use the knowledge base of the large language model to answer the user's question, and document the expanded user question and the answer of the large language model to obtain the rewritten user question for subsequent similarity calculation.
3. The method for document search and academic report generation according to claim 1, characterized in that: The step S30 is specifically as follows: S301: Obtain relevant literature from a search engine according to the rewritten user question and obtain a literature summary; S302: Generate embedding vectors for two types of text: user questions and document summaries using a pre-trained word embedding model; S303: The language similarity between the two types of text is obtained by calculating the cosine similarity of the embedding vectors between the two types of text, specifically: Where CosineSimilarity() is the cosine similarity; A and B are the embedding vectors of the two texts respectively; · represents the dot product; ‖A‖ and ‖B‖ represent the modulus of the two embedding vectors respectively; S304: Perform a preliminary screening of documents according to a set semantic similarity threshold.
4. The method for document search and academic report generation according to claim 1, characterized in that: The step S40 is specifically as follows: S401: Obtain detailed information on the literature after the initial screening, including the number of citations of the paper, the impact factor of the journal in which the paper is published, and the similarity between the references of the paper; S402: Setting the weight of each information factor in the detailed information of the document by using the rank sum ratio method, and calculating the weighted rank sum ratio of each information element. The calculation formula of the weighted rank sum ratio is specifically: Where, WRSR i is the weighted rank sum ratio of the i-th paper; n is the total number of papers, i∈[1,n]; p is the total number of information factors, j∈[1,p]; W j is the weight of the jth evaluation index; R ij is the rank of the information factor in the jth column of the ith paper; S403: Re-rank the documents obtained by the initial screening by combining the weighted rank sum ratio and the score obtained by adding the similarities, so as to obtain recommended related papers.
5. The method for document search and academic report generation according to claim 1, characterized in that: The step S50 is specifically as follows: According to the re-ranked related papers, select the TOP-N related papers and obtain all the references and citations of the TOP-N related papers. Select the omitted papers and re-rank them according to the weighted rank sum ratio and similarity, and then screen the top M papers for re-recommendation.
6. The method for document search and academic report generation according to claim 1, characterized in that: The step S51 is specifically as follows: S511: According to the semantic information and completeness of the paper, the document content is divided into several text blocks; S512: Calculate the similarity between the text block and the user's question according to the question raised by the user; S513: taking out TOP-N text blocks according to the calculated similarity as the background information of the user's question and inputting them into the large language model, and setting different keywords to require the large language model to answer the user's question; S514: Summarize the answers to the questions obtained by the large language model based on different text blocks, so as to obtain the final answer to the question, and assist the user in understanding the paper being consulted.
7. The method for document search and academic report generation according to claim 1, characterized in that: The step S52 is specifically as follows: S521: Obtain the TOP-N related papers in step S50, and extract information from the papers in blocks according to the parts generated by the academic report template, wherein the information includes corresponding topic background information, research progress in the current field, and outstanding contributions of the experiment; S522: Combine the set keywords and purify the obtained information again through the large language model to prevent harmful information from affecting the final output result; S523: The purified information is merged with the user's question, and combined with the set academic report keywords, the information is passed into the large language model to generate an academic report for user reference.
8. A document search and academic report generation system, characterized in that: The system is used to implement the document search and academic report generation method as described in any one of claims 1 to 7, including a user question processing module, a document retrieval and preliminary screening module, a deep retrieval module, a paper understanding auxiliary module and an academic report generation module; A user question processing module is used to obtain user questions input by users in the form of natural language dialogue and rewrite the questions according to the user questions; The literature search and preliminary screening module is used to obtain relevant literature from the search engine according to the rewritten user questions, calculate the similarity between the literature abstract and the user questions to perform preliminary screening of the literature, obtain detailed information of the literature after the preliminary screening, and re-rank the literature obtained by the preliminary screening by combining the rank sum ratio method and similarity, so as to obtain recommended related papers; The deep search module is used to select the TOP-N related papers for deep search and re-recommendation based on the re-ranked related papers; The paper understanding assistance module is used for users to select papers from all recommended papers for review, and assists users in understanding the reviewed papers by asking questions based on a large language model; The academic report generation module is used to obtain the TOP-N relevant papers selected in the deep search module, extract relevant and useful information from the papers and input it into the large language model to generate academic reports for user reference.
9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the document search and academic report generation method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for document search and academic report generation as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Subject-based document retrieval prediction method
WO2024078141A1
Cited By
Literature review generation method, electronic equipment, storage medium and program product
CN120781846A
Document review generation method, electronic device, storage medium, and program product
CN120781846B
Document deep traceability system based on multi-agent cooperation
CN121071126A