Document retrieval method, device, equipment and storage medium
By using large language models for semantic understanding and topic information generation, the problem of inaccurate keyword selection and search-based construction in existing literature search is solved, and a higher-quality literature search effect is achieved.
Patent Information
- Application Number
- CN202411261363.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-09
AI Technical Summary
In the existing literature search methods, improper keyword selection or inaccurate search method construction leads to poor search quality and it is difficult to effectively improve the accuracy and efficiency of literature search.
Using a method based on a pre-constructed large language model, we understand the search information input by the user through semantics, generate topic information related to the search information, including topics, topic keywords, short descriptions of topics and subjects involved, and literature search is carried out based on these topic information.
The quality of literature search has been improved, and through accurate generation and utilization of subject information, high-quality literature related to the search information can be more effectively screened out.
Smart Images

Figure CN119066154B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a document retrieval method, device, equipment and storage medium. Background Art
[0002] Literature retrieval refers to the process of screening and obtaining literature related to the research topic through a series of methods and techniques in the process of academic research. It is an indispensable part of scientific research and is of great significance for achieving research goals and improving research quality.
[0003] At present, when conducting a literature search, it is necessary to first clarify the search subject and keywords, select appropriate academic databases and resources, and use Boolean logic, phrase search and other strategies to conduct a preliminary extensive search. For example, under the theme of "cancer treatment", keywords and Boolean logic can be used to construct the search formula "cancer AND treatment NOT surgery" for a preliminary search. Next, the preliminary searched documents are screened and filtered according to the title, abstract and keywords of the document to obtain literature materials that are highly relevant to the research topic. However, during the search process, improper keyword selection or inaccurate search formula construction may occur, resulting in poor search quality. Therefore, how to improve the search quality of documents is a technical problem that needs to be urgently solved by technicians in this field. Summary of the invention
[0004] In view of this, the present disclosure proposes a document retrieval method, apparatus, device and storage medium, which can improve the retrieval quality of documents.
[0005] According to a first aspect of the present disclosure, there is provided a document retrieval method, comprising:
[0006] Get the search information entered by the user;
[0007] Performing semantic understanding on the search information based on a pre-built first language model to generate topic information related to the search information;
[0008] Performing a literature search based on the subject information to obtain literature materials related to the search information;
[0009] The subject information includes at least one of the subject, subject keywords, a brief description of the subject, and the subjects involved in the subject.
[0010] In a possible implementation, when semantic understanding of the search information is performed based on the pre-built first language model to generate topic information related to the search information, the method includes:
[0011] Performing semantic analysis on the search information to obtain key semantic elements corresponding to the search information;
[0012] Performing context analysis on the search information to obtain context information corresponding to the search information;
[0013] Based on the key semantic elements and the context information, constructing a semantic relationship corresponding to the search information;
[0014] Based on the semantic relationship, topic information related to the search information is generated.
[0015] In a possible implementation, when there is multiple subject information related to the search information, the method further includes:
[0016] Calculate the topic score corresponding to each topic information;
[0017] Based on the topic scores corresponding to each topic information, target topic information is screened out from multiple topic information;
[0018] A document search is performed based on the target subject information to obtain document materials related to the search information.
[0019] In a possible implementation, when performing a literature search based on the subject information to obtain literature materials related to the search information, the process includes:
[0020] Based on the subject information, generating a general search formula corresponding to the subject information;
[0021] Based on the general search formula, generating a target search formula for searching a target data source;
[0022] Based on the target search formula, literature materials related to the search information are retrieved from the target data source.
[0023] In a possible implementation, when generating a general search formula corresponding to the subject information based on the subject information, the process includes:
[0024] Extracting topic keywords from the topic information;
[0025] Obtaining standard subject keywords, synonyms and near-synonyms corresponding to the subject keywords, as well as limited fields and time ranges set by the user;
[0026] Based on the standard subject keywords, synonyms and antonyms corresponding to the subject keywords and the limited fields and time ranges set by the user, a preset search formula generation algorithm is used to generate a general search formula corresponding to the subject information.
[0027] In a possible implementation, after obtaining the literature related to the search information, the method further includes:
[0028] Based on a preset document metadata template, construct metadata corresponding to each of the document materials;
[0029] Deduplication of each of the document materials is performed based on the metadata corresponding to each of the document materials;
[0030] From the duplicated document data, a set number of document data with the highest comprehensive scores are screened out as final document data related to the search information.
[0031] In a possible implementation, after obtaining the final document materials, the method further includes:
[0032] Analyze each of the final document materials based on the pre-built second language model to obtain a document summary of each of the final document materials;
[0033] Based on the search information, subject information, metadata and document overview related to each of the final document materials, a literature review of each of the final document materials is generated.
[0034] According to a second aspect of the present disclosure, there is provided a document retrieval device, comprising:
[0035] A search information acquisition module is used to acquire search information input by a user;
[0036] A topic generation module, used for performing semantic understanding on the search information based on a pre-built first language model, and generating topic information related to the search information;
[0037] A search module, used to search for documents based on the subject information to obtain documents related to the search information;
[0038] The subject information includes at least one of the subject, subject keywords, a brief description of the subject, and the subjects involved in the subject.
[0039] According to a third aspect of the present disclosure, a document retrieval device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in the first aspect of the present disclosure.
[0040] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions, when executed by a processor, implement the method described in the first aspect of the present disclosure.
[0041] The present disclosure provides a document retrieval method, device, equipment and storage medium, the method comprising: obtaining retrieval information input by a user; semantically understanding the retrieval information based on a pre-built first language model to generate subject information related to the retrieval information; performing document retrieval based on the subject information to obtain document materials related to the retrieval information; wherein the subject information includes at least one of a subject, subject keywords, a brief description of the subject and a subject-related discipline. In the method of the present disclosure, the user only needs to input the retrieval information, and the system can accurately output the subject information related to the retrieval information based on the first language model, and then retrieve high-quality document materials based on the accurate subject information, thereby improving the retrieval quality of the document retrieval.
[0042] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0044] Figure 1 A flowchart of a document retrieval method according to an embodiment of the present disclosure is shown.
[0045] Figure 2 A schematic block diagram of a document retrieval device according to an embodiment of the present disclosure is shown.
[0046] Figure 3 A schematic block diagram of a document retrieval device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0047] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0048] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0049] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0050] <Method Example>
[0051] Figure 1 FIG. 1 is a flowchart of a document retrieval method according to an embodiment of the present disclosure. Figure 1 As shown, the method includes steps S1100-S1300.
[0052] S1100, obtaining search information input by the user. The search information is a text used to describe the key content of the document to be searched. For example, the search information can be an abstract of the document to be searched, or a research topic corresponding to the document to be searched, or other key text content of the document to be searched, which is not specifically limited here.
[0053] In a possible implementation, after obtaining the search information input by the user, the search information will also be subjected to at least one preprocessing operation of text cleaning, language detection and translation, standardization, stop word filtering, format verification, sensitive information processing and length control. Among them, text cleaning: used to remove special characters, extra spaces and line breaks in the search information; used to unify the encoding format of the search information, for example, to unify the encoding format of the search information into UTF-8; and also used to correct common spelling errors in the search information. Language detection and translation: used to detect the language used in the search information, and when the language used is not the system default language, provide a translation option to facilitate the user to translate the input search information into the system default language. Standardization: used to detect the uppercase letters included in the search information, determine whether the detected uppercase letters meet the preset conversion conditions, and convert them to the corresponding lowercase letters if they meet the conversion conditions; used to detect dates, numbers, etc. in the search information, and convert the tested dates and numbers into standard formats. Stop word filtering: used to find stop words in the search information according to a preset stop word list, and delete the retrieved stop words. Sensitive information processing: used to identify whether the search information contains pre-set sensitive information. If so, the sensitive information is desensitized. Length control: used to detect the text length of the search information and determine whether the text length is greater than the pre-set length threshold. If it is greater than the length threshold, the search information is segmented to obtain multiple search information segments that meet the length threshold, and in the subsequent search process, each search information segment is searched separately.
[0054] By performing at least one of the above preprocessing operations, the search information input by the user can be standardized, thereby improving the accuracy and efficiency of subsequent searches.
[0055] S1200: semantically understand the search information based on the pre-built first language model, and generate subject information related to the search information, wherein the subject information includes at least one of a subject, subject keywords, a brief description of the subject, and a subject-related discipline.
[0056] In a possible implementation, when semantic understanding of the search information is performed based on the pre-built first language model to generate topic information related to the search information, the following steps may be included:
[0057] First, semantic analysis is performed on the search information to obtain the key semantic elements corresponding to the search information. Among them, the key semantic elements are keywords, key phrases or key sentences that reflect the core content of the search information. The key semantic elements include the key semantic elements of the main topic and the key semantic elements of the action or goal.
[0058] Specifically, in order to enable the first language model to have this capability, a set number of retrieval information samples need to be acquired in advance. For each retrieval information sample, the key semantic elements of the corresponding main topic and the key semantic elements of the action or target are annotated, thereby obtaining a semantic analysis training sample. After obtaining the semantic analysis training sample corresponding to each retrieval information sample, the first language model is trained using the retrieval information sample in each training sample as input data and the annotated data as standard output data. After the training is completed, the retrieval information is input into the first language model, and the first language model performs the semantic analysis function, that is, it can output the key semantic elements corresponding to the retrieval information.
[0059] For example, when the user inputs the search information "Explore the application of artificial intelligence in climate change prediction, especially how machine learning models can improve the prediction accuracy of long-term climate patterns", the first language model performs semantic analysis and obtains the following key semantic elements:
[0060] - Main topics: Artificial intelligence, climate change prediction
[0061] - Action / Goal: Explore applications, improve predictions
[0062] Second, context analysis is performed on the search information to obtain context information corresponding to the search information, including research fields, research objectives, research scope, technical focus, potential challenges and other context information corresponding to the project.
[0063] Continuing with the above embodiment, when the search information input by the user is "explore the application of artificial intelligence in climate change prediction, especially how machine learning models can improve the prediction accuracy of long-term climate patterns", the first language model performs a context analysis function and obtains the following context information:
[0064] - Research field: Interdisciplinary (Artificial Intelligence + Climate Science)
[0065] - Research objective: Apply AI technology to improve the accuracy of climate prediction
[0066] - Research scope: Focuses on long-term climate patterns rather than short-term weather forecasts
[0067] - Technical focus: Special attention to machine learning models
[0068] - Potential challenges: Complexity in dealing with long-term climate data
[0069] Third, based on key semantic elements and contextual information, the semantic relationship corresponding to the retrieved information is constructed.
[0070] Specifically, in order to enable the first language model to have the ability to analyze semantic relationships, the following steps need to be performed:
[0071] 1. Obtain a certain amount of search information samples.
[0072] 2. Label each sample:
[0073] a) Marking key semantic elements
[0074] b) Annotate context information
[0075] c) Annotate the semantic relationship between key semantic elements and contextual information
[0076] 3. Build training dataset:
[0077] - Input data: retrieve key semantic elements and contextual information from information samples
[0078] - Output data: corresponding semantic relationship annotation
[0079] 4. Use the constructed training dataset to fine-tune the first language model so that it can perform semantic relationship analysis functions.
[0080] 5. After training is completed, the model usage process is as follows:
[0081] a) Input: key semantic elements and context information corresponding to the retrieval information
[0082] b) Processing: The model performs semantic relationship analysis
[0083] c) Output: semantic relationship corresponding to the retrieved information
[0084] In this way, the first language model will be able to analyze the key semantic elements and contextual information of the input and output the semantic relations between them, thus providing a basis for subsequent topic information generation.
[0085] Continuing with the above embodiment, when the key semantic elements and context information corresponding to the search information “explore the application of artificial intelligence in climate change prediction, especially how machine learning models can improve the prediction accuracy of long-term climate patterns” are as described above, the semantic relationship constructed by the first language model will be as follows:
[0086] - Artificial intelligence as a tool for climate change prediction
[0087] - Machine learning models are a subset of artificial intelligence
[0088] - Forecast accuracy is a goal that needs to be improved
[0089] - Long-term climate models are the subject of predictions
[0090] Fourth, based on semantic relations, generate topic information related to the retrieved information.
[0091] Specifically, first generate multiple topics related to the retrieval information based on the generated semantic relations. Traverse each topic, and for the current topic that has been traversed, perform semantic understanding of the current topic to obtain the topic keywords of the current topic; then, based on the current topic and the topic keywords of the current topic, identify the research purpose, main research method or direction, expected research results, potential research contribution, and the disciplines involved in the topic; then, fill the identified research purpose, main research method or direction, expected research results, and potential research contribution into the pre-built description template to generate a brief description of the topic of the current topic. At this point, all the topic information of the current topic is obtained. After the traversal is completed, the topic information of each topic can be obtained. Among them, the description template is as follows:
[0092] "This study aims to [research purpose] and [expected research results] through [main research method or direction]. This work is expected to [potential research contribution]."
[0093] In order to enable the large language model to have the functions of the above steps, it is also necessary to refer to the above construction of training data for each step (wherein the training data includes the input data of each step and the standard output data annotated for the input data), so that the first large language model can perform the functions of the above steps through training with the training data. Continuing with the above embodiment, when the semantic relationship shown above is constructed, the generated topic information related to the search information is as follows:
[0094] a) - Topic: "Comparative Study of Machine Learning Algorithms in Long-term Climate Model Prediction"
[0095] - Topic keywords: machine learning, climate model, comparison of prediction algorithms
[0096] - Topic Short Description: This study aims to compare the performance of different machine learning algorithms in long-term climate model predictions and improve the accuracy of long-term climate predictions by applying a variety of advanced machine learning techniques. This work is expected to provide climate scientists with more reliable prediction tools and enhance our understanding of the complex climate system.
[0097] - Subjects covered: artificial intelligence, climate science, statistics
[0098] b) - Topic: "Deep Learning Networks for Improving the Accuracy of Global Climate Models"
[0099] - Topic keywords: deep learning, global climate model, accuracy improvement
[0100] - Topic Short Description: This research aims to explore the potential of deep learning networks to improve the accuracy of global climate models. By integrating advanced deep learning techniques into existing climate models, we aim to significantly improve the accuracy of long-term climate predictions. This work is expected to provide more precise tools for climate change research and support more effective environmental policy making.
[0101] - Subjects covered: Deep learning, climate modeling, big data analysis
[0102] c) - Topic: "AI-driven climate change prediction model based on time series analysis"
[0103] - Theme keywords: time series analysis, AI, climate change prediction
[0104] - Topic Short Description: This research aims to develop an AI-driven climate change prediction model based on time series analysis. By combining advanced artificial intelligence techniques with traditional time series analysis methods, we expect to be able to more accurately capture and predict long-term climate change trends. This work is expected to provide a more precise and reliable tool for understanding and predicting the complex dynamics of the climate system.
[0105] - Subjects covered: Artificial Intelligence, Time Series Analysis, Climate Science, Data Science
[0106] d) - Theme: "Application of Multi-source Data Fusion in AI-assisted Climate Prediction"
[0107] - Keywords: data fusion, AI, climate prediction, multi-source data
[0108] - Short description of the topic: This study explores the application of multi-source data fusion technology in AI-assisted climate prediction. By integrating data from multiple sources such as satellites, ground stations, and ocean buoys, and analyzing them using advanced AI algorithms, we aim to improve the comprehensiveness and accuracy of long-term climate model predictions. This work is expected to provide climate scientists with a richer and more reliable data foundation, thereby improving the prediction and understanding of climate change.
[0109] - Subjects covered: Data fusion, artificial intelligence, climatology, geoinformation science
[0110] In a possible implementation, after obtaining the brief description of each topic, at least one preprocessing operation of language optimization, length adjustment and consistency check is performed on the brief description of each topic.
[0111] When optimizing the language of the brief description of each topic, natural language processing technology can be used to optimize the language expression to ensure the fluency and professionalism of the brief description of each topic.
[0112] When adjusting the length of each topic brief description, it includes: judging whether the word count of the topic brief description is within the preset word count range, and if it is not within the preset word count range, adjusting the topic brief description so that the adjusted word count is within the preset word count range. The word count range may be (lower word count limit, upper word count limit), and the lower word count limit and the upper word count limit may be set according to the specific application scenario. For example, the lower word count limit may be set to 50, and the upper word count limit may be set to 100. When the word count of the topic brief description is lower than the lower word count limit, the word count is adjusted to the preset word count range by adding relevant details of the topic brief description; when the word count of the topic brief description is higher than the upper word count limit, the topic brief description is compressed by using the summary method so that the word count is adjusted to the preset word count range.
[0113] In a possible implementation, the following steps may be included when adding details related to the topic brief description:
[0114] First, calculate the length of the topic brief description, and determine the number of words to be added based on the length of the topic brief description and the preset word range.
[0115] Second, analyze the possible directions for expansion in the brief description of the topic, which include at least one of the specific steps of the main research method, more details of the expected research results, potential research impact or application, and connection with existing research.
[0116] Third, the number of words to be added, the direction that can be expanded, and the reference information are input into the pre-trained large language model, so that the large language model can determine the detailed information related to the short description of the topic that can be added. The reference information includes at least one of the current short description, the topic corresponding to the current short description, the topic keywords, the subject involved in the topic, the search information input by the user, and the context information.
[0117] Fourth, the detailed information related to the topic brief description is integrated with the topic brief description to obtain an expanded topic brief description.
[0118] Furthermore, after obtaining the expanded topic brief description, it also includes operations such as verifying whether the word count of the expanded topic brief description is within a preset word count range, whether the content is coherent, and whether it is consistent with the original topic brief description, so as to improve the reliability of the expanded topic brief description.
[0119] For example, assuming that the lower limit of the word count is 100 words, the original topic brief description is "This study aims to develop an AI-driven climate change prediction model based on time series analysis. By combining advanced artificial intelligence technology and traditional time series analysis methods, we expect to be able to more accurately capture and predict long-term climate change trends." A total of 80 words, in this embodiment, after expansion by the above method, a 120-word expanded brief description will be obtained "This study aims to develop an AI-driven climate change prediction model based on time series analysis. By combining advanced artificial intelligence technology and traditional time series analysis methods, we expect to be able to more accurately capture and predict long-term climate change trends. The model will integrate multi-source climate data, including satellite observations, ground station records, and historical climate archives, and use deep learning algorithms to extract complex time-dependent features. This innovative method is expected to improve the accuracy and time span of climate predictions, and provide a reliable basis for formulating climate adaptation strategies."
[0120] The method of this embodiment can effectively add relevant details such as the specific steps of the research method, more details of the expected research results, potential research impact or application, and connection with existing research, so that the content of the brief description is richer and the preset word count requirement can be met.
[0121] When checking the consistency of each topic brief description, it includes: checking whether the topic brief description is consistent with the topic and topic keywords. If there is inconsistency, adjust the topic brief description.
[0122] In a possible implementation, when detecting whether the topic brief description is consistent with the topic and the topic keywords, evaluation may be performed from the following aspects:
[0123] First, perform keyword matching to check whether the topic keywords appear in the topic brief description.
[0124] Second, semantic similarity analysis is performed to calculate the similarity between the topic keywords and the topic brief description.
[0125] Third, conduct a topic consistency check to check whether the topic brief description reflects the topic.
[0126] Fourth, conduct a contextual relevance check to check whether the topic brief description covers the core content of the topic.
[0127] Fifth, subject matching, check whether the brief description of the topic reflects the relevant disciplines involved in the topic.
[0128] The score of the topic short description is determined based on the evaluation results of the above items, and it is judged whether the score of the topic short description is greater than the preset score threshold. When it is greater than or equal to the preset score threshold, it is judged that the topic short description is consistent with the topic and topic keywords; when it is less than the preset score threshold, it is judged that the topic short description is inconsistent with the topic and topic keywords.
[0129] In the event that the subject brief description is judged to be inconsistent with the subject and subject keywords, the subject brief description will be adjusted to ensure that it is consistent with the subject and subject keywords, thereby improving the accuracy of subsequent literature retrieval.
[0130] In a possible implementation, in order to improve the search efficiency, when multiple subject information related to the search information is obtained, a preset number of subject information with the highest relevance to the search information is screened out from the multiple subject information as target subject information, and the target subject information is used to search for literature in subsequent search steps. The specific steps are as follows:
[0131] First, the topic score corresponding to each topic information is calculated. The topic score is calculated based on the similarity score, keyword matching score, subject relevance score, novelty score and interdisciplinary score.
[0132] The similarity score is used to evaluate the similarity between the subject in the subject information and the search information. The more similar, the higher the score. The calculation formula of the similarity score is as follows:
[0133] Similarity_Score=F1*cosine_similarity(user_input_vector,topic_vector)
[0134] Where, Similarity_Score is the similarity score, F1 is the highest score of the preset similarity score, cosine_similarity() is the cosine similarity function, user_input_vector is the vector representation of the search information entered by the user, and topic_vector is the vector representation of the topic in the topic information.
[0135] The keyword matching score is used to evaluate the keyword matching between the subject keywords in the subject information and the keywords specified by the user for the search information. The more matching keywords, the higher the score. The calculation formula of the keyword matching score is as follows:
[0136] Keyword_Score = F2* (matched_keywords / total_keywords) * (1 + log(1+ matched_importance))
[0137] In the formula, Keyword_Score is the keyword matching score, F2 is the maximum score of the preset keyword matching score, matched_keywords is the number of matched keywords, total_keywords is the total number of keywords specified by the user for the search information, matched_importance is the sum of the importance scores of the matched keywords, and the importance score of each matched keyword can be calculated using its corresponding TF-IDF value.
[0138] The subject relevance score is used to evaluate the match between the subject involved and the subject of interest specified by the user when entering the search information. The higher the match, the higher the score. The calculation formula for the subject relevance score is as follows:
[0139] Discipline_Score = F3 * (matched_disciplines / max(user_disciplines,topic_disciplines))
[0140] In the formula, Discipline_Score is the discipline relevance score, F3 is the highest score of the preset discipline relevance score, matched_disciplines is the number of matches between the subject-related disciplines and the subjects of interest, user_disciplines is the number of subjects of interest specified by the user, and topic_disciplines is the number of subjects involved in the topic.
[0141] The novelty score is used to assess the novelty of the topic compared to the title or abstract of a recent journal article. The higher the score, the higher the novelty. The novelty score calculation formula is as follows:
[0142] Novelty_Score = F4 * (1 - similarity_to_existing_research)
[0143] Where Novelty_Score is the novelty score, F4 is the highest score of the preset novelty score, and similarity_to_existing_research is the highest similarity between the topic and the title or abstract of the most recent journal article.
[0144] The interdisciplinary score is used to assess whether the topic involves multiple disciplines. The more disciplines involved, the higher the score. The formula for calculating the interdisciplinary score is as follows:
[0145] Interdisciplinary_Score =F5 * (1 - 1 / log(1 + num_disciplines))
[0146] In the formula, Interdisciplinary_Score is the interdisciplinary score, F1 is the highest score of the preset interdisciplinary score, and num_disciplines is the number of disciplines involved in the topic.
[0147] It should be noted that F1, F2, F3, F4 and F5 are determined according to the importance of the evaluation items in the subject scoring process. The higher the importance, the greater the corresponding score, and the total score is 100. In a specific example, F1 can be set to 40 points, F2 to 25 points, F3 to 20 points, F4 to 10 points, and F5 to 5 points.
[0148] The calculation formula of the corresponding topic score can be shown as follows:
[0149] Total_Score = Similarity_Score + Keyword_Score + Discipline_Score +Novelty_Score + Interdisciplinary_Score
[0150] In the formula, Total_Score is the topic score.
[0151] The above calculation formula can be used to calculate the topic score corresponding to each topic information.
[0152] Second, based on the topic scores corresponding to each topic information, target topic information is screened out from multiple topic information. Specifically, each topic information can be sorted in descending order according to the topic scores, and N1 topic information with the highest topic scores are screened out as target topic information. The number of N1 can be set according to the specific scenario, and preferably, the value range of N1 can be set to 3 to 5.
[0153] Third, conduct literature search based on the target subject information to obtain literature materials related to the search information.
[0154] Furthermore, after obtaining the target topic information, each target topic information and the topic score corresponding to each target topic information are organized into structured data in JSON format to facilitate the storage and reading of each target topic information. The structured data in JSON format is as follows:
[0155] {
[0156] "topic_id": "Unique identifier",
[0157] "main_topic": "Topic",
[0158] "keywords": ["Topic keyword 1", "Topic keyword 2", "..."],
[0159] "description": "Short description of the topic",
[0160] "relevance_score": topic score
[0161] "discipline": ["Topic involves discipline 1", "Topic involves discipline 2"],
[0162] }
[0163] S1300, perform document search based on subject information to obtain document materials related to the search information. It should be noted that the steps of performing document search based on subject information and performing document search based on target subject information are the same, and the following takes performing document search based on subject information as an example to describe the steps in detail.
[0164] In a possible implementation, when performing a document search based on subject information to obtain document materials related to the search information, the following steps may be included:
[0165] First, based on the subject information, a general search formula corresponding to the subject information is generated. Specifically, the subject keywords are first extracted from the subject information; then the standard subject keywords, synonyms and near synonyms corresponding to the subject keywords, as well as the limited fields and time ranges set by the user are obtained; then, based on the standard subject keywords, synonyms and near synonyms corresponding to the subject keywords, as well as the limited fields and time ranges set by the user, a preset search formula generation algorithm is used to generate a general search formula corresponding to the subject information. Among them, the search formula generation algorithm is used to use Boolean operators to connect the standard subject keywords, synonyms and near synonyms corresponding to the subject keywords, as well as the limited fields and time ranges set by the user, so as to generate a general search formula corresponding to the subject information.
[0166] In a possible implementation, the system includes a plurality of search formula generation algorithms of different complexity. The higher the complexity, the more complex the generated general search formula; the lower the complexity, the simpler the generated general search formula. In order to ensure that the complexity of the generated general search formula is within a reasonable range, after the general search formula is generated, the following steps are further included:
[0167] First, the complexity score of the general search formula is calculated. Specifically, the calculation formula of the complexity score can be as follows:
[0168] Complexity_Score = (Keyword_Count * F6) + (Operator_Count * F7) +(Nesting_Level * F8)
[0169] In the formula, Complexity_Score is the complexity score, Keyword_Count is the number of keywords in the general search formula, which is equal to the sum of the number of standard subject keywords, synonyms and near synonyms corresponding to the subject keywords, Operator_Count is the number of Boolean operators, Nesting_Level is the highest nesting level of brackets in the general search formula, F6 is the weight coefficient corresponding to the number of keywords, F7 is the weight coefficient corresponding to the number of operators, and F8 is the weight coefficient corresponding to the highest nesting level. Among them, F6, F7 and F8 can be set according to the specific application scenario. For example, F6 can be set to 0.5, F7 to 0.3, and F8 to 0.2.
[0170] The complexity score of the general search formula can be calculated according to the above calculation formula.
[0171] Secondly, it is determined whether the complexity score of the general search formula is within the preset complexity range. The preset load symbol range may be (complexity lower limit, complexity upper limit), and the complexity lower limit and complexity upper limit may be set according to the specific application scenario. For example, the complexity lower limit may be set to 3, and the complexity upper limit may be set to 15.
[0172] Next, when it is determined that the complexity score of the general search formula is not within the preset complexity range, the general search formula is adjusted until the complexity score of the adjusted general search formula is within the preset complexity range. It should be noted here that when the complexity score of the general search formula is lower than the lower complexity limit, it means that the search of the general search formula is too broad, and when the complexity score of the general search formula is higher than the upper complexity limit, it means that the search of the general search formula is too narrow. Therefore, it is necessary to control the complexity score of the general search formula within the preset complexity range.
[0173] When it is determined that the complexity score of a general search formula is lower than the complexity lower limit, it is necessary to select a search formula generation algorithm with a higher level of complexity than the currently executed search formula generation algorithm to regenerate the general search formula, and repeat the above-mentioned general search formula complexity evaluation process until the complexity score of the generated general search formula is within a reasonable range.
[0174] When it is determined that the complexity score of a general search formula is higher than the complexity upper limit, it is necessary to select a search formula generation algorithm with a lower complexity level than the currently executed search formula generation algorithm to regenerate the general search formula, and repeat the above general search formula complexity evaluation process until the complexity score of the generated general search formula is within a reasonable range.
[0175] In order to ensure the accuracy of the general search formula, in a possible implementation method, when the complexity score of the general search formula is within a reasonable range, the general search formula will also be used for retrieval testing to obtain the retrieval results corresponding to the general search formula, and evaluate whether the retrieval results meet the preset fine-tuning conditions. If the fine-tuning conditions are met, the general search formula will be input into a pre-built general search formula fine-tuning model to fine-tune the general search formula through the retrieval formula fine-tuning model, and based on the fine-tuned general search formula, a target search formula for retrieving the target data source is generated.
[0176] In this embodiment, the fine-tuning condition may include at least one of the following:
[0177] 1. The number of search results is not ideal:
[0178] - Too many: usually more than 1,000 results (can be adjusted based on specific needs).
[0179] - Too few: Usually less than 50 results (can be adjusted according to specific needs).
[0180] 2. Low relevance of search results: Through sampling inspection, the proportion of documents relevant to the search information in the retrieved literature is less than 60%.
[0181] 3. Uneven distribution of search results by time: Within the time range set by the user, there are too many or too few documents in a specific time period.
[0182] 4. Insufficient coverage of specific fields or methods: The retrieved literature does not cover the main research methods or directions in the brief description of the topic.
[0183] It should be noted here that before performing fine-tuning of a general search formula, it is necessary to first construct a general search formula fine-tuning model with the general search formula fine-tuning function. Specifically, a set number of general search formula samples are obtained, and for each general search formula sample, the fine-tuned general search formula is annotated, thereby obtaining a search formula fine-tuning training sample. After obtaining all the search formula fine-tuning training samples, the general search formula samples in each search formula fine-tuning training sample are used as input data, and the annotated fine-tuned general search formula is used as standard output data to train the second largest language model. After the training is completed, a general search formula fine-tuning model constructed based on the second largest language model can be obtained.
[0184] Second, based on the general search formula, a target search formula is generated for searching the target data source.
[0185] It should be noted that, in order to meet the user's preference for data sources, when the user enters the search information, a data source list will be provided to the user, and the data source list includes multiple optional data sources. The user can select at least one data source of interest as the target data source from the data source list. The data source list may include the following optional data sources:
[0186] 1. PubMed
[0187] - Description: Biomedical literature database provided by the National Center for Biotechnology Information (NCBI)
[0188] - Main fields: life sciences, biomedicine, clinical medicine
[0189] - Features: Provides a large number of peer-reviewed academic articles and is updated frequently
[0190] 2. ArXiv
[0191] - Description: Open access preprint repository operated by Cornell University
[0192] - Main fields: Physics, Mathematics, Computer Science, Quantitative Biology, Quantitative Finance, Statistics
[0193] - Features: Contains a large number of latest research results, but some articles may not have been peer-reviewed
[0194] 3. Crossref
[0195] - Description: Non-profit academic publishing registration organization
[0196] - Main field: Interdisciplinary
[0197] - Features: Provides DOI (Digital Object Identifier) and metadata services with wide coverage
[0198] 4. EuropePMC
[0199] - Description: European life science and biomedical literature database
[0200] - Main fields: Life sciences, biomedicine
[0201] - Features: In addition to the full text of the article, it also provides data mining services
[0202] 5. PubMed Central (PMC)
[0203] - Description: Free full-text archive provided by the National Institutes of Health (NIH)
[0204] - Main field: Biomedicine and life sciences
[0205] - Features: Provides a large number of open access full-text articles
[0206] 6. DOAJ (Directory of Open Access Journals)
[0207] - Description: Directory of open access journals
[0208] - Main field: Interdisciplinary
[0209] - Features: Focus on high-quality, open access, peer-reviewed journals
[0210] 7. BioMed Central
[0211] - Description: Open Access Publisher
[0212] - Main fields: Biology, Medicine
[0213] - Features: All articles are open access, focusing on rapid publication and innovative research
[0214] Furthermore, the system sets different target search formula generation algorithms for each data source. After determining the target data source that the user is interested in, the target search formula generation algorithm corresponding to the target data source will be used to convert the general search formula into a target search formula for searching the target data source.
[0215] When there are multiple target data sources that the user is interested in, the target search formula generation algorithm corresponding to each target data source will be used to generate a target search formula for searching each target data source. That is, if there are several target data sources, several target search formulas will be generated, and then based on different target search formulas, relevant literature will be retrieved in different target data sources.
[0216] Third, based on the target search formula, literature related to the search information is retrieved from the target data source.
[0217] In a possible implementation, when there are multiple target search terms, an asynchronous search mechanism will be used to search for literature. The asynchronous search mechanism is a search method that does not block the execution of the main program. It allows the system to send search requests to multiple data sources at the same time without waiting for each request to be completed before proceeding to the next one. Specifically, the system creates an asynchronous search task for each target data source based on each target search term. All tasks are started at the same time, and the system does not wait for all tasks to be completed, but processes the results immediately when any task is completed. When all tasks are completed or the set timeout period is reached, the system summarizes all the results obtained and returns the summarized results to the user.
[0218] The asynchronous retrieval mechanism can significantly improve the system's response speed and efficiency, especially when dealing with multiple data sources with different possible response times. It allows the system to better utilize available resources and provide a faster retrieval experience.
[0219] In a possible implementation, after obtaining the literature related to the search information, the following steps may be further included:
[0220] First, based on the preset document metadata template, the metadata corresponding to each document is constructed. The document metadata template includes multiple fields that may appear in the document, such as DOI, title, author, publication year, journal / conference name, abstract, keywords, references and page numbers. For each retrieved document, the field information of each field in the document metadata template is extracted from the document, and the extracted field information is filled into the corresponding document metadata template, so as to obtain the metadata corresponding to each document.
[0221] Second, based on the metadata corresponding to each document, each document is deduplicated. Specifically, the metadata corresponding to each document is traversed, and for the current metadata traversed, the repetition scores between the current metadata and other metadata are scored in turn. When the repetition score between the two is greater than or equal to the preset repetition upper limit, the document corresponding to the other metadata calculated by the parameters and the current metadata are determined as duplicate documents, and the document duplicated with the document corresponding to the current metadata is deleted; when the repetition score between the two is less than the preset repetition lower limit, the document corresponding to the other metadata calculated by the parameters and the current metadata are determined as different documents, and the document different from the document corresponding to the current metadata is retained; when the repetition score between the two is greater than or equal to the preset repetition lower limit and less than the preset repetition upper limit, the document corresponding to the other metadata calculated by the parameters and the current metadata is determined as a possible duplicate document. This time, the user is prompted to review. If it is a duplicate file after review, the document duplicated with the document corresponding to the current metadata is deleted. If it is a different document after review, the document different from the document corresponding to the current metadata is retained. After the traversal is completed, all the remaining document is used as the deduplicated document. The lower limit of the repetition degree and the upper limit of the repetition degree may be set according to a specific application scenario. For example, the lower limit of the repetition degree may be set to 80, and the upper limit of the repetition degree may be set to 90.
[0222] In one possible implementation, the calculation formula of the repetition score is as follows:
[0223] Duplication score = DOI_Score + Title_Score + Author_Score
[0224] In the formula, DOI_Score is the duplication score of DOI in metadata, Title_Score is the duplication score of title in metadata, and Author_Score is the duplication score of author in metadata.
[0225] The calculation formula for the duplication score of this DOI is as follows:
[0226] DOI_Score = F9 * 100 if DOI matches else 0
[0227] In the formula, F9 is the weight coefficient corresponding to the duplication score of the DOI. If the DOIs in the two metadata match, 100 multiplied by the weight coefficient is used as the duplication score of the DOI. If the DOIs in the two metadata do not match, 0 multiplied by the weight coefficient is used as the duplication score of the DOI.
[0228] The formula for calculating the repetition score for this title is as follows:
[0229] Title_Score =F10 * min(100, title_similarity * 100)
[0230] Where F10 is the weight coefficient corresponding to the title repetition score, and title_similarity is the title similarity calculated using the selected string similarity algorithm (such as the cosine similarity algorithm).
[0231] The formula for calculating the duplication score for this author is as follows:
[0232] Author_Score=F11*(First_Author_Score + Corresponding_Author_Score +Other_Authors_Score)
[0233] In the formula, F11 is the weight coefficient corresponding to the author's duplication score, First_Author_Score is the matching score of the first author. Specifically, if the first author matches, First_Author_Score is 40 points, otherwise First_Author_Score is 0 points. Corresponding_Author_Score is the matching score of the corresponding author. Specifically, if the corresponding author matches, Corresponding_Author_Score is 40 points, otherwise Corresponding_Author_Score is 0 points. Other_Authors_Score is the matching score of other authors. The calculation formula of the matching score of other authors is as follows:
[0234] Other_Authors_Score = min(20, 20 * (number of other authors matched / total number of other authors))
[0235] It should be noted here that F9, F10 and F11 can be set according to specific application scenarios. For example, F9 can be set to 0.5, F10 can be set to 0.3, and F11 can be set to 0.2.
[0236] In a possible implementation, for duplicate document materials, it is also determined whether the duplicate document materials are from the same data source. If they are from the same data source, the operation of deleting duplicate documents described above is performed. If they are from different data sources, the integrity of the metadata of the two document materials is evaluated, and the metadata of the duplicate files is merged based on the integrity of the metadata of the two document materials, and the document materials corresponding to the metadata with lower integrity are deleted. When merging the metadata of duplicate files based on the integrity of the metadata of the two document materials, the following steps may be included:
[0237] First, the metadata corresponding to each data source is evaluated. Each field in the metadata is checked to see if it meets the integrity criteria set for the field, and the integrity score of the metadata is calculated based on the number of fields that meet the integrity criteria and the total number of fields in the metadata. The calculation formula for the integrity score is as follows:
[0238] Completeness score = (number of fields that meet the completeness criteria / total number of fields in the metadata) * 100
[0239] The integrity standards set for each field in the metadata are shown in Table 1:
[0240] Table 1
[0241]
[0242] Second, compare the integrity scores of the metadata corresponding to different data sources and select the data source with the highest integrity score as the primary data source. If the integrity scores are consistent, select the data source with the highest credibility as the primary data source.
[0243] Third, based on the metadata corresponding to the main data source, the metadata is merged. The specific steps are as follows: compare each field information in the metadata of the main data source with the corresponding field information in the metadata of other data sources. If the field information in the main data source meets the integrity standard, the field information in the main data source is retained. If the field information in the main data source does not meet the integrity standard, but the field information in other data sources meets the integrity standard, the field information in the other data sources is used to replace the corresponding field information in the main data source. If the field information in multiple data sources (including the main data source) meets the integrity standard, the field information in the data source with higher credibility is selected to replace the corresponding field information in the main data source, where the data source with higher credibility can be an official data source. If the field information in multiple data sources meets the integrity standard and the reliability of multiple data sources is close, the field information in the main data source is retained.
[0244] Furthermore, during the merging process, some special field information needs to be specially processed. Specifically, for fields such as keywords and references, the corresponding field information in multiple data sources can be merged and duplicates can be removed to obtain more complete field information. For abstracts, if the language versions provided in multiple data sources are different, the abstract information of each language version can be retained. For full-text links to references, the full-text links of all references provided in multiple data sources can be retained.
[0245] Furthermore, after completing the merging of metadata from multiple data sources, the merged metadata will be processed as follows: First, the merged metadata will be standardized. Specifically, the formats of dates, author names, etc. will be unified to ensure that all merged data follow consistent format standards. Second, the original data source will be marked for each field in the merged metadata, which will help with subsequent traceability and quality control. Next, the merged metadata will be quality checked to ensure that there are no obvious errors or inconsistencies in the merged metadata.
[0246] Furthermore, during the process of merging metadata, if there is a contradiction in the same field information in different data sources (such as different publication dates), the reliability ranking of the data sources is checked and the information is cross-validated. If the accurate field information cannot be determined, the field information in the primary data source is retained and the field is marked as uncertain.
[0247] It should be noted here that in the process of document deduplication, the two documents involved in the duplication score calculation may be different language versions of the same document, or they may be the preprint and the formally published version of the same document. For different language versions of the same document, the commonly used language version needs to be retained. For the preprint and the formally published version of the same document, the formally published version needs to be retained.
[0248] In a possible implementation, when two documents participating in the duplication score calculation use different language versions, the multi-language version scores of the two documents are directly calculated. When the calculated multi-language version score is greater than a preset multi-language version score threshold, the two documents are judged to be different language versions of the same document. The multi-language version score threshold can be set according to the specific application scenario. For example, the multi-language version score threshold can be set to 0.8. The calculation formula of the multi-language version score is as follows:
[0249] Multilingual version score = F12 * title similarity + F13 * author matching + F14 * abstract similarity + F15 * publication information matching
[0250] Wherein, F12, F13, F14 and F15 are weight coefficients corresponding to title similarity, author matching, abstract similarity and publication information matching, respectively. These weight coefficients can be set according to specific application scenarios. For example, F12, F13, F14 and F15 can be set to 0.3, 0.3, 0.2 and 0.2, respectively.
[0251] Among them, when calculating the title similarity in this formula, it is necessary to translate the non-English title into English first, and then calculate the cosine similarity between the English titles as the title similarity. When calculating the author matching in this formula, it is necessary to transliterate the author in non-Latin characters, and then determine the transliterated author to calculate the author matching. The algorithm formula for the author matching can be found in Author_Score, which will not be repeated here. When calculating the abstract similarity in this formula, it is necessary to translate the non-English abstract into English first, and then calculate the cosine similarity between the English abstracts as the abstract similarity. When calculating the publication information matching, the journal name matching, volume number matching and page matching of the two documents are calculated respectively, and the total score of each matching is calculated, and the calculated total score is used as the publication information matching. The calculation method of each matching degree can be found in Author_Score, which will not be repeated here.
[0252] In a possible implementation, when the repetition score between the current metadata and other metadata is greater than or equal to the preset repetition upper limit, the preprint matching score continues to be calculated. When the preprint matching score is greater than the preset preprint matching score threshold, the two documents involved in the comparison are determined to be a preprint version and a formally published version. The preprint matching score threshold can be set according to the specific application scenario. For example, the preprint matching score threshold can be set to 0.85. The calculation formula of the preprint matching score is as follows:
[0253] Preprint matching score = F16 * title similarity + F17 * author list overlap + F18 * content similarity + F19 * timeline plausibility
[0254] Wherein, F16, F17, F18 and F19 are weight coefficients corresponding to title similarity, author list overlap, content similarity and timeline rationality respectively. Each weight coefficient can be set according to the specific application scenario. For example, F16, F17, F18 and F19 can be set to 0.3, 0.2, 0.3 and 0.2 respectively.
[0255] In a possible implementation, the following steps may be included when calculating the rationality score of the time priority:
[0256] First, the data sources of the two documents involved in the comparison are used to determine which of the two documents is the preprint version and which is the formally published version. Specifically, some data sources provide preprint versions, and some data sources provide formally published versions. Therefore, the data sources can be used to determine whether the two compared documents are preprint versions or formally published versions.
[0257] Secondly, the timeline rationality score of the two comparative documents is calculated according to the following algorithm. Among them, the basic principle of timeline rationality assessment is that the release date of the preprint should be earlier than the release date of the official publication, and the time interval should be within a reasonable range. The timeline rationality score ranges from 0 to 1, with 1 representing the most reasonable timeline and 0 representing a completely unreasonable timeline.
[0258] When scoring the rationality of the timeline: a) If the preprint date is later than the official publication date, the timeline rationality score is 0, that is, the timeline is unreasonable; b) If the preprint date is earlier than the official publication date, calculate the time difference between the publication date of the preprint version and the official publication version, and determine the timeline rationality score based on the time difference. Specifically, after calculating the time difference, determine the range to which the time difference belongs, and the corresponding timeline rationality score can be found based on the range to which the time difference belongs. Among them, the system has a preset correspondence table between the time difference range and the corresponding timeline rationality score, and the correspondence table is shown in Table 2:
[0259] Table 2
[0260]
[0261] It should be noted here that if the date information of the preprint or the formally published version is missing, the timeline rationality score can be a default value, which can be set to 0.5.
[0262] Third, from the deduplicated literature, select a set number of literature with the highest comprehensive scores as the final literature related to the search information. The comprehensive score calculation formula of the literature is as follows:
[0263] Score = (B_norm * W_b) + (T_norm * W_t) + (C_norm * W_c) + (I_norm *W_i)
[0264] In the formula, Score is the comprehensive score of the document, B_norm is the BGE-Reranker-Large score, T_norm is the score of the publication time of the document, C_norm is the score of the number of citations of the document, I_norm is the journal impact factor score of the document, W_b, W_t, W_c and W_i are the weight coefficients of the corresponding items. Each weight coefficient can be set according to specific needs. For example, W_b, W_t, W_c and W_i can be set to 0.7, 0.15, 0.1 and 0.05 respectively.
[0265] Among them, the BGE-Reranker-Large score is to input the literature data and search information into the BGE-Reranker-Large model, which will calculate the similarity score between the two and normalize the similarity score output by the model to the range of 0-1. The normalized score is the BGE-Reranker-Large score of the document.
[0266] The calculation formula of the literature publication time score T_norm is as follows:
[0267] T_norm = (current year - year of publication) / (current year - year of earliest publication)
[0268] Among them, the current year and the year of search, the earliest document year is the earliest publication year of all literature materials.
[0269] The calculation formula of the citation score C_norm of the literature is as follows:
[0270] C_norm = log(1 + number of citations) / log(1 + maximum number of citations)
[0271] Among them, the number of citations is the number of citations of the currently evaluated literature, and the maximum citation test is the maximum number of citations in all literature.
[0272] The calculation formula for the journal impact factor score I_norm of the literature is as follows:
[0273] I_norm = log(1 + impact factor) / log(1 + maximum impact factor)
[0274] The impact factor is the impact factor of the journal corresponding to the currently evaluated literature, and the maximum impact factor is the impact factor of the largest journal corresponding to all literature.
[0275] Based on the above-mentioned comprehensive score calculation formula of the literature, the comprehensive score corresponding to each literature can be calculated, and then a set number of literatures with the highest comprehensive scores can be screened out as the final literatures related to the search information. Among them, the set number here can be set according to the specific application scenario. For example, the set number here can be set to 30, that is, the 30 articles with the highest comprehensive scores will be screened out from the deduplicated literature as the final literatures.
[0276] After obtaining the final literature data, the following steps are also included:
[0277] First, the final literature data are analyzed based on the pre-built second largest language model to obtain the literature summary of each final literature data. Specifically, before executing this step, it is necessary to try to use multiple open access databases (such as PubMedCentral, DOAJ) to download the full text of each literature data, and then input the full text of each literature data into the pre-built second largest language model in turn, so as to obtain the literature summary of each literature data outputted in turn by the second largest language model.
[0278] In order to enable the second largest language model to have the function of generating document summaries, it is necessary to first obtain a set number of document data samples, and for each document data sample, mark its corresponding document summary, so as to obtain a document summary training sample. After obtaining all document summary training samples, the document data samples in each document summary training sample are used as input data, and the marked document summary is used as standard output data to train the large language model. After the training is completed, the final document data is input into the second largest language model to obtain the document summary corresponding to the final document data.
[0279] Second, based on the search information, subject information, metadata and literature overview related to each final document, generate a literature review of each final document. The specific steps are as follows:
[0280] First, the following information is integrated: the search information entered by the user, the generated subject information related to the search information, and the document overview and metadata of the final document.
[0281] Secondly, a large language model (such as GPT-4) is used to generate a literature review framework. At the same time, the detailed information of each field in the literature review framework is generated by integrating the search information entered by the user, the subject information related to the search information, the literature overview and metadata of the final document, and other information, and then filled into the corresponding fields in the literature review framework. Among them, the literature review framework can include the following fields:
[0282] - Research Background
[0283] - Main research directions
[0284] - Methodology Overview
[0285] - Key Findings
[0286] - Research trends
[0287] - Interdisciplinary connections
[0288] - Future Outlook
[0289] Next, apply open source citation management and formatting algorithms to ensure that the citation format meets academic standards. The specific steps include: 1. Data preparation. Extract necessary citation information from the literature metadata and convert the citation information into a format supported by the selected library (such as BibTeX or JSON). 2. Style selection. Load or define the required citation style (such as APA, MLA). For libraries that support CSL, you can use a standard CSL file. 3. Citation generation. Generate in-text citations and reference entries using the selected library and style. 4. Format application. Insert the generated in-text citations into the body of the literature review, and collect and sort the reference entries. 5. Consistency check. Ensure the consistency of in-text citations and reference lists. Handle special cases such as multi-author citations, multiple documents by the same author, etc. 6. Output formatting. Integrate the formatted citations and references into the final document. Adjust the output format (such as HTML, LaTeX or plain text) as needed.
[0290] Finally, an interactive reference list is generated, including DOI links and database source information. The final output of the literature review includes the following:
[0291] - Main body of the review (including automatically inserted citations)
[0292] - Interactive reference list
[0293] - Visualize research topic network diagram
[0294] - Timeline of key documents
[0295] - Main author and institution statistics
[0296] The present disclosure provides a document retrieval method, including: obtaining retrieval information input by a user; semantically understanding the retrieval information based on a pre-built first language model to generate subject information related to the retrieval information; performing document retrieval based on the subject information to obtain document materials related to the retrieval information; wherein the subject information includes at least one of a subject, subject keywords, a brief description of the subject, and a subject-related discipline. In the method of the present disclosure, the user only needs to input the retrieval information, and the system can accurately output the subject information related to the retrieval information based on the first language model, and then retrieve high-quality document materials based on the accurate subject information, thereby improving the retrieval quality of the document retrieval.
[0297] <Device Example>
[0298] Figure 2 FIG. 1 is a schematic block diagram of a document retrieval device according to an embodiment of the present disclosure. Figure 2 As shown, the document retrieval device 100 includes:
[0299] The search information acquisition module 110 is used to acquire the search information input by the user;
[0300] A topic generation module 120, configured to perform semantic understanding on the search information based on a pre-built first language model, and generate topic information related to the search information;
[0301] A search module 130 is used to perform a document search based on the subject information to obtain document materials related to the search information;
[0302] The subject information includes at least one of the subject, subject keywords, a brief description of the subject, and the subjects involved in the subject.
[0303] <Equipment Embodiment>
[0304] Figure 3 FIG. 1 is a schematic block diagram of a document retrieval device according to an embodiment of the present disclosure. Figure 3 As shown, the document retrieval device 200 includes: a processor 210 and a memory 220 for storing executable instructions of the processor 210. The processor 210 is configured to implement any of the above-mentioned document retrieval methods when executing the executable instructions.
[0305] Here, it should be noted that the number of processors 210 may be one or more. Meanwhile, in the document retrieval device 200 of the embodiment of the present disclosure, an input device 230 and an output device 240 may also be included. The processor 210, the memory 220, the input device 230 and the output device 240 may be connected via a bus or in other ways, which are not specifically limited here.
[0306] The memory 220 is a computer-readable storage medium that can be used to store software programs, computer executable programs, and various modules, such as the program or module corresponding to the document retrieval method of the embodiment of the present disclosure. The processor 210 executes various functional applications and data processing of the document retrieval device 200 by running the software programs or modules stored in the memory 220.
[0307] The input device 230 may be used to receive input numbers or signals. The signals may be key signals related to user settings and function control of the device / terminal / server. The output device 240 may include display devices such as display screens.
[0308] <Storage Medium Embodiment>
[0309] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is also provided, on which computer program instructions are stored. When the computer program instructions are executed by the processor 210, any of the document retrieval methods described above is implemented.
[0310] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A document retrieval method, characterized in that: include: Get the search information entered by the user; Performing semantic understanding on the search information based on a pre-built first language model to generate topic information related to the search information; Performing a literature search based on the subject information to obtain literature materials related to the search information; The subject information includes the subject, subject keywords, a brief description of the subject, and at least one of the subjects involved in the subject; When a document search is performed based on the subject information to obtain document materials related to the search information, it includes: Based on the subject information, generating a general search formula corresponding to the subject information; Based on the general search formula, generating a target search formula for searching a target data source; Based on the target search formula, retrieving literature materials related to the search information from the target data source; When generating a general search formula corresponding to the subject information based on the subject information, it includes: Extracting topic keywords from the topic information; Obtaining standard subject keywords, synonyms and near-synonyms corresponding to the subject keywords, as well as limited fields and time ranges set by the user; Based on the standard subject keywords, synonyms and near-synonyms corresponding to the subject keywords and the limited fields and time ranges set by the user, a preset search formula generation algorithm is used to generate a general search formula corresponding to the subject information; The system includes a variety of search formula generation algorithms with different complexities. The higher the complexity, the more complex the generated general search formula; the lower the complexity, the simpler the generated general search formula. In order to ensure that the complexity of the generated general search formula is within a reasonable range, after the general search formula is generated, the following steps are also included: Calculate the complexity score of common search expressions; Determine whether the complexity score of the general search formula is within a preset complexity range, where the complexity range is expressed as (complexity lower limit, complexity upper limit); When it is determined that the complexity score of a general search formula is lower than the lower complexity limit, a search formula generation algorithm with a higher complexity level than the currently executed search formula generation algorithm is selected to regenerate the general search formula. When it is determined that the complexity score of a general search formula is higher than the upper complexity limit, a search formula generation algorithm with a lower complexity level than the currently executed search formula generation algorithm is selected to regenerate the general search formula.
2. The method according to claim 1, characterized in that When semantic understanding of the search information is performed based on the pre-built first language model to generate topic information related to the search information, it includes: Performing semantic analysis on the search information to obtain key semantic elements corresponding to the search information; Performing context analysis on the search information to obtain context information corresponding to the search information; Based on the key semantic elements and the context information, constructing a semantic relationship corresponding to the search information; Based on the semantic relationship, topic information related to the search information is generated.
3. The method according to claim 1, characterized in that When there are multiple subject information related to the search information, it also includes: Calculate the topic score corresponding to each topic information; Based on the topic scores corresponding to each topic information, target topic information is screened out from multiple topic information; A document search is performed based on the target subject information to obtain document materials related to the search information.
4. The method according to claim 1, characterized in that After obtaining the literature related to the search information, it also includes: Based on a preset document metadata template, construct metadata corresponding to each of the document materials; Deduplication of each of the document materials is performed based on the metadata corresponding to each of the document materials; From the deduplicated document materials, a set number of document materials with the highest comprehensive scores are screened out as final document materials related to the search information.
5. The method according to claim 4, characterized in that After obtaining the final literature materials, it also includes: Analyze each of the final document materials based on the pre-built second language model to obtain a document summary of each of the final document materials; Based on the search information, subject information, metadata and document overview related to each of the final document materials, a literature review of each of the final document materials is generated.
6. A document retrieval device, characterized in that: include: A search information acquisition module is used to acquire search information input by a user; A topic generation module, used for semantically understanding the search information based on a pre-built first language model, and generating topic information related to the search information; A search module, used to search for documents based on the subject information to obtain documents related to the search information; The subject information includes the subject, subject keywords, a brief description of the subject, and at least one of the subjects involved in the subject; When a document search is performed based on the subject information to obtain document materials related to the search information, it includes: Based on the subject information, generating a general search formula corresponding to the subject information; Based on the general search formula, generating a target search formula for searching a target data source; Based on the target search formula, retrieving literature materials related to the search information from the target data source; When generating a general search formula corresponding to the subject information based on the subject information, it includes: Extracting topic keywords from the topic information; Obtaining standard subject keywords, synonyms and near-synonyms corresponding to the subject keywords, as well as limited fields and time ranges set by the user; Based on the standard subject keywords, synonyms and near-synonyms corresponding to the subject keywords and the limited fields and time ranges set by the user, a preset search formula generation algorithm is used to generate a general search formula corresponding to the subject information; The system includes a variety of search formula generation algorithms with different complexities. The higher the complexity, the more complex the generated general search formula; the lower the complexity, the simpler the generated general search formula. In order to ensure that the complexity of the generated general search formula is within a reasonable range, after the general search formula is generated, the following steps are also included: Calculate the complexity score of common search expressions; Determine whether the complexity score of the general search formula is within a preset complexity range, where the complexity range is expressed as (complexity lower limit, complexity upper limit); When it is determined that the complexity score of a general search formula is lower than the lower complexity limit, a search formula generation algorithm with a higher complexity level than the currently executed search formula generation algorithm is selected to regenerate the general search formula. When it is determined that the complexity score of a general search formula is higher than the upper complexity limit, a search formula generation algorithm with a lower complexity level than the currently executed search formula generation algorithm is selected to regenerate the general search formula.
7. A document retrieval device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 5 when executing the executable instructions.
8. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Document retrieval method and device and storage medium
CN110516157A
Literature review generation method based on large language model
CN117709306A
Medical question and answer text generation method and equipment based on retrieval enhancement generation
CN117763114A