A document mining system and method based on scenario driving and correlation analysis
By using a scenario-driven and association analysis-based literature mining system, which utilizes authoritative journal dictionaries and large language models to dynamically screen high-value literature and perform structured analysis, the system solves the problems of low efficiency and unstable results in academic literature retrieval, and achieves efficient and interpretable research decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG AGRI UNIV
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-21
AI Technical Summary
Existing academic literature retrieval systems are inefficient when faced with massive amounts of literature, making it difficult to guarantee the authority and quality of the literature sources. Furthermore, the results of artificial intelligence analysis fluctuate greatly, making it difficult to reliably support high-level scientific research decision-making tasks.
A literature mining system based on scenario-driven and association analysis is adopted. Through a predefined authoritative journal dictionary and large language model, the screening rules are dynamically adjusted to generate a subset of high-value documents and conduct association structure analysis. Combined with the scientific research decision-making scenario, stable scientific research decision analysis results are generated.
It improves the efficiency and quality of literature retrieval, ensures the relevance and interpretability of analysis results, realizes automated support for scientific research decision-making, and significantly enhances the stability and consistency of scientific research analysis.
Smart Images

Figure CN122432330A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology and information technology, and specifically relates to a document mining system and method based on scenario-driven and association analysis. Background Technology
[0002] Currently, the number of academic documents worldwide is experiencing exponential growth. Researchers face the following challenges when confronted with this massive volume of literature: First, general academic databases (such as Web of Science and Scopus) aim for comprehensive coverage, resulting in a large number of low-relevance or low-quality documents mixed in with the search results. Researchers must spend a significant amount of time manually screening these documents, leading to inefficiency. Second, existing search ranking mechanisms largely rely on keyword matching or citation frequency, lacking fundamental guarantees of the academic authority of documents from the journal source, making it difficult to ensure the source quality of core references. Third, although some literature management or visualization tools (such as EndNote and VOSviewer) provide basic services, and emerging tools are beginning to integrate question-answering functions based on large language models, these methods typically analyze raw, unwashed search results directly. This results in high input noise, leading to significant fluctuations and poor reproducibility in AI analysis results, making it difficult to reliably support high-level research decision-making tasks such as research gap identification, methodological diagnosis, and future trend prediction. Summary of the Invention
[0003] To address the challenge of effectively reducing noise and ensuring stable input of literature while maintaining the authority of the source, and to achieve deep collaborative optimization of literature screening and intelligent analysis, this invention proposes a literature mining system and method based on scenario-driven and correlation analysis.
[0004] A literature mining system based on scenario-driven and association analysis, which achieves one of the objectives of this invention, includes a retrieval module and a decision analysis module; The retrieval module is used to obtain the corresponding initial document set from the filtered journal dictionary according to the user's retrieval instructions; The decision analysis module is used to identify research decision-making scenarios based on user search instructions, adjust the parameters of preset screening rules according to the research decision-making scenarios, and filter the initial document set to obtain a high-value document subset according to the preset screening rules after parameter adjustment; perform association structure analysis on the high-value document subset to generate document association information; and perform research decision analysis on the high-value document subset and the document association information based on a large language model to generate research decision analysis results that match the user's search needs.
[0005] The predefined journal dictionary uses journal names and their corresponding International Standard Serial Numbers (ISSNs) as core identifiers. For example, in social science fields such as management and economics, the widely recognized FT50 (Financial Times 50 Top Journals) list can be used as an example of this set, thereby enabling convenient tracking and efficient processing of the highest quality literature sources in the field.
[0006] Furthermore, it also includes a user interaction module for receiving user input, displaying search results and analysis reports, and providing interactive functions. It employs a graphical interface, which includes: a keyword input box, a journal multi-selector, a year range slider, and a search start button; search results are displayed in list format, supporting sorting by citation frequency, publication year, title, and other methods, and features pagination, abstract expansion / collapse, abstract translation (e.g., Chinese-English), and data export (e.g., CSV format); decision analysis results are presented in category tabs or card lists, with each analysis conclusion displayed independently.
[0007] Furthermore, the retrieval module receives retrieval instructions from the user interaction module, which include at least: keywords, a range of publication years, and a list of target journals selected from the journal dictionary of the data processing module. The retrieval module converts the journal list into ISSN-based filtering conditions, combines them with other user instructions to construct a structured query request, and obtains an initial set of documents from an external database. This initial set of documents includes: title, author, publication year, source journal, citation frequency, digital object identifier (DOI), abstract, author affiliation, and keywords.
[0008] Furthermore, the decision analysis module embeds or can invoke multiple large language model service interfaces (such as DeepSeek, GPT-4) to execute a series of configurable research decision analysis tasks based on this high-value literature subset. Specific task types include, but are not limited to: research topic gap mining, interdisciplinary theoretical transfer possibility analysis, research methodology audit and risk warning, prediction of future potential hot research directions, and suggestions for the management and practical transformation of research results. Each analysis task is designed to produce structured outputs.
[0009] Furthermore, methods for obtaining the corresponding initial literature set include: The journal names selected by the user are standardized and mapped to the corresponding ISSN list. Using the ISSN list as a filter, a structured query request is constructed based on the keywords and publication year range in the search command. An initial set of documents that meet the filter conditions and query request is retrieved from the external database.
[0010] Furthermore, the system has a standardized authoritative journal mapping table pre-stored locally. This mapping table uses the full name of the official standard journal as the index item and the unique valid ISSN number as the mapping value. The data comes from the official directory of authoritative journals and the publicly available data of the ISSN China National Center. In the mapping table, invalid journal entries that have been renamed, discontinued, or merged are pre-removed. The journal names selected by the user are standardized and normalized to remove journal abbreviations, capitalization differences, and special symbols, and converted into the standard full journal name format consistent with the above mapping table; Based on the standardized full name of the journal, precise key-value matching is performed in the authoritative journal mapping table to retrieve the corresponding ISSN number one by one. Perform batch verification and standardization on the matched ISSN numbers, and remove duplicate ISSNs, ISSNs with incorrect format, and invalid ISSNs that have not passed the official checksum verification. The verified ISSN numbers are aggregated to form a unique and valid list of ISSN numbers, which serves as the core filtering condition for the retrieval module to initiate queries to external databases.
[0011] Furthermore, adjusting the parameters of the preset filtering rules includes: In the research gap mining scenario, the weight coefficients of citation frequency, publication year, and journal score are assigned according to their importance in the research gap mining scenario, and the specified number of high-value literature subsets is set to 30%~50% of the total initial literature set; In the methodological evaluation scenario, the weight coefficients of citation frequency, publication year, and journal score are assigned according to their importance in the methodological evaluation scenario, and the specified number of high-value literature subsets is set to 10%~20% of the total initial literature set. In the scenario of predicting innovative topics, the citation frequency threshold is adjusted to the average citation frequency of all documents within N years prior to the retrieval time, and the range of publication years is expanded to the set number of years prior to the retrieval time. The weight coefficients of citation frequency, publication year, and journal score are assigned according to the importance of citation frequency, publication year, and journal score in the scenario of predicting innovative topics. The specified number of high-value document subsets is set to 40% to 60% of the total initial document set.
[0012] The research decision-making scenarios described in this invention include a custom research gap mining scenario, a methodology evaluation scenario, and an innovative topic prediction scenario. The research gap mining scenario refers to a research analysis scenario used to identify under-researched areas within the current research field. The methodology evaluation scenario refers to a research analysis scenario used to audit and identify risks in existing research methods and model frameworks within the field. The innovative topic prediction scenario refers to a research analysis scenario used to predict potential emerging research directions and hot topics in the future.
[0013] Furthermore, the preset screening rules of the decision analysis module include: calculating a comprehensive score for each document in the initial document set using a weighted scoring method, arranging them in descending order of comprehensive score, and selecting a specified number of documents at the top of the sorted list to obtain a high-value document subset; the calculation method of the weighted scoring method includes: document score = w1 × normalized citation frequency + w2 × normalized publication year + w3 × normalized journal score, where w1 is the weight coefficient of citation frequency, w2 is the weight coefficient of publication year, w3 is the weight coefficient of journal score, and the sum of the weight coefficients w1, w2, and w3 is 1. The default values are 0.6, 0.3, and 0.1, respectively.
[0014] Furthermore, it also includes a journal selection module, which is used to filter the journals in the journal dictionary to obtain a filtered journal dictionary. The filtering method includes: retaining only the journals that rank in the top set proportion in each subject area, and the citation half-life of the journal (i.e. the time required for the number of citations of a paper published in the journal to decrease by half from publication) ≥ a set time limit.
[0015] Furthermore, it also includes quality verification of high-value subsets, with verification methods including determining whether the following conditions are met: The average citation frequency of the literature is greater than or equal to the set number of times; the proportion of core authors is greater than or equal to the first set proportion; the proportion of authoritative journals is greater than or equal to the second set proportion; If any of the above conditions are not met, this document will be removed from the high-value document subset, and the decision analysis module will be called again to select high-value documents to supplement the high-value document subset until the set number of documents is met.
[0016] Furthermore, the decision analysis module is also used to perform quality adjustment on the generated scientific research decision analysis results. The quality adjustment method includes: evaluating the completeness, consistency, and confidence of the scientific research decision analysis results respectively; when the completeness is lower than a preset threshold, expanding the size of the high-value literature subset; when the consistency is lower than a preset threshold, increasing the weight coefficient of citation frequency and expanding the range of publication years; when the confidence is lower than a preset threshold, adjusting the parameters of the association structure analysis; the decision analysis module regenerates the high-value literature subset and the scientific research decision analysis results according to the adjusted relevant parameters, and performs quality adjustment again, iteratively executing until the consistency, completeness, and confidence all reach the corresponding preset thresholds.
[0017] The research decision analysis results refer to the structured information that the system ultimately outputs to users to directly support their specific research decision-making tasks. These results include at least the following three components: (1) Analysis conclusions: Specific conclusions or suggestions given directly based on users' search needs and the identified scientific research decision-making scenarios; for example, research found that there is a gap in field A, or suggested adopting method C and paying attention to risks D. (2) Supporting evidence: Key evidence supporting the above conclusions. This evidence should be linked back to one or more specific documents in the high-value document subset, or core nodes, clusters, etc. in the document association information; (3) Confidence index: A quantitative score output by the quality adjustment module that characterizes the credibility of the conclusion. The above three parts together constitute a complete scientific research decision analysis result, and are presented in a structured format such as JSON, XML, or category tabs.
[0018] The quality assessment indicators for the scientific research decision analysis results described in this invention include consistency, completeness, and confidence. Consistency refers to the semantic similarity between the analysis results obtained after repeatedly executing the same analysis task. Completeness refers to the coverage of the actual number of fields output by the analysis results with the number of fields required by the task. Confidence refers to the credibility score corresponding to the output analysis results of the large language model.
[0019] Furthermore, the process of performing association structure analysis on a subset of high-value documents and generating document association information includes: Extract the reference list and cited list of each document in the high-value document subset, and construct a directed citation network among the documents; Based on the directed citation network, the link analysis algorithm is used to calculate the citation influence of each document, and the core document nodes are determined according to the citation influence. Based on the directed citation network, several document clusters are obtained by dividing the network using a community detection algorithm, and the shortest path algorithm is used to extract the key citation paths between the document clusters. Aggregate and statistically analyze the core literature nodes, literature clusters, and citation paths between clusters obtained from the analysis to form a structured description text, which serves as literature association information.
[0020] The structured description text shall at least include: the total number of nodes, the total number of citation relationships, a list of core literature nodes and their citation influence scores, a list of literature clusters and their theme tags and representative literature, key citation paths between clusters and their weights. However, considering that claims should be as concise as possible, it can also be specified only in the specification.
[0021] Furthermore, the method for generating a scientific research decision-making analysis result matching the user's retrieval requirement includes: According to the identified scientific research decision-making scenario, match the corresponding scenario template from a pre-stored task template library; at least one analysis subtask required to complete the scenario analysis is predefined in the scenario template, and each analysis subtask includes: a prompt word template for guiding the analysis of the large language model, the output format required for the large language model to return, and the prerequisite task dependency of other subtasks that need to wait before this subtask is executed; Extract all analysis subtasks from the matched scenario template to form an analysis subtask set; According to the prerequisite task dependency in each analysis subtask, analyze the execution order of each subtask and determine the execution mode: serial execution is adopted for subtasks with prerequisite task dependencies, and parallel execution is adopted for subtasks without prerequisite task dependencies; Execute each analysis subtask sequentially or in parallel according to the parsed execution order and execution mode. When executing any analysis subtask, input the high-value literature subset, the literature association information, the prompt word template of the current analysis subtask, and the output format requirement into the large language model together, and call the large language model to generate an original analysis result that meets the output format requirement, which serves as the subtask original analysis result of this subtask; After all analysis subtasks are executed, splice the original analysis results of each subtask in sequence according to the execution order of the subtasks, and then integrate them in the preset JSON array, classification label page, or card list format to generate the final scientific research decision-making analysis result.
[0022] Furthermore, the method for identifying a scientific research decision-making scenario based on the user's retrieval requirement includes:
[0023] Furthermore, methods for generating research decision analysis results that match user search needs include: Extract keyword features from user search instructions, perform word segmentation on the search instruction text, remove stop words and punctuation marks, retain content words as keyword features, match the keyword features with a predefined database of multiple scenario keywords, and identify scenarios where at least one keyword is matched as the identified scientific research decision-making scenarios. Based on the identified scientific research decision-making scenario, the corresponding set of analysis subtasks is retrieved and decomposed from the pre-stored task template library. Each analysis subtask in the set of analysis subtasks includes prompt word templates, output format requirements, recommended large language models, and dependencies of preceding tasks. Based on the dependencies of the preceding tasks in each analysis subtask, the execution order of each subtask is resolved and the execution mode is determined. Subtasks with dependencies on preceding tasks are executed serially, while subtasks without dependencies on preceding tasks are executed in parallel. The abstract and title text of each document in the high-value document subset are extracted. The long text is cut into text blocks that do not exceed the length of the current large language model context window by using text segmentation and semantic compression algorithms. Meaningless punctuation and repeated sentences are removed and concatenated into a continuous string according to the document sorting order. According to the parsed execution order and execution mode, each analysis subtask is executed sequentially or in parallel. When executing any analysis subtask, the continuous string, the literature association information, the prompt word template of the current analysis subtask, and the output format requirements are sent to the large language model server. The recommended large language model is called to generate the original analysis result that meets the output format requirements as the original analysis result of the subtask. After all the analysis subtasks have been completed, the raw analysis results of each subtask are concatenated in the order of execution of the subtasks, and then integrated according to the preset JSON array, category tab, or card list format to generate the final scientific research decision analysis results.
[0024] The large language model described in this invention refers to a neural network model based on deep learning technology, pre-trained on a large-scale corpus, capable of understanding and generating natural language. It has the ability to accept text prompts as input and return text or structured data output that meets the instructions. Exemplarily, the large language model includes, but is not limited to, DeepSeek series models, GPT series models, GLM series models, and LLaMA series models, which can provide services through API calls or local deployment. In the system of this invention, the decision analysis module embeds or can call one or more service interfaces of the large language model, selecting the corresponding model to execute the analysis task based on the recommended large language model fields in the analysis sub-task.
[0025] A document mining method based on scenario-driven and association analysis to achieve the second objective of this invention includes: The initial set of documents is retrieved from the filtered journal dictionary based on the user's search instructions. The system identifies research decision-making scenarios based on user search instructions, determines parameters for preset screening rules based on these scenarios, and filters the initial document set to obtain a high-value document subset based on the adjusted preset screening rules. It then performs association structure analysis on the high-value document subset to generate document association information. Finally, it conducts research decision analysis on the high-value document subset and the document association information based on a large language model to generate research decision analysis results that match the user's search needs.
[0026] A computer program product for achieving the third objective of the present invention includes a computer program / instruction that, when executed by a processor, implements the steps of the document mining method based on scene-driven and association analysis.
[0027] The beneficial effects of this invention include: This invention obtains initial literature through a predefined set of authoritative journals, ensuring the basic quality of the literature retrieval sources. By dynamically adjusting the screening rules according to the research decision-making scenario, the selected high-value literature subset is deeply adapted to subsequent analysis tasks, improving the relevance and accuracy of the analysis. Through association structure analysis of the high-value literature subset, the structured relationships between the documents are introduced into a large language model, compensating for the semantic deficiencies of pure text input and improving the interpretability and completeness of the judgment results. By establishing a feedback loop between the quality assessment of the analysis results and the parameters of the screening rules, the system achieves self-optimization capabilities, significantly improving the consistency and stability of multi-round analysis. This invention can automate the entire process from literature retrieval, scenario-adaptive screening, association structure analysis to intelligent judgment without manual intervention, significantly improving the efficiency of scientific research analysis. The final output results can directly support research gap identification and other scientific research decision-making scenarios, effectively improving the quality of literature analysis. Attached Figure Description
[0028] Figure 1 This is a schematic diagram of the system described in this invention; Figure 2 This is a schematic diagram of document retrieval in an embodiment of the method described in this invention; Figure 3 This is a schematic diagram of the user interface in an embodiment of the method described in this invention. Detailed Implementation
[0029] The following detailed embodiments are provided to explain the technical solutions of the present invention, so that those skilled in the art can understand the present invention. The scope of protection of the present invention is not limited to the following specific embodiments. Any modifications or improvements made by those skilled in the art that incorporate the technical solutions of the present invention but differ from the following detailed embodiments are also within the scope of protection of the present invention.
[0030] This embodiment describes in detail a literature mining system based on scenario-driven and association analysis, which utilizes scenario-driven dynamic filtering, association structure analysis, and feedback closed-loop optimization. See details... Figure 1 The system includes a retrieval module, a decision analysis module, and a user interaction module, and the modules exchange data through a network or inter-process communication.
[0031] During system initialization, a predefined journal dictionary is loaded. This dictionary uses the journal names of authoritative journals as keys and their corresponding ISSN numbers as values. The dictionary data is derived from a preset journal directory. The system iterates through this authoritative journal dictionary, generates a checkbox for each journal name, and loads this checkbox into the journal selection area of the user interaction module. For example, a simplified FT50 journal dictionary is as follows: FT50_DB={ "Academy of Management Journal": "0001-4273", "Strategic Management Journal":"0143-2095", "Journal of Finance": "0022-1082", "Marketing Science":"θ732-2399", "MIS Quarterly": "0276-7783", #...Other Journals } In one embodiment, the system further includes a journal selection module for filtering journals in the journal dictionary to obtain a filtered journal dictionary. The filtered journal dictionary refers to a subset obtained by filtering the original journal dictionary according to preset rules. These preset rules include: retaining only the top 25% of authoritative journals in each subject area, journals with a citation half-life ≥ 3 years, and journals with no academic fraud records in the past 5 years. The authoritative journals are those in the JCR Q1 (Journal Citation Reports, Quartile 1) category. The JCR categorizes all academic journals in the same subject into four levels (Q1-Q4) based on their impact / quality, from highest to lowest. The Q1 category comprises the top 25% of journals in that subject.
[0032] In one embodiment, the data processing module further includes a journal update module for updating journals. The update rules include: periodically (e.g., monthly) synchronizing the official journal directory, synchronizing updates for scenarios such as journal renaming, discontinuation, and merger, deleting invalid ISSNs, and adding valid ISSNs.
[0033] In one embodiment, the system iterates through the authoritative journal dictionary, generates an interactive selection box for each journal name, and loads the journal selection box into the journal selection area of the user interaction module, such as... Figure 3 As shown; the built-in ISSN uniqueness verification algorithm ensures that a journal name corresponds to only one valid ISSN, avoiding mapping conflicts.
[0034] The search module is one of the core execution units of the system. When a user clicks the "Start Search" button in the graphical interface of the user interaction module, this module performs the following operations: Collect user input data, including: keywords (such as "Digital Transformation"), start_year and end_year selected by the slider (such as 2020-2024), and journal names selected by the user from the journal list.
[0035] The system calls a preset journal dictionary to batch map the journal names selected by the user to their corresponding ISSN numbers, generates an ISSN filter list issn_list, and automatically removes invalid or duplicate ISSNs.
[0036] A structured query request is constructed based on the OpenAlex API specification. The core filtering condition of this query request is: filter=primary_location.source.issn:{issn_list}|publication_year:{start_year}-{end_year}, which is combined with the search keywords.
[0037] Call the " / works" interface of OpenAlex in a paged manner and loop until all eligible literature records are obtained or the preset acquisition limit (such as 1000 articles) is reached.
[0038] For each piece of literature, parse and extract fields such as title, author, source journal, publication year, citation frequency, DOI (Digital Object Identifier), abstract, author institution, and keywords. Exclude invalid literature with missing core fields, empty abstracts, or invalid DOIs. Map the fields from different source databases to a standard format, organize them into structured DataFrame data, and transfer this DataFrame data to the decision analysis module.
[0039] The decision analysis module starts after receiving the literature DataFrame data transmitted by the retrieval module and performs the following operations: (I) Scenario recognition and determination of dynamic screening rules The decision analysis module extracts keyword features from the user's retrieval instructions transmitted by the user interaction module. The specific extraction method is as follows: Use a word segmentation tool to segment the retrieval instruction text, remove stop words (the preset stop word list includes: of, for, the, a, an, in, is, with, and, etc.) and punctuation marks, and retain Chinese characters with a length greater than 1 or English characters with a length greater than 2 as keyword features. For example, for the retrieval instruction "Analyze the research gap in the field of digital transformation", the extracted keyword features are analyze, digital transformation, and research gap. The system predefines several scientific research decision scenarios (such as research gap analysis scenario, methodology evaluation scenario, innovation topic prediction scenario, interdisciplinary research concept scenario, etc.), and each scenario corresponds to an exclusive core keyword library. The keyword library for each scenario consists of the core terms of that scenario. For example, the keyword library for the research gap analysis scenario includes {gap, blank, deficiency, lack, to be studied, unexplored}; the keyword library for the methodology evaluation scenario includes {method, methodology, model, algorithm, framework, empirical, simulation}; the keyword library for the innovation topic prediction scenario includes {future, trend, frontier, emerging, hot spot, prediction}. Match the extracted keyword features with the keyword libraries of each scenario, and determine the scenarios with at least one keyword hit as candidate scientific research decision scenarios. One embodiment also includes: Using a pre-trained Sentence-BERT semantic encoding model, encode the user's retrieval instruction text and the standard semantic descriptions of each candidate scenario and calculate the cosine similarity, and determine the scenarios with a similarity reaching the preset trigger threshold (such as 0.7) as the identified scientific research decision scenarios.
[0040] (II) Dynamically screen and generate a subset of high-value literature In one embodiment, based on the parameters determined above, the decision analysis module performs multi-dimensional sorting on the initial literature set. By default, the literature is sorted in descending order according to the primary "Citations" field (in descending order), the secondary field is the publication year (in descending order), and the third field is the journal score (in descending order). The weight and priority of each field are dynamically adjusted according to the scenario.
[0041] In another embodiment, a weighted scoring method is used to calculate a comprehensive score for each document in the initial document set. The documents are then sorted in descending order of comprehensive score, and a specified number of documents at the top of the sorting are selected to obtain a high-value document subset. The calculation method of the weighted scoring method includes: document score = w1 × normalized citation frequency + w2 × normalized publication year + w3 × normalized journal score, where w1, w2, and w3 are weight coefficients dynamically adjusted according to the scientific research decision-making scenario, with default values of 0.6, 0.3, and 0.1, respectively. The documents are then sorted in descending order of comprehensive score.
[0042] In one embodiment, the method further includes scoring the documents using a weighted scoring method and arranging them in descending order of score; the scoring formula includes: Literature score = w1 × normalized citation frequency + w2 × normalized publication year + w3 × normalized journal score; The weighting coefficients w1, w2, and w3 are 0.6, 0.3, and 0.1, respectively, and are dynamically configured according to the scenario.
[0043] The normalization method is minimum-maximum normalization, that is, normalized value = (original value of indicator - minimum value of indicator) / (maximum value of indicator - minimum value of indicator).
[0044] After sorting, the top N documents are selected according to the system's preset quantity threshold N (e.g., N=100) to form a high-value document subset. The number of documents in the subset can be dynamically adjusted according to the total number of searches. When the search volume is < N, all documents are selected.
[0045] (III) Relationship Structure Analysis This step performs association structure analysis on a subset of high-value documents to generate document association information; specifically, it includes: The reference and cited list of each document in the high-value document subset is extracted to construct a directed citation network among the documents. This network provides the topological data foundation for subsequent analysis and computation. Based on this network, a link analysis algorithm (such as PageRank) is used to calculate the citation influence of each document, identifying core document nodes that meet certain criteria. In this embodiment, the criteria are the top 10% of the most influential documents. Simultaneously, based on the network, a community detection algorithm is used to divide the document into several clusters, and a shortest path algorithm is employed to extract the key citation paths between these clusters, thereby characterizing the evolution direction of the research trajectory in the field.
[0046] The core document nodes, document clusters, and citation paths between clusters obtained from the above analysis are aggregated and statistically analyzed to form a structured descriptive text. This structured descriptive text, as document association information, is input into a large language model along with a subset of high-value documents for research decision analysis.
[0047] In one embodiment, the method further includes quality verification of the high-value subset, wherein the verification method includes determining whether the following conditions are met: The average citation frequency of the literature is ≥50 times; the proportion of core authors is ≥30%; and the proportion of authoritative journals is ≥60%. In this embodiment, authoritative journals, i.e. JCR Q1 journals, will be removed from the high-value literature subset if any of the above conditions are not met. The decision analysis module will then be called again to select high-value literature to supplement the high-value literature subset until the set number of literatures is met.
[0048] (iv) Scientific research decision analysis based on large language models The decision analysis module predefines several scientific research decision analysis tasks, each corresponding to a prompt word template and structured output format requirements, such as... Figure 2 As shown, the process begins by extracting the title and abstract of each document from the high-value document subset. These are then concatenated into a continuous string in sorted order. If the length of the concatenated text exceeds the maximum context window length of the current large language model, a text segmentation and semantic compression algorithm is used to cut the long text into text blocks no longer than the window length. Each text block is then used as the analysis context. Simultaneously, the citation network description generated by the association structure analysis is used as the common context. Then, based on the research decision-making scenario identified in the first step, the corresponding set of analysis subtasks is matched from the task template library.
[0049] Before sending the prompts to the large language model, the system automatically combines the concatenated continuous strings with the reference network description generated by association structure analysis, appending this as a common context before the prompts for all analysis subtasks. The combination format is as follows: Assuming the network analysis results are generated according to the template in step 4 of Example 3, the initial part of the combined prompt word finally sent to the large language model will be: [Core Content of the Documents] The following is a concatenated text of the title and abstract of each document in the current high-value document subset: Title: Title A. Abstract: Contents of Abstract A... Title: Title B. Abstract: Contents of Abstract B... [Document Association Structure Information] Total number of nodes: 156, total number of references: 423.
[0050] Core Literature Nodes (Top 10 in Citation Influence): Title: "Title A", Author: "Author A", Citation Impact Score: 0.087 Title: "Title C", Author: "Author C", Citation Impact Score: 0.065 Document clustering (3 clusters in total): Cluster 1: Tags [Deep Learning, Neural Networks, Representation Learning], contains 45 articles, representative articles: [Title A, Title D].
[0051] Cluster 2: Tags [model compression, edge computing, lightweight], contains 38 documents, representative documents: [title E, title F].
[0052] Cluster 3: Tags: Fairness, Explainability, Ethical Alignment, containing 32 articles, with representative articles: Title G, Title H.
[0053] Key inter-cluster reference paths (research context): Path 1: Cluster 1 → Cluster 2 (Citations: 23) Path 2: Cluster 1 → Cluster 3 (Citations: 15) Path 3: Cluster 2 → Cluster 3 (Citations: 8) Please make full use of the above core content and structural information when conducting the following analysis.
[0054] In another embodiment, the prompt word templates for each analysis subtask include: (1) For the research gap identification task, the template is as follows: Based on the core content and citation network structure information of the above literature, and in combination with the literature content, identify three core research gaps in the current research field. Return the results in JSON array format, with each entry containing the gap name, cause, and suggested research design fields.
[0055] (2) For the methodology audit task, the template is as follows: Based on the core content and citation network structure information of the above literature, audit the research methods commonly used in the literature and identify the concentrated areas of methodological risk. Return the results in JSON array format, with each entry containing the method name, the subject cluster it belongs to, the potential risk, and the improvement suggestion fields.
[0056] (3) For the task of predicting innovative research topics, the template is as follows: Based on the core content and citation network structure information of the above-mentioned literature, predict the research directions that may emerge in the next 1-2 years. Return the results in JSON array format, with each entry containing the direction name, innovation point, recommended journal, and prediction basis fields.
[0057] (4) For interdisciplinary theory transfer analysis tasks, the template is as follows: Based on the core content and citation network structure information of the above literature, identify which other disciplines the core theoretical framework of the current field can be transferred to. Return the results in JSON array format, with each entry containing the target discipline, transferable theory, transfer path, and potential research question fields.
[0058] (5) For the task of analyzing practical or policy implications, the template is as follows: Based on the core content and citation network structure information of the above-mentioned literature, analyze the implications of the research results for practice or policy making. Return the results in JSON array format, with each entry containing the implications, applicable scenarios, implementation suggestions, and supporting literature source fields.
[0059] The system calls a recommended large language model API (such as the DeepSeek API) to send the combined suggestion words to the model and then to the large language model server. It receives the JSON-formatted analysis results returned by the large language model, performs data validity verification and field parsing, and generates the original analysis results for the subtask. If multiple analysis subtasks exist, the system executes them in parallel or sequentially based on task dependencies. The parsed structured analysis results are cached in a local cache and then transmitted to the user interaction module.
[0060] (v) Quality Adjustment The decision analysis module is also used to adjust the quality of the generated scientific research decision analysis results, and the evaluation indicators include: Consistency: The same analysis was performed three times, and the semantic similarity between the results was calculated (using Sentence-BERT). The average value was taken as the consistency score. Completeness: Check whether the results cover all output fields required by the subtask, and calculate the ratio of the actual number of output fields to the required number of fields; Confidence score: Extracted from the logits returned by the large language model or the confidence score configured internally by the system.
[0061] When the consistency score is lower than the preset threshold of 0.8, the completeness score is lower than the preset threshold of 0.9, or the confidence score is lower than the preset threshold of 0.7, the quality assessment is deemed to have failed.
[0062] The specific calculation methods for the above evaluation indicators are as follows: (a) Completeness assessment The system reads the preset "output format requirements" field from the corresponding analysis subtask template, parses it to obtain all the field names that the task requires to be output, and forms the field set F. required Then, the analysis results actually returned by the large language model are parsed, and the non-empty fields actually contained therein are extracted to form the actual output field set F. actual The formula for calculating completeness C is: C = |F| actual | / |F required | Where |·| represents the number of elements in the set. When C is lower than a preset threshold (e.g., 0.9), it is determined to be insufficient in completeness.
[0063] (ii) Consistency Assessment The system repeatedly calls the large language model three times for the same analysis task and the same input data, obtaining three independent analysis result texts R1, R2, and R3. Subsequently, a semantic similarity algorithm is used to calculate the cosine similarity between R1 and R2, R2 and R3, and R1 and R3, respectively, yielding three similarity scores S. 12 S 23 S 13 Consistency score I is calculated by taking these three scores S. 12 S 23 S 13 Arithmetic mean: When I falls below a preset threshold (e.g., 0.85), it is considered insufficiently consistent. This mechanism ensures the stability and reproducibility of the analysis results.
[0064] (III) Confidence Assessment When calling the large language model API, the system proactively enables the return of logprobs. For the output text generated by the model, the system extracts the log probability of each generated token, calculates the arithmetic mean of the log probabilities of all tokens, and then converts this mean into a probability value between 0 and 1 using an exponential function. This value is used as the confidence score K for this analysis result. When K is lower than a preset threshold (e.g., 0.7), the confidence level is considered insufficient. For models that do not support returning logprobs, the system can alternatively add a command to the prompt template: "Please add a field named confidence_score, with a value between 0 and 1, representing your level of certainty about the results of this analysis, to the end of your returned JSON result." The system parses this field from the returned result as the confidence score.
[0065] When the evaluation result is lower than the preset threshold, at least one of the following adjustments will be automatically executed and the high-value literature subset and scientific research decision analysis results will be regenerated until the quality evaluation meets the standard; (1) Adjust the size parameter of the high-value literature subset, such as increasing the size of the high-value literature subset by 50% (e.g., from 100 articles to 150 articles); (2) Adjust the weight coefficient of citation frequency, such as increasing the weight of citation frequency by 10 percentage points; (3) Adjust the upper and lower limits of the publication year range, such as expanding the publication year range by 1-2 years; (4) Adjust the clustering threshold of the association structure analysis. The magnitude of the above adjustments will be dynamically determined according to the degree of non-compliance: when the completeness is lower than the threshold, the subset size will be increased first; when the consistency is lower than the threshold, the weight of citation frequency will be increased first and the publication year range will be expanded first; when the confidence is lower than the threshold, the parameters of the association structure analysis will be adjusted first to enhance the reliability of the input information.
[0066] After adjustment, the system automatically repeats steps (ii) to (iv) to perform a quality assessment again until the assessment result reaches the preset threshold or the preset maximum number of retries (e.g., 3 times). Finally, the research decision analysis results that meet the quality assessment standards are transmitted to the user interaction module for display.
[0067] The user interaction module is implemented using a web framework (such as Streamlit). The interface is divided into a left-side control panel and a central main display area. The initial literature collection is displayed as an expandable list, with each result showing information such as title, year, citation count, journal, and author. Users can click "Expand" to view the complete abstract, and click the "Translate Abstract" button to retrieve the Chinese translation from a cloud-based translation service. The interface also provides an "Export CSV" button, which can export all metadata from the initial literature collection to a CSV file. The structured analysis results output by the decision analysis module are displayed in tabs, with each tab corresponding to an analysis task (such as "Gap Mining" or "Methodology Audit"). The interface also includes a "Citation Network Visualization" area, which displays the citation relationships between documents in the form of a force-directed graph. After the user clicks the "Start Analysis" button within a tab, the analysis results for that task will be presented in a structured card list format.
[0068] Example 2 This embodiment describes a document mining method based on the system described in Embodiment 1. The method includes the following steps: The initial set of documents is retrieved from the filtered journal dictionary based on the user's search instructions. The system identifies research decision-making scenarios based on user search instructions, adjusts the parameters of preset filtering rules according to these scenarios, and filters the initial document set to obtain a high-value document subset based on the adjusted preset filtering rules. It then performs association structure analysis on the high-value document subset to generate document association information. Finally, it conducts research decision analysis on the high-value document subset and the document association information based on a large language model to generate research decision analysis results that match the user's search needs.
[0069] In one embodiment, the user's search command input method includes: the user accesses a graphical interactive interface, enters the keyword "Artificial Intelligence Ethics" in the sidebar, sets the year range to 2019-2024 using the slider, selects 10 journals such as "Journal of Business Ethics" and "MIS Quarterly" in the journal selection box, and clicks "Start Search".
[0070] In one embodiment, the system converts the selected journals into an ISSN list, combines it with keywords and year ranges, and sends a request to the OpenAlex API. After multiple rounds of pagination requests, a total of 235 articles matching the criteria are retrieved.
[0071] In one embodiment, the method for scene identification and dynamic filtering includes: the system analyzes user search instructions to identify research decision-making scenarios as research gap exploration scenarios and methodological evaluation scenarios (multiple scenarios overlapping). The subset size is set to 120 articles. 235 articles are sorted using a weighted scoring method, and the top 120 articles are selected to form a high-value literature subset.
[0072] Association Structure Analysis: The system obtains the reference list of each document in the high-value document subset via API, constructs a citation network with documents as nodes and citation relationships as directed edges. The PageRank algorithm is used to calculate the citation influence of each document, and core document nodes are determined based on this influence. Based on the directed citation network, several document clusters are obtained using the Louvain community detection algorithm, and the Dijkstra shortest path algorithm is used to extract key citation paths between document clusters. The core document nodes, key citation paths, and document clusters are transformed into structured descriptive text according to a preset template, for example: the current document set contains X documents, forming Y citation relationships. The top 10% of core documents in terms of citation influence are: [Document ID and corresponding citation influence value]. The system is divided into K document clusters using the community detection algorithm, with the size and topic tags of each cluster as follows: [Cluster 1: N1 documents, topic tag 'xx'; Cluster 2: N2 documents…]. The main citation relationships between clusters are: [Cluster A → Cluster B: C citation edges].
[0073] Multi-stage AI assessment: Users click on different tabs under the "AI Decision Brain" area. For example, clicking the start button under the "Gap Mining" tab. The system concatenates the titles and abstracts of the first 100 documents, along with the association structure analysis results (citation network description), into background text, and sends it to the DeepSeek-V3 API along with preset gap mining prompts. After a period of time, the system receives and parses the returned JSON data, generates and displays the returned research gap cards in the interface, and each card contains a specific gap description, cause analysis, and research design suggestions. For example, if the returned JSON results contain 4 research gaps, each gap is marked with its association with core literature and topic clusters (e.g., gap 1 originates from the non-overlapping domain between the fairness algorithm cluster and the privacy protection cluster). The system performs a quality assessment, achieving a consistency score of 0.85, a completeness score of 0.95, and a confidence score of 0.82, all meeting the standards, and directly outputs a report.
[0074] Results Interaction: Users can browse the search results list, expand abstracts of articles of interest, and translate them. Simultaneously, users can switch between different AI analysis tabs to obtain suggestions from various dimensions, such as methodological audits and future research topics. They can switch to the citation network visualization tab to view the force-directed graph. Finally, users can download a CSV file containing metadata and correlation analysis results for all 235 documents for later use.
[0075] Example 3 The initial set of documents is retrieved from the filtered journal dictionary based on the user's search instructions. The system identifies research decision-making scenarios based on user search instructions, adjusts the parameters of preset filtering rules according to these scenarios, and filters the initial document set to obtain a high-value document subset based on the adjusted preset filtering rules. It then performs association structure analysis on the high-value document subset to generate document association information. Finally, it conducts research decision analysis on the high-value document subset and the document association information based on a large language model to generate research decision analysis results that match the user's search needs.
[0076] In one embodiment, the method for obtaining the corresponding initial document set includes: The journal names selected by the user are standardized and mapped to the corresponding ISSN list. Using the ISSN list as a filter, a structured query request is constructed based on the keywords and publication year range in the search command. An initial set of documents that meet the filter conditions and query request is retrieved from the external database.
[0077] In one embodiment, the preset screening rules include: calculating a comprehensive score for each document in the initial document set using a weighted scoring method, arranging them in descending order of comprehensive score, and selecting a specified number of documents at the top of the sorting to obtain a high-value document subset; the calculation method of the weighted scoring method includes: document score = w1 × normalized citation frequency + w2 × normalized publication year + w3 × normalized journal score, where w1, w2, and w3 are weight coefficients, with default values of 0.6, 0.3, and 0.1, respectively.
[0078] In one embodiment, the parameters for adjusting the preset filtering rules include: In the research gap mining scenario, the weighting coefficients for citation frequency, publication year, and journal score are assigned based on the importance of citation frequency, publication year, and journal score. In this embodiment, the weighting coefficients are adjusted to w1=0.3, w2=0.5, and w3=0.2, and the specified number of high-value literature subsets is set to 30%~50% of the total initial literature set. In the methodological evaluation scenario, the weighting coefficients for citation frequency, publication year, and journal score are assigned based on the importance of citation frequency, publication year, and journal score. In this embodiment, the weighting coefficients are adjusted to w1=0.4, w2=0.2, and w3=0.4, and the specified number of high-value literature subsets is set to 10%~20% of the total initial literature set. In the scenario of predicting innovative topics, the citation frequency threshold is adjusted to the average citation frequency of all documents within N years prior to the retrieval time, and the publication year range is expanded to N years prior to the retrieval time. The weight coefficient of citation frequency is assigned according to the importance of citation frequency, publication year, and journal score in the scenario of predicting innovative topics. In this embodiment, the weight coefficients are adjusted to w1=0.4, w2=0.4, and w3=0.2, and the specified number of high-value document subsets is set to 40%~60% of the total initial document set.
[0079] It is important to emphasize that the aforementioned configurations of weighting coefficients w1, w2, and w3, and the size of high-value literature subsets for different research decision-making scenarios, are not simple empirical values or conventional designs. Rather, they are configuration schemes with specific technical significance established by this invention based on a quantitative analysis of the contribution of literature evaluation indicators in each scenario. Specifically: For research gap discovery, the core objective is to identify high-value unexplored areas within a current research field. A truly valuable research gap is characterized by its cutting-edge nature and timeliness; that is, the research direction corresponding to the gap should be actively discussed recently but has not yet formed a mature system. Therefore, publication year (w2) is assigned the highest weight of 0.5 to prioritize capturing the latest literature dynamics from the past 2-3 years, ensuring the discovered gap has timely value. Citation frequency (w1) has the next highest weight, used to filter literature that has gained some attention in a short period, demonstrating the potential influence of the direction, but avoiding over-reliance on long-term accumulated high-citation literature and missing emerging directions. Journal score (w3) has the lowest weight of 0.2, because breakthrough gaps often appear in highly innovative non-top-tier journals. The subset size is set at 30%~50% to ensure sufficient samples for pattern identification while eliminating marginal noise.
[0080] In the context of methodological evaluation, the core objective is to audit the maturity, reliability, and potential risks of existing research methods, models, or frameworks. Methodologies that are frequently cited typically represent extensive validation within the academic community, thus receiving the highest weight (w1). Methodological literature published in top-tier journals often undergoes more rigorous peer review and possesses higher credibility, hence the high weight (w3). Furthermore, mature methodologies have a long evolutionary cycle, making the most recent publication year relatively less important. Therefore, this invention places citation frequency and journal score on an equal footing in this scenario (w1=w3), while reducing the weight of publication year (w2). The subset size is set at 10%–20%, aiming to focus on a small number of the most classic and authoritative documents for in-depth diagnostics, avoiding diluting the analytical focus with a large number of non-core documents.
[0081] For the scenario of predicting innovative research topics, the core objective is to predict potential research directions that may become hot topics within the next 1-3 years. A reliable innovative research topic typically requires a certain academic foundation (characterized by citation frequency, reflecting its influence in the field) and must also possess significant cutting-edge nature (characterized by publication year, reflecting its novelty). Therefore, in this scenario, citation frequency and publication year are equally important (w1=w2), aiming to screen high-potential literature that has both a certain level of recognition and is recently published. The journal score weight w3 is relatively low. Furthermore, to capture a broader range of early signals, this scenario dynamically adjusts the citation frequency threshold to the average of the past three years and expands the range of publication years, while also increasing the subset size to 40%~60% to maximize the coverage of diverse literature that may represent future trends.
[0082] Through the refined and asymmetric weight configurations for different scenarios described above, the system of this invention can more accurately filter out high-value literature subsets that are highly compatible with specific scientific research tasks from the source, thereby providing optimal input for subsequent intelligent judgment based on large language models, which cannot be achieved by conventional and unified sorting and filtering rules.
[0083] In one embodiment, the method further includes: quality adjustment of the generated scientific research decision analysis results. The quality adjustment method includes: the decision analysis module is further used to adjust the quality of the generated scientific research decision analysis results, and the quality adjustment method includes: evaluating the completeness, consistency, and confidence of the scientific research decision analysis results respectively; when the completeness is lower than a preset threshold, expanding the size of the high-value literature subset; when the consistency is lower than a preset threshold, increasing the weight coefficient of citation frequency and expanding the range of publication years; when the confidence is lower than a preset threshold, adjusting the parameters of the association structure analysis; the decision analysis module regenerates the high-value literature subset and the scientific research decision analysis results according to the adjusted relevant parameters, and performs quality adjustment again, iteratively executing until the consistency, completeness, and confidence all reach the corresponding preset thresholds.
[0084] In one embodiment, the method further includes filtering the journals in the journal dictionary to obtain a filtered journal dictionary. The filtering method includes retaining only the journals that rank in the top set proportion in each subject area, and the citation half-life of the journals is greater than or equal to a set time limit.
[0085] In one embodiment, the process of performing association structure analysis on a subset of high-value documents and generating document association information includes: Extract the reference list and cited list of each document in the high-value document subset, and construct a directed citation network among the documents; Based on the directed citation network, the link analysis algorithm is used to calculate the citation influence of each document, and the core document nodes are determined according to the citation influence. Based on the directed citation network, several document clusters are obtained by dividing the network using a community detection algorithm, and the shortest path algorithm is used to extract the key citation paths between the document clusters. The core document nodes, document clusters, and citation paths between clusters obtained from the analysis are aggregated and statistically analyzed to form a structured descriptive text, which serves as document association information.
[0086] This embodiment takes a specific document mining task as an example and performs the following steps on a subset of high-value documents (assuming it contains 120 documents): Step 1: Construct a directed reference network The system retrieves the reference list and cited list (i.e., which documents cited the document) for each document in the high-value document subset via the OpenAlex or CrossRef API. All documents appearing in the references or cited documents are treated as network nodes, denoted as V, with a total of N = |V| nodes. For each document... i References j Establish a reference relationship from node j Pointing to node i The directed edge e j→i This forms a set of directed edges E. The final directed reference network graph is G=(V,E).
[0087] Step 2: Calculate citation impact and identify core literature nodes This embodiment uses a Personalized PageRank algorithm with restart to calculate the citation influence of each document, in order to more accurately reflect its relative importance within a limited set of documents. The specific steps are as follows: Set the damping factor α = 0.85 (meaning there is an 85% probability of continuing along the outgoing chain and a 15% probability of jumping back to the initial seed node set).
[0088] Define a personalized vector v: designate each document in the high-value document subset as a seed node and assign it an equal initial weight, i.e., for the seed node set S (size m=120). v i =1 / m, if i∈S, otherwise v i =0.
[0089] Iteratively calculate the PageRank value r until convergence: r (t+1) =α×r (t) ×M+(1-α)×v Where M is the transition probability matrix (if node i Excessive d i Then Mi→j =1 / d i For dangling nodes that have not left the chain, jump evenly to all nodes.
[0090] After convergence, the top k documents with the highest values in r are selected as core document nodes. In this embodiment, k = max(10, ⌈0.1×N⌉), which means at least 10 documents or 10% of the total number of nodes. The titles, authors, and citation impact scores of these core document nodes are recorded as structured data.
[0091] Step 3: Divide the literature into clusters and extract key citation paths This embodiment uses the Louvain community discovery algorithm to segment literature clusters in order to identify research subfields: Initially, each node is treated as an independent community. The first stage is modularity optimization, which includes: traversing each node and calculating the modularity increment ΔQ after moving it to any neighboring community; the modularity calculation formula is: Where L is the total number of sides, A ij For elements of the adjacency matrix (directed edges are considered bidirectional), d i For nodes i The degree (in-degree + out-degree), δ(c) i ,c j Indicator node i and j Determine if the node is in the same community. Move the node to the community that maximizes ΔQ; otherwise, leave it in its current position. Repeat this iteration until the modularity no longer increases.
[0092] The second phase is network cohesion, which includes: compressing the communities obtained in the first phase into supernodes, constructing a new weighted network, and returning to the first phase for further optimization.
[0093] After convergence, several document clusters are obtained, each containing a group of documents closely related in terms of citation. The system automatically generates topic tags for each cluster: extracting high-frequency words from the titles and abstracts of all documents within the cluster (after filtering for stop words), and taking the three keywords with the highest word frequency as cluster tags.
[0094] Each cluster is treated as a supernode, and the citation relationships between clusters are calculated: if there is a citation edge from any document in cluster A to any document in cluster B, a directed edge from A to B is added to the hypergraph, with a weight equal to the total number of citations between clusters. Then, Dijkstra's shortest path algorithm is used to calculate the shortest path between each pair of supernodes (using hop count or the reciprocal of the weight as the distance), and the top 5 paths with the highest betweenness centrality connecting different clusters are extracted as key citation paths. Each path is formatted as: cluster label 1 → cluster label 2 → ... → cluster label n, representing the evolution or migration of research ideas.
[0095] Step 4: Generate structured description text The system aggregates the above analysis results into structured descriptive text according to the following template for use by large language models: [Document Association Structure Information] Total number of nodes: {N}, Total number of references: {E}.
[0096] Core literature nodes (Top {k} of citation influence): 1. Title: [Title 1], Author: [Author 1], Citation Impact Score: [Score 1] 2. Title: [Title 2], Author: [Author 2], Citation Impact Score: [Score 2] ... Document clustering (a total of {C} clusters): Cluster 1: Tags [Tag A, Tag B, Tag C], containing {size1} documents, representing documents: [List 2-3 core document titles].
[0097] Cluster 2: Tags [Tag D, Tag E, Tag F], containing {size2} documents, representing documents: ... ... Key inter-cluster reference paths (research context): Path 1: {Cluster A} → {Cluster B} → {Cluster C} (Citation count: {weight}) Path 2: ... This structured text, along with the title and abstract of a subset of high-value documents, serves as the input context for a large language model, used for subsequent research decision analysis.
[0098] In one embodiment, the method for generating research decision analysis results that match the user's search needs includes: Extract keyword features from user search instructions, perform word segmentation on the search instruction text, remove stop words and punctuation marks, retain content words as keyword features, match the keyword features with a predefined database of multiple scenario keywords, and identify scenarios where at least one keyword is matched as the identified scientific research decision-making scenarios. Based on the identified scientific research decision-making scenario, the corresponding set of analysis subtasks is retrieved and decomposed from the pre-stored task template library. Each analysis subtask in the set of analysis subtasks includes prompt word templates, output format requirements, recommended large language models, and dependencies of preceding tasks. Based on the dependencies of the preceding tasks in each analysis subtask, the execution order of each subtask is resolved and the execution mode is determined. Subtasks with dependencies on preceding tasks are executed serially, while subtasks without dependencies on preceding tasks are executed in parallel. The abstract and title text of each document in the high-value document subset are extracted. The long text is cut into text blocks that do not exceed the length of the current large language model context window by using text segmentation and semantic compression algorithms. Meaningless punctuation and repeated sentences are removed and concatenated into a continuous string according to the document sorting order. According to the parsed execution order and execution mode, each analysis subtask is executed sequentially or in parallel. When executing any analysis subtask, the continuous string, the literature association information, the prompt word template of the current analysis subtask, and the output format requirements are sent to the large language model server. The recommended large language model is called to generate the original analysis result that meets the output format requirements as the original analysis result of the subtask. After all the analysis subtasks have been completed, the raw analysis results of each subtask are concatenated in the order of execution of the subtasks, and then integrated according to the preset JSON array, category tab, or card list format to generate the final scientific research decision analysis results.
[0099] Example 4 A computer program product includes a computer program / instructions that, when executed by a processor, implement the various steps of the method described in this invention.
[0100] The contents not described in detail in this specification are existing technologies known to those skilled in the art.
Claims
1. A document mining system based on scenario-driven and association analysis, characterized in that, It includes a retrieval module and a decision analysis module; The retrieval module is used to obtain the corresponding initial document set from the filtered journal dictionary according to the user's retrieval instructions; The decision analysis module is used to identify scientific research decision-making scenarios based on user search instructions, adjust the parameters of preset screening rules based on the scientific research decision-making scenarios, and filter the initial literature set to obtain a subset of high-value literature based on the preset screening rules after parameter adjustment. The high-value document subset is subjected to association structure analysis to generate document association information; Based on a large language model, a research decision analysis is performed on the subset of high-value documents and the related information of those documents to generate research decision analysis results that match the user's search needs.
2. The document mining system based on scenario-driven and association analysis as described in claim 1, characterized in that, Methods for obtaining the corresponding initial literature set include: The journal names selected by the user are standardized and mapped to the corresponding ISSN list. Using the ISSN list as a filter, a structured query request is constructed based on the keywords and publication year range in the search command. An initial set of documents that meet the filter conditions and query request is retrieved from the external database.
3. The document mining system based on scenario-driven and association analysis as described in claim 1, characterized in that, The parameters for adjusting the preset filtering rules include: In the research gap mining scenario, the weight coefficients of citation frequency, publication year, and journal score are assigned according to their importance in the research gap mining scenario, and the specified number of high-value literature subsets is set to 30%~50% of the total initial literature set; In the methodological evaluation scenario, the weight coefficients of citation frequency, publication year, and journal score are assigned according to their importance in the methodological evaluation scenario, and the specified number of high-value literature subsets is set to 10%~20% of the total initial literature set. In the scenario of predicting innovative topics, the citation frequency threshold is adjusted to the average citation frequency of all documents within N years prior to the retrieval time, and the range of publication years is expanded to N years prior to the retrieval time. The weight coefficients of citation frequency, publication year, and journal score are assigned according to the importance of citation frequency, publication year, and journal score in the scenario of predicting innovative topics. The specified number of high-value document subsets is set to 40% to 60% of the total initial document set.
4. The document mining system based on scenario-driven and association analysis as described in claim 3, characterized in that, The preset screening rules of the decision analysis module include: calculating a comprehensive score for each document in the initial document set using a weighted scoring method, arranging them in descending order of comprehensive score, and selecting a specified number of documents at the top of the sort to obtain a high-value document subset; the calculation method of the weighted scoring method includes: document score = w1 × normalized citation frequency + w2 × normalized publication year + w3 × normalized journal score, where w1 is the weight coefficient of citation frequency, w2 is the weight coefficient of publication year, w3 is the weight coefficient of journal score, and the sum of the weight coefficients w1, w2, and w3 is 1.
5. The document mining system based on scenario-driven and association analysis as described in claim 3, characterized in that, The decision analysis module is also used to perform quality adjustment on the generated scientific research decision analysis results. The quality adjustment method includes: evaluating the completeness, consistency, and confidence of the scientific research decision analysis results respectively; when the completeness is lower than a preset threshold, expanding the size of the high-value literature subset; when the consistency is lower than a preset threshold, increasing the weight coefficient of citation frequency and expanding the range of publication years; when the confidence is lower than a preset threshold, adjusting the parameters of the association structure analysis; the decision analysis module regenerates the high-value literature subset and the scientific research decision analysis results according to the adjusted relevant parameters, and performs quality adjustment again, iterating until the consistency, completeness, and confidence all reach the corresponding preset thresholds.
6. The document mining system based on scenario-driven and association analysis as described in claim 1, characterized in that, It also includes a journal selection module, which is used to filter the journals in the journal dictionary to obtain a filtered journal dictionary. The filtering method includes: retaining only the journals that rank in the top set proportion in each subject area, and the citation half-life of the journals is greater than or equal to the set time limit.
7. The document mining system based on scenario-driven and association analysis as described in claim 1, characterized in that, The process of performing association structure analysis on a subset of high-value documents and generating document association information includes: Extract the reference list and cited list of each document in the high-value document subset, and construct a directed citation network among the documents; Based on the directed citation network, the link analysis algorithm is used to calculate the citation influence of each document, and the core document nodes are determined according to the citation influence. Based on the directed citation network, several document clusters are obtained by dividing the network using a community detection algorithm, and the shortest path algorithm is used to extract the key citation paths between the document clusters. The core document nodes, document clusters, and citation paths between clusters obtained from the analysis are aggregated and statistically analyzed to form a structured descriptive text, which serves as document association information.
8. The document mining system as described in claim 1, characterized in that, Methods for generating research decision analysis results that match user search needs include: Extract keyword features from user search instructions, perform word segmentation on the search instruction text, remove stop words and punctuation marks, retain content words as keyword features, match the keyword features with a predefined database of multiple scenario keywords, and identify scenarios where at least one keyword is matched as the identified scientific research decision-making scenarios. Based on the identified scientific research decision-making scenario, the corresponding set of analysis subtasks is retrieved and decomposed from the pre-stored task template library. Each analysis subtask in the set of analysis subtasks includes prompt word templates, output format requirements, recommended large language models, and dependencies of preceding tasks. Based on the dependencies of the preceding tasks in each analysis subtask, the execution order of each subtask is resolved and the execution mode is determined. Subtasks with dependencies on preceding tasks are executed serially, while subtasks without dependencies on preceding tasks are executed in parallel. The abstract and title text of each document in the high-value document subset are extracted. The long text is cut into text blocks that do not exceed the length of the current large language model context window by using text segmentation and semantic compression algorithms. Meaningless punctuation and repeated sentences are removed and concatenated into a continuous string according to the document sorting order. According to the parsed execution order and execution mode, each analysis subtask is executed sequentially or in parallel. When executing any analysis subtask, the continuous string, the literature association information, the prompt word template of the current analysis subtask, and the output format requirements are sent to the recommended large language model server. The recommended large language model is called to generate the original analysis result that meets the output format requirements as the original analysis result of the subtask. After all the analysis subtasks have been completed, the raw analysis results of each subtask are concatenated in the order of execution. Then, the concatenated results are integrated according to a preset JSON array, category tag page, or card list format to generate the final scientific research decision analysis result.
9. A document mining method based on scenario-driven and association analysis according to the system described in claim 1, characterized in that, include: The initial set of documents is retrieved from the filtered journal dictionary based on the user's search instructions. The research decision-making scenario is identified based on the user's search instructions. The parameters of the preset screening rules are adjusted according to the research decision-making scenario. The initial literature set is then screened according to the preset screening rules after the parameters are adjusted to obtain a subset of high-value literature. The high-value document subset is subjected to association structure analysis to generate document association information; Based on a large language model, a research decision analysis is performed on the subset of high-value documents and the related information of those documents to generate research decision analysis results that match the user's search needs.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the document mining method based on scene-driven and association analysis as described in claim 9.