Research process planning method and device, electronic equipment and readable storage medium

Through the understanding and analysis of key research information and article retrieval, an index structure is constructed to determine the research direction, and the problems of low efficiency and poor accuracy of literature screening and research process planning in scientific research are solved, achieving more efficient and intelligent scientific research support.

CN120086353APending Publication Date: 2025-06-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510238110.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the field of scientific research, researchers face difficulties in finding, screening and understanding a large number of documents. Traditional methods are time-consuming and labor-intensive and are prone to missing important information. Research process planning requires a comprehensive consideration of a variety of factors, and professional knowledge and experience requirements are high.

Method used

By obtaining key research information, comprehension analysis is carried out to determine search phrases, search articles based on search phrases, construct an index structure, determine the target research direction, and plan the research process based on this direction.

Benefits of technology

It has improved the intelligence and efficiency of scientific research process planning, provided comprehensive and accurate research support, accelerated the scientific research process, and improved scientific research efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086353A_ABST
    Figure CN120086353A_ABST
Patent Text Reader

Abstract

The invention discloses a research process planning method and device, electronic equipment and a readable storage medium, and relates to the technical field of scientific research, and the method comprises the steps: carrying out the understanding analysis of key research information, carrying out the article retrieval according to a determined search phrase, and obtaining a target article set, and determining a target research direction according to the key research information and the target article set after the structured hierarchical representation, and planning a research process based on the target research direction, so that the problem that the research process is difficult to carry out during literature review and research process planning in the professional technical field can be solved. The scientific research process planning method solves the technical problems that the scientific research process planning method depends on manual retrieval and analysis of researchers and has high professional quality requirements on the researchers, and the technical effects of improving the intelligence degree and efficiency of scientific research process planning, providing comprehensive and accurate research support for the researchers, accelerating the scientific research process and improving the scientific research efficiency and quality are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of scientific research, and particularly to a method, apparatus, electronic device, and readable storage medium for planning research processes. Background Art

[0002] In today's scientific research field, the number of academic publications has been growing exponentially. This trend has brought many challenges to scientific research work, especially in literature reviews and research process planning. The huge number of literatures makes it extremely difficult for researchers to search, screen, and understand relevant research results. In traditional literature review methods, researchers need to manually search academic databases, read a large number of literatures, and screen out information related to their own research. This process not only consumes a large amount of time and energy but also easily misses important information. When planning research processes, researchers also need to comprehensively consider various factors, including experimental design, data analysis methods, etc., which requires extremely high professional knowledge and experience of researchers. Summary of the Invention

[0003] This application provides a method, apparatus, electronic device, and readable storage medium for planning research processes to at least solve the problems of low efficiency and poor accuracy in related technologies.

[0004] This application provides a method for planning a research process, including:

[0005] Obtaining key research information, understanding and analyzing the key research information, and determining a search phrase based on the result of the understanding and analysis;

[0006] Performing article retrieval based on the search phrase to obtain a set of target articles, and constructing a corresponding index structure based on each article in the set of target articles;

[0007] Determining a target research direction based on the key research information and the index structure, and performing planning based on the target research direction to obtain a target research process.

[0008] This application also provides an apparatus for planning a research process, including:

[0009] An information acquisition module, configured to obtain key research information, understand and analyze the key research information, and determine a search phrase based on the result of the understanding and analysis;

[0010] An article search module, configured to perform article retrieval based on the search phrase to obtain a set of target articles, and construct a corresponding index structure based on each article in the set of target articles;

[0011] A process planning module, configured to determine a target research direction based on the key research information and the index structure, and perform planning based on the target research direction to obtain a target research process.

[0012] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above research process planning methods when executing the computer program.

[0013] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above research process planning methods when executed by a processor.

[0014] The present application also provides a computer program product including a computer program, which implements the steps of any of the above research process planning methods when executed by a processor.

[0015] Through the present application, by understanding and analyzing key research information, retrieving articles according to the determined search phrases to obtain a set of target articles, determining a target research direction based on the key research information and the set of target articles represented in a structured hierarchy, and planning a research process based on the target research direction, it is possible to solve the technical problem of relying on manual retrieval and analysis by researchers and having high requirements for the professional qualities of researchers when facing literature reviews and research process planning in the field of professional technology, achieving the technical effects of improving the intelligence level and efficiency of scientific research process planning, providing comprehensive and accurate research support for scientific researchers, accelerating the scientific research process, and improving the efficiency and quality of scientific research. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of a research process planning method provided by an embodiment of the present application;

[0018] Figure 2 It is a schematic flowchart of another research process planning method provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic diagram of the ARIA composition structure of another research process planning method provided by an embodiment of the present application;

[0020] Figure 4 It is a schematic flowchart of yet another research process planning method provided by an embodiment of the present application;

[0021] Figure 5 It is a structural block diagram of a research process planning device provided by an embodiment of the present application;

[0022] Figure 6 It is a schematic diagram of the hardware structure of the computer device provided by the embodiments of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0024] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0025] With the development of artificial intelligence technology, large language models (LLMs) have gradually become powerful tools. An LLM is an artificial intelligence model based on deep learning, usually constructed based on the Transformer architecture, and is designed to understand and generate natural language. The Transformer architecture is a deep learning model architecture that performs well in natural language processing tasks. The Transformer architecture uses the attention mechanism to capture the relationships between different parts of the input text, so as to better understand and generate language. LLMs learn knowledge about grammar, semantics, and pragmatics of language through unsupervised learning on large-scale text data, and thus have powerful information processing and generation capabilities, and can help researchers handle complex tasks to a certain extent. In some general knowledge Q&A tasks, LLMs perform well and can quickly and accurately answer various questions. Therefore, technical personnel in the scientific research field use LLMs for literature retrieval and research process planning.

[0026] However, the scientific research field is highly specialized, involving a large number of professional terms and complex concepts. When dealing with tasks in these professional fields, due to the relatively scarce professional data in its training dataset, which is mostly general content, it is difficult for the LLM to accurately understand and generate content involving professional terms and concepts, so there are limitations. To address the above issues, related technologies have tried various techniques to improve the application of the LLM in professional fields, such as In-Context Learning (ICL), Fine-Tuning (FT), and Retrieval-Augmented Generation (RAG), etc. Among them, ICL learns to perform specific tasks by using the given input-output examples as context, so it relies on carefully designed prompts to provide rich context for the LLM. However, the accuracy and structuring degree of the prompts are required to be relatively high, and it is affected by the model context limit; FT refers to further training the model on the basis of a pre-trained model to make it adapt to specific tasks or datasets. However, it requires a large amount of computing resources, suitable training datasets, and access to model weights or fine-tuning APIs, with high costs and risks such as overfitting; RAG is a model that combines language model and information retrieval technologies. When generating text or answering questions, it will first retrieve relevant information from a large collection of documents, and then use this information to guide text generation to improve the quality and accuracy of prediction. Therefore, it requires a predefined dataset, and this dataset needs to be frequently maintained and updated. At the same time, it has relatively high requirements for context processing capabilities when dealing with complex documents.

[0027] In addition, aiming at the deficiencies of the LLM in professional field applications, related technologies have also been committed to developing domain-specific automated tools and systems. For example: ① By customizing prompts and domain-specific training, enabling the LLM to generate Boolean search queries for system evaluation. Although it improves the capabilities of the LLM to a certain extent, it highly depends on the construction of manual prompts, and the quality of the prompts directly affects the accuracy and effectiveness of the results. Moreover, this method is often limited to specific academic databases and fields and lacks generality.

[0028] ②Utilize other models in combination with LLM to improve the capabilities of LLM. For example, in related research, LLM-based frameworks focus on complex technical tasks in specific fields, such as Building Information Modeling (BIM), chip design, and chemical synthesis. Frameworks like BIM and ChipGPT integrate LLM into specialized modules and combine additional resources and tools to achieve functions such as automatic prompt generation, enhanced information retrieval, and code generation. This can effectively reduce processing time and manual workload, but they have poor generality and are difficult to apply to other fields. For interdisciplinary researchers, they need to master specific frameworks in multiple different fields, which undoubtedly increases the complexity and cost of research. Moreover, these frameworks still have problems such as low efficiency and inability to fully integrate multi-source information when dealing with large-scale literature retrieval and complex research process planning.

[0029] In summary, the related tools and systems in the prior art have problems such as poor generality, reliance on manual operations, inability to efficiently process large-scale literature and complex research tasks. Therefore, the research process planning method provided in the embodiments of the present application constructs an intelligent, efficient, and general multi-LLM framework to achieve the effect of automatically and accurately retrieving and screening literature, converting literature information into an operable research process, and providing comprehensive and accurate research support for researchers.

[0030] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the research process planning method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0032] The embodiments of the present application provide a research process planning method, which can be used in computers, etc. Figure 1 It is a flowchart of the research process planning method according to the embodiments of the present application. As Figure 1 shown, the process includes the following steps:

[0033] Step S101, obtain key research information, conduct understanding and analysis on the key research information, and determine a search phrase based on the results of the understanding and analysis.

[0034] Specifically, in the embodiments of the present application, based on the Artificial Intelligence Research Assistant (ARIA), a multi-LLM framework based on four agents is constructed, aiming to simulate the collaboration of an expert team and systematically replicate the human work research process, so as to achieve the automation of literature retrieval, screening, processing, and research process generation. Among them, the four agents are: Conversation Agent, Retriever Agent, Processor Agent, and Suggester Agent. The Conversation Agent is used to interact with the user, obtain the key research information input by the user, or present the finally planned research process to the user; the Retriever Agent is used to conduct article retrieval to obtain literature related to the research; the Processor Agent is used to process the retrieved articles to maintain the consistency of the article content; the Suggester Agent is used to generate the final research process. Each of the four agents performs its own duties and collaborates closely, and can comprehensively handle each link from user demand understanding to research process generation. This multi-agent collaboration mode improves the intelligence and adaptability of the system.

[0035] In some alternative embodiments, the Conversation Agent serves as the main interface for the user to interact with ARIA and drives the execution of ARIA through four main functions, including: information acquisition function, understanding and analysis function, search phrase creation function, and visualization display function. In the information acquisition function, the user elaborates on the research topic, research objectives, and key research information of interest through the guided text input box provided by the Conversation Agent. In the understanding and analysis function, the LLM deployed by the Conversation Agent processes and analyzes the obtained key research information to generate an understanding and analysis, thereby providing a context basis for the subsequent ARIA process. In the search phrase creation function, ARIA uses the generated understanding and analysis results to provide a series of natural language search phrases, which can capture the essence of the user's key research questions and lay a foundation for creating Boolean queries applicable to academic databases. At the end of literature retrieval and analysis, when ARIA generates a research process, the visualization display function of the Conversation Agent will also provide the user with a series of research specifications and parameter lists according to the research process and deliver the research process to the user.

[0036] In some alternative embodiments, taking the research on enhancing dropwise condensation through electro-wetting and shear gas flow techniques in surface engineering as an example, the user first interacts with the dialogue agent, inputting the research topic, research objectives, key technologies, etc., such as "Application of dropwise condensation in surface engineering", "Enhancing dropwise condensation to improve heat transfer efficiency", "Achieving droplet detachment through electro-wetting and shear gas flow, and studying electrode design and arrangement", etc. The dialogue agent generates an understanding analysis based on these inputs and provides search phrases, such as "Improving the efficiency of droplet detachment in condensation using electro-wetting technology", "Influence of electrode design and arrangement in electro-wetting on condensation applications", etc. These are only examples and not limited thereto.

[0037] Step S102: Conduct an article retrieval based on the search phrases to obtain a set of target articles, and construct a corresponding index structure for each article in the set of target articles.

[0038] Specifically, in the embodiments of the present application, the retrieval agent in ARIA uses the search phrases obtained by the dialogue agent as a basis to interface with multiple literature databases for article retrieval and screening. During the article retrieval process, the search phrases are equivalent to keywords, and by gradually matching with the article title, abstract, keyword fields, and even the full text, the literature that best matches the current user's research topic or research objectives, etc., is obtained. For example, the retrieval agent performs expansion and Boolean query conversion based on the above search phrases such as "Improving the efficiency of droplet detachment in condensation using electro-wetting technology", "Influence of electrode design and arrangement in electro-wetting on condensation applications", etc., and then conducts a retrieval in multiple academic databases. During this period, the user can also select the determined search phrases or input custom phrases again.

[0039] In some alternative embodiments, after the retrieval is completed and the retrieval agent can obtain the full text information of the target articles and store them, in order to achieve fast and efficient query of specific information in the technical field, the processing agent processes the literature. The processing agent in the embodiments of the present application creates an index structure for each article through LlamaIndex, and this index provides a structured hierarchical representation for the literature. Among them, LlamaIndex (also known as GPT Index) is a powerful Python framework designed to help developers use large language models (LLMs) more easily. The framework provides various types of index structures to adapt to different application scenarios and query requirements. For example, a vector index (such as GPT Simple VectorIndex) converts text data into vector embeddings for easy similarity search; a tree index (GPT Tree Index) can hierarchically organize data and is suitable for handling situations that require summarization or hierarchical queries; a graph index (GPT Graph Index) can handle complex relational data to construct a knowledge graph, etc.

[0040] Step S103: Determine the target research direction based on the key research information and the index structure, and perform planning based on the target research direction to obtain the target research process.

[0041] Specifically, in the embodiment of the present application, after uniformly constructing the index structure for the retrieved articles, taking the original key research information as the problem and the articles in the index structure as the materials, it is recommended that the agent sort out according to the mining of the problem and materials, classify and organize the collected literature, and summarize it according to dimensions such as research theme, research method, and research conclusion, so as to obtain different research directions. Further, after determining the research direction, the built-in LLM can be used to directly perform deeper analysis according to the research direction to obtain the research methods, research steps, research tools, parameters of interest, etc. corresponding to the research direction, so as to generate a complete research process. Users can directly carry out subsequent scientific research work according to the research process. For example, after retrieving in multiple academic databases according to search phrases such as "using electrowetting technology to improve the droplet detachment efficiency in condensation" and "the influence of electrode design and arrangement in electrowetting on condensation applications", the processing agent indexes the articles, and the recommendation agent generates a research process according to the processed information, including steps such as the selection of materials and equipment, experimental setup, data collection and analysis, etc. This is only an example and is not limited thereto.

[0042] In some alternative embodiments, ARIA provided by the embodiments of the present application can understand the research needs in different fields through the interaction between the dialogue agent and the user, and use multi-agent collaboration and a flexible index structure to provide effective support for research in different fields, with good versatility. Whether it is research in surface engineering, chemical synthesis or other fields, the embodiments of the present application can retrieve relevant literature according to the user's needs and generate corresponding research processes, meet the personalized needs of different researchers, provide more targeted research support, help researchers quickly understand the basic knowledge of the research field, obtain relevant research resources, and generate feasible research processes, reduce the learning cost and technical threshold of researchers in cross-field research, reduce the dependence of researchers on professional knowledge, and enable non-professional researchers to carry out research work more easily.

[0043] Through the present application, by understanding and analyzing the key research information, retrieving articles according to the determined search phrases to obtain the target article set, determining the target research direction based on the key research information and the target article set represented by the structured hierarchy, and planning the research process based on the target research direction, it is possible to solve the technical problem of relying on manual retrieval and analysis by researchers and having high requirements for the professional quality of researchers when facing literature reviews and research process planning in professional technical fields, achieve the technical effects of improving the intelligent level and efficiency of scientific research process planning, providing comprehensive and accurate research support for scientific research personnel, accelerating the scientific research process, and improving the scientific research efficiency and quality.

[0044] Embodiments of the present application also provide a method for planning a research process, which can be used in computers and the like. Figure 2 It is a flowchart of the method for planning a research process according to an embodiment of the present application, as Figure 2 shown. The process includes the following steps:

[0045] Step S201, obtain key research information, understand and analyze the key research information, and determine a search phrase based on the understanding and analysis results. For details, please refer to Figure 1 step S101 of the embodiment shown, which will not be elaborated here.

[0046] Step S202, perform article retrieval based on the search phrase to obtain a set of target articles, and construct a corresponding index structure based on each article in the set of target articles.

[0047] Specifically, the above step S202 includes:

[0048] Step S2021, retrieve in the local database according to the search phrase to obtain a first set of articles and the digital object unique identifier or title corresponding to each article.

[0049] Specifically, in the embodiments of the present application, the search phrase is the keyword for article retrieval. However, the same object may have different expression forms in different articles. Therefore, the search phrase is extended to ensure the accuracy of article retrieval.

[0050] In some alternative embodiments, the above step S2021 includes:

[0051] Step a1, generate semantic extension phrases according to the search phrase, and determine a phrase list based on the semantic extension phrases.

[0052] Step a2, generate a retrieval expression according to the phrase list, and perform retrieval in the local database based on the retrieval expression.

[0053] Specifically, in the embodiments of the present application, the retrieval agent uses the built-in LLM for query expansion and syntax conversion. First, the LLM expands according to the search phrase to obtain variants of the search phrase, that is, semantic extension phrases, so as to obtain the final phrase list for article retrieval. Second, these phrases are converted into retrieval expressions. The embodiments of the present application perform article retrieval according to the Boolean query method. Therefore, the phrases are converted into Boolean expressions, but not limited thereto. Among them, the Boolean query is an information retrieval query method based on Boolean logic, which is commonly used in systems such as databases and search engines. By combining multiple retrieval terms through logical operators, the required information can be more accurately located and filtered. The logical operators include: AND, OR, and NOT.

[0054] In some alternative embodiments, after determining the Boolean expression, as Figure 3 shown, a preliminary search is first performed in the local database. The local database is pre-constructed for the user or the institution where the user belongs. By representing relevant literature in the field as vectors, its own vector library is constructed. Retrieving using the expanded Boolean expression is equivalent to performing semantic retrieval, which can query relevant literature in the local database to obtain the first article set and the Digital Object Unique Identifier (DOI) or title of the articles. Among them, DOI is a mechanism for identifying digital resources, aiming to solve problems such as the identification, location, and persistent linking of digital resources.

[0055] In some alternative embodiments, the construction process of the local database includes: ① Literature data collection and preprocessing. Collect literature related to the target field from various academic databases, institutional repositories, industry report websites, etc., remove noise information in the literature, such as advertisements, irrelevant special characters, garbled codes, etc. For HTML tags, Markdown format tags, etc. in the text, appropriate processing is also required to convert them into plain text form, convert all characters in the text into a unified case form, and perform word segmentation on the text; ② Select a suitable vector representation model, such as the bag-of-words model, TF-IDF (Term Frequency-Inverse Document Frequency), word embedding model, pre-trained language model, etc.; ③ Based on the selected vector representation model, convert the text into vectors, select a suitable database, and store the generated literature vectors together with the relevant metadata of the literature. At the same time, for the convenience of subsequent retrieval and management, a unique identifier can also be assigned to each vector. In addition, the local database needs to be updated and maintained to keep it synchronized with the development of technology.

[0056] Step S2022, retrieve in the first open-source database according to the Digital Object Unique Identifier or title, obtain the full-text information of each article in the first article set, and determine the target article set according to the full-text information.

[0057] Specifically, in the embodiment of the present application, because the scope of the local database is limited, the full-text information of each article in the first article set is retrieved in the first open-source database according to the DOI or title. The first open-source database includes: ScienceDirect, Web of Science, and Scopus, but is not limited thereto. Based on the full-text information of the articles obtained in the preliminary search, the retrieval agent improves the retrieval accuracy through a three-stage algorithm: the snowball stage, the article frequency ranking stage, and the semantic screening stage.

[0058] In some alternative embodiments, the above step S2022 includes:

[0059] Step b1: Determine the citations and / or references of each article in the first article set according to the full-text information, and determine the second article set according to the citations and / or references.

[0060] Step b2: Rank the cross-reference frequencies of the articles in the second article set, and filter out the articles in the second article set whose cross-reference frequencies exceed a preset frequency threshold to obtain a third article set.

[0061] Step b3: Perform semantic similarity analysis on the articles in the third article set according to the phrase list, and filter out the articles in the third article set whose similarity exceeds a preset similarity threshold to obtain a target article set.

[0062] Specifically, in the embodiment of the present application, in the snowball stage, according to the full-text information of each article in the first article set, trace the citations and references of the article, so as to include the citations and references in the retrieval scope, expand the literature search scope, and obtain the second article set.

[0063] In some alternative embodiments, in the article frequency ranking stage, the citation frequencies of each article are counted. Therefore, different articles may cite the same article at the same time, so there will be a situation of cross-reference. Since the number of times an article is cited can evaluate its influence and quality. Generally speaking, the more times an article is cited, the higher the degree of attention and recognition it receives in this field. For example, in academic research, highly cited papers are usually considered to be research results with important contributions. Therefore, in the embodiment of the present application, by ranking the cross-reference frequencies of each article, the articles with higher cross-reference frequencies are filtered out, so as to obtain the third article set.

[0064] In some alternative embodiments, in the semantic screening stage, based on the initial phrase list used for retrieval, the SciBERT model is used to perform semantic similarity ranking on the abstracts of each article in the third article set, and articles with higher similarity are screened out to obtain the final target article set. Among them, SciBERT is a pre-trained language model specifically designed for scientific text processing. Based on the BERT architecture, SciBERT usually adopts a multi-layer bidirectional Transformer encoder. By encoding the input text, it generates context-related word vector representations, which can capture the semantic information of words in different contexts and provide strong support for subsequent natural language processing tasks. By expanding the search scope, the embodiments of the present application can simulate the literature retrieval thinking of researchers, perform automatic and accurate article retrieval in multiple academic databases, and shorten the time for literature reviews. At the same time, through two screenings, the redundancy of article retrieval results can be avoided, and the most accurate, most similar, and most valuable reference documents can be provided for users, providing strong support for research work. In addition, in the article retrieval process of the embodiments of the present application, not only can the search phrases be expanded, converted into Boolean queries, and retrieved in multiple academic databases, but also through multi-stage algorithms such as snowball screening, frequency ranking, and semantic screening, it is ensured that the most relevant literature can be obtained. This comprehensive retrieval and screening strategy can obtain literature information more comprehensively and accurately, reducing information omission and redundancy.

[0065] Step S2023, perform full-text retrieval on each article in the target article set based on the second open-source database to obtain the full-text information of each article in the target article set.

[0066] Specifically, in the embodiments of the present application, the retrieval agent finally performs full-text retrieval on the articles in the target article set through the Elsevier Full-Text Retrieval API and the arXiv API to obtain the full-text information of each article. The Elsevier Full-Text Retrieval API and the arXiv API are two different academic literature retrieval services. The database corresponding to the Elsevier Full-Text Retrieval API contains journal articles in multiple disciplinary fields such as computer science, engineering technology, energy science, environmental science, mathematics, physics, chemistry, astronomy, medicine, life science, and social science; arXiv is an open-access academic preprint repository that mainly collects papers in the fields of physics, mathematics, computer science, biology, and mathematical economics. A preprint refers to the first draft of a paper that has not yet undergone peer review and may not have been published in a formal academic journal. The arXiv API allows users to retrieve these preprint documents programmatically and is suitable for researchers who need to quickly obtain the latest research results.

[0067] In step S2024, perform data cleaning on the full-text information of each article in the target article set; generate a corresponding combined graph index based on the full-text information after data cleaning.

[0068] Specifically, in the embodiments of the present application, as Figure 3 shown, after full-text article retrieval and storage, the processing agent performs two main tasks: data cleaning and data indexing. In the data cleaning stage, the processing agent removes unnecessary elements (such as metadata, tags, etc.), optimizes storage, and ensures the consistency of all stored article content. To enable fast and efficient querying of specific information in the technical field, the LlamaIndex data management framework is used to process the literature. A queryable, RAG-based combined graph index is created through LlamaIndex, which provides a structured hierarchical representation for the literature, including an abstract representation and a detailed representation. Each filtered article is assigned an abstract representation and a corresponding detailed representation. The abstract representation serves as a high-level abstraction of the article, and the detailed representation encodes all the content of the article as a vector embedding. Among them, RAG is Retrieval Augmented Generation, a technology that combines information retrieval and language generation, aiming to enable the model to utilize information from external knowledge sources when generating text, thereby improving the quality, accuracy, and reliability of the generated content. In the embodiments of the present application, ARIA deeply analyzes the filtered literature and uses the LlamaIndex tool to create a dynamic and customized dataset to achieve user-targeted, specific-domain RAG. This index structure can effectively manage and retrieve large-scale scientific texts, enabling fast positioning and retrieval from high-level abstracts to detailed content, thereby improving the efficiency and accuracy of information retrieval and providing a solid data foundation for the generation of research processes.

[0069] In step S203, determine the target research direction based on the key research information and the index structure, and perform planning based on the target research direction to obtain the target research process. For details, please refer to Figure 1 step S103 of the embodiment shown herein, which will not be elaborated herein.

[0070] Through the present application, by understanding and analyzing the key research information, retrieving articles according to the determined search phrases to obtain the target article set, determining the target research direction based on the key research information and the target article set with a structured hierarchical representation, and planning the research process based on the target research direction, it is possible to solve the technical problem of relying on researchers' manual retrieval and analysis and having high requirements for researchers' professional qualities when facing literature reviews and research process planning in the professional technical field, achieving the technical effects of improving the intelligence level and efficiency of scientific research process planning, providing comprehensive and accurate research support for scientific researchers, accelerating the scientific research process, and improving the efficiency and quality of scientific research.

[0071] Embodiments of the present application also provide a method for planning a research process, which can be used in computers and the like. Figure 4 It is a flowchart of the method for planning a research process according to an embodiment of the present application, as Figure 4 shown. The process includes the following steps:

[0072] Step S401: Obtain key research information, conduct an understanding and analysis of the key research information, and determine a search phrase based on the results of the understanding and analysis. For details, please refer to Figure 2 Step S201 of the embodiment shown herein, which will not be elaborated herein.

[0073] Step S402: Conduct an article search based on the search phrase to obtain a set of target articles, and construct a corresponding index structure based on each article in the set of target articles. For details, please refer to Figure 2 Step S202 of the embodiment shown herein, which will not be elaborated herein.

[0074] Step S403: Determine a target research direction based on the key research information and the index structure, and conduct planning based on the target research direction to obtain a target research process.

[0075] Specifically, the above step S403 includes:

[0076] Step S4031: Construct a research framework based on the key research information, and generate an instruction prompt based on the research framework.

[0077] Specifically, in the embodiments of the present application, it is recommended that the agent integrate the key research information provided by the user into the research blueprint as research elements to build a research framework, which serves as a centralized information source. The process is as follows: First, extract the core research purpose from the key research information, refine the research purpose into specific research questions, which will guide the direction of the entire research. Second, sort out the existing research results related to the research questions in chronological order, that is, the retrieved articles, clarify the main research results and development trends at different stages, and pay attention to the current research hotspots and frontier issues. For example, it can be judged from the year distribution of literature publication, the concentrated directions of research institutions and scholars, etc., so as to obtain the research status, existing theories and methods in this field. Classify and summarize the literature, such as by research theme, research method, research object, etc., to extract the theoretical basis involved in the literature, analyze the core viewpoints, application scopes and limitations of different theories, draw a theoretical framework diagram or list for comparison, clearly present the relationships and characteristics of each theory, and thus determine the advantages and disadvantages of the existing research. At the same time, summarize the research methods adopted in the literature, such as experimental method, survey method, literature research method, etc., analyze the specific operation processes, advantages and disadvantages and application scenarios of each method, and compare the application effects of different methods in the research, so as to provide a theoretical basis and research gaps for the user's research. Furthermore, construct a special instruction prompt based on the theoretical basis and research gaps, etc., which is used as an input query to guide the main query engine to perform the RAG process.

[0078] Step S4032, traverse and retrieve the index structure based on the instruction prompt to determine the target research direction.

[0079] Specifically, in the embodiments of the present application, the obtained instruction prompt includes the guidance of using blueprint context, identifying innovative methods, and synthesizing the retrieved information into the final research process. The main query engine traverses the graph hierarchy of the combined graph index according to the instruction prompt, searches and retrieves relevant technical information, and uses the built-in LLM (such as OpenAI GPT model: gpt-4) to synthesize the information returned from the graph into comprehensive research tools, parameters of interest and their ranges, etc., to determine the final research process. The specific process is a conventional technical means in this field and will not be elaborated here.

[0080] In some alternative embodiments, it is recommended that the agent construct a special instruction prompt based on the retrieved information to guide the main query engine to search and retrieve relevant information in the index, and synthesize this information into a detailed and comprehensive research process, which includes key information such as materials, equipment, steps required for the experiment, and data analysis methods. For example, in the case study of dropwise condensation, the generated research process covers all aspects from material selection to data characterization, providing clear research guidance for researchers and helping to improve the accuracy and reliability of the research.

[0081] In some alternative embodiments, the ARIA provided by the present application has good versatility and can be applied to a variety of technical fields, such as:

[0082] ① In the field of medical research, the literature retrieval and analysis capabilities of ARIA can help researchers quickly obtain the latest medical research results, understand the pathogenesis of diseases, treatment methods, and drug R & D progress. By screening and integrating a large number of medical literatures, it can provide comprehensive research background information for medical researchers and assist them in designing more effective experimental plans and clinical studies. During the drug R & D process, researchers can use ARIA to retrieve relevant literatures on drug target research, drug clinical trial data, etc., providing a basis for the design and development of new drugs. At the same time, the ability of ARIA to generate research processes can help medical researchers plan experimental steps, select appropriate experimental methods and data analysis means, improving the efficiency and quality of medical research.

[0083] ② In the field of materials science, researchers need to continuously explore new material properties and applications. ARIA can help researchers retrieve and analyze literatures in the field of materials science, understanding the relationships between the structures, properties, and preparation methods of different materials. Through in - depth analysis of the literatures, ARIA can provide research ideas and experimental plans for materials science researchers. For example, it can guide researchers in selecting appropriate material components and preparation processes to obtain materials with specific properties. When researching new superconducting materials, ARIA can retrieve relevant literatures, analyze the characteristics and preparation methods of existing superconducting materials, providing references for the R & D of new superconducting materials and accelerating the R & D process of new materials.

[0084] ③ In the field of environmental science, ARIA can be used to retrieve and analyze literatures in the field of environmental science, understanding the sources, transmission routes, and treatment methods of environmental pollution. By integrating a large number of environmental science literatures, ARIA can provide comprehensive research background information for environmental science researchers, helping them design more effective environmental monitoring and treatment plans. When researching air pollution control, ARIA can retrieve relevant literatures on air pollutant emission data, the impact of meteorological conditions on pollutant diffusion, etc., providing a basis for formulating air pollution control strategies. At the same time, the ability of ARIA to generate research processes can help environmental science researchers plan experimental steps, select appropriate monitoring methods and data analysis means, improving the efficiency and quality of environmental science research.

[0085] ④In the field of social sciences, researchers are faced with a large amount of literature and complex research questions. ARIA can help researchers quickly retrieve and screen relevant literature, understand the essence and laws of social phenomena. Through the analysis of social science literature, ARIA can provide research ideas and methods for social science researchers. For example, it can guide researchers in designing questionnaires, selecting appropriate research objects, and data analysis methods. When studying social inequality issues, ARIA can retrieve relevant sociological and economic literature, analyze the impact of different factors on social inequality, and provide references for formulating relevant policies.

[0086] The ARIA provided by this application can complete the key literature review task within one hour through an automated workflow, utilizing multi-agent collaboration and multi-database retrieval, significantly improving the efficiency of the literature review. Moreover, it can not only expand the database query scope selected by the user, but also query multiple database APIs simultaneously, screen thousands of articles through the developed screening algorithm, shorten the time of the literature review, enable researchers to obtain relevant literature information faster, and provide strong support for research work.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0088] The embodiments of this application also provide a device for planning a research process. This device is used to implement the above embodiments and preferred implementation methods, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware is also possible and contemplated.

[0089] This embodiment provides a device for planning a research process, as Figure 5 shown, including:

[0090] An information acquisition module 501, configured to acquire key research information, perform understanding and analysis on the key research information, and determine a search phrase based on the results of the understanding and analysis.

[0091] An article search module 502, configured to perform article retrieval based on the search phrase, obtain a set of target articles, and construct a corresponding index structure based on each article in the set of target articles.

[0092] A process planning module 503, configured to determine a target research direction based on the key research information and the index structure, and perform planning based on the target research direction to obtain a target research process.

[0093] In some alternative embodiments, the article search module 502 includes:

[0094] An article identifier retrieval unit, configured to retrieve in a local database according to a search phrase, so as to obtain a first article set and a digital object unique identifier or title corresponding to each article.

[0095] An article screening unit, configured to retrieve in a first open-source database according to the digital object unique identifier or title, obtain the full-text information of each article in the first article set, and determine a target article set according to the full-text information.

[0096] A full-text retrieval unit, configured to perform full-text retrieval on each article in the target article set based on a second open-source database, so as to obtain the full-text information of each article in the target article set.

[0097] In some alternative embodiments, the article identifier retrieval unit includes:

[0098] A phrase expansion subunit, configured to generate a semantic expansion phrase according to the search phrase, and determine a phrase list based on the semantic expansion phrase.

[0099] An expression construction subunit, configured to generate a retrieval expression according to the phrase list, and perform a retrieval in the local database based on the retrieval expression.

[0100] In some alternative embodiments, the article screening unit includes:

[0101] A scope expansion subunit, configured to determine the citations and / or references of each article in the first article set according to the full-text information, and determine a second article set according to the citations and / or references.

[0102] A first screening subunit, configured to rank the cross-reference frequencies of the articles in the second article set, and screen out the articles in the second article set whose cross-reference frequencies exceed a preset frequency threshold, so as to obtain a third article set.

[0103] A second screening subunit, configured to perform semantic similarity analysis on the articles in the third article set according to the phrase list, and screen out the articles in the third article set whose similarity exceeds a preset similarity threshold, so as to obtain a target article set.

[0104] In some alternative embodiments, the apparatus further includes: an article processing module, and the article processing module includes:

[0105] A data cleaning unit, configured to perform data cleaning on the full-text information of each article in the target article set.

[0106] An index construction unit, configured to generate a corresponding combined graph index according to the full-text information after data cleaning.

[0107] In some alternative embodiments, the index construction unit includes:

[0108] A vector representation subunit, configured to divide the full-text information into abstract information and detailed information, and perform vector representations on the abstract information and the detailed information respectively.

[0109] A combined graph index construction subunit, configured to construct a combined graph index based on the generated abstract representation and detailed representation.

[0110] In some alternative embodiments, the process planning module 503 includes:

[0111] A framework construction unit, configured to construct a research framework according to the key research information, and generate an instruction prompt based on the research framework.

[0112] A literature traversal unit, configured to traverse and retrieve the index structure based on the instruction prompt to determine the target research direction.

[0113] For the description of the features in the corresponding embodiments of the apparatus for planning the research process, reference may be made to the relevant description in the corresponding embodiments of the method for planning the research process, which will not be elaborated herein one by one.

[0114] An embodiment of the present application further provides an electronic device, as Figure 6 shown, including a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above embodiments of the method for planning the research process.

[0115] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the method for planning the research process when running.

[0116] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0117] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the method for planning the research process are implemented.

[0118] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the research process planning method.

[0119] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0120] The above has introduced in detail a research process planning method, apparatus, electronic device, and readable storage medium provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for planning a research process, characterized in that: include: Obtaining key research information, performing comprehension analysis on the key research information, and determining search phrases based on the comprehension analysis results; Performing article retrieval based on the search phrase to obtain a target article set, and constructing a corresponding index structure based on each article in the target article set; A target research direction is determined according to the key research information and the index structure, and planning is performed based on the target research direction to obtain a target research process.

2. The research process planning method according to claim 1, characterized in that: The article retrieval based on the search phrase to obtain a target article set includes: Searching the local database according to the search phrase to obtain a first article set and a digital object unique identifier or title corresponding to each article; Searching a first open source database according to the digital object unique identifier or the title to obtain full-text information of each article in the first article collection, and determining the target article collection according to the full-text information; Full-text retrieval is performed on each article in the target article set based on the second open source database to obtain full-text information of each article in the target article set.

3. The research process planning method according to claim 2, characterized in that: The searching in the local database according to the search phrase includes: generating a semantically expanded phrase according to the search phrase, and determining a phrase list based on the semantically expanded phrase; A search expression is generated according to the phrase list, and a search is performed in the local database based on the search expression.

4. The research process planning method according to claim 3, characterized in that: Determining the target article set according to the full-text information includes: Determine citations and / or references of each article in the first article set according to the full-text information, and determine a second article set according to the citations and / or references; Ranking the cross-reference frequency of each article in the second article set, and screening out articles in the second article set whose cross-reference frequency exceeds a preset frequency threshold, to obtain a third article set; A semantic similarity analysis is performed on each article in the third article set according to the phrase list, and articles in the third article set whose similarity exceeds a preset similarity threshold are screened out to obtain the target article set.

5. The research process planning method according to claim 2, characterized in that: The step of constructing a corresponding index structure based on each article in the target article set includes: Performing data cleaning on the full-text information of each article in the target article set; Generate the corresponding combined graph index based on the full-text information after data cleaning.

6. The method for planning a research process according to claim 5, characterized in that: The generating of the corresponding combined graph index according to the full-text information after data cleaning includes: Convert the full text information into summary information and detailed information, and respectively perform vector representation on the summary information and the detailed information; The combined graph index is constructed based on the generated summary representation and detailed representation.

7. The research process planning method according to claim 1, characterized in that: Determining the target research direction according to the key research information and the index structure includes: constructing a research framework according to the key research information, and generating instruction prompts based on the research framework; The index structure is traversed and searched based on the instruction prompt to determine the target research direction.

8. A planning device for a research process, characterized in that: include: An information acquisition module, used to acquire key research information, perform understanding and analysis on the key research information, and determine a search phrase based on the understanding and analysis results; An article search module, used to perform article retrieval based on the search phrase, obtain a target article set, and construct a corresponding index structure based on each article in the target article set; The process planning module is used to determine the target research direction according to the key research information and the index structure, and to plan based on the target research direction to obtain the target research process.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the planning method of the research process as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the planning method for a research process as claimed in any one of claims 1 to 7.

Citation Information

Cited By

  • Auxiliary method and auxiliary device for scientific research in medical specialized field

    CN120873154A

  • Method and system for quantitative research scheme based on big language model conception

    CN122021948A

  • A method and system for conceiving a quantitative research scheme based on a large language model

    CN122021948B