Scientific research assistance system and method based on large language model
Through a scientific research assistance system based on a large language model, the metadata layer and basic computing layer are used to provide basic data, the experimental simulation layer is used to construct experimental data, and the paper generation layer is used to write the first draft of the paper. This solves the problem of lack of innovation and logic in papers in the existing system and realizes efficient and scientific scientific research assistance.
Patent Information
- Application Number
- CN202410943811.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-07-15
AI Technical Summary
The papers generated by existing scientific research assistance systems lack independent viewpoints and innovation, have poor logic and coherence, are highly templated, and cannot provide in-depth logical reasoning and theoretical support.
A scientific research assistance system based on a large language model is adopted, including a metadata layer, a basic computing layer, an experimental simulation layer, and a paper generation layer. The large language model is used to construct experimental data and write the first draft of the paper, and iterative feedback and model fine-tuning are carried out in combination with principle validity measurement and effect evaluation indicators.
The generated draft of the paper is scientific and authentic, conforms to the paper specifications, can achieve efficient content generation, and provide scientific research assistance for in-depth discussion and rigorous reasoning.
Smart Images

Figure CN118940833B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of scientific research assistance systems, and in particular to a scientific research assistance system and method based on a large language model. Background Art
[0002] Research support systems are crucial tools for assisting researchers in their scientific research. With the advancement of technology, various types of research support systems have emerged, covering multiple research processes, from literature retrieval, experimental design, and data analysis to paper writing and knowledge management. High-performance research support systems can provide reliable support for scientific research and improve its efficiency.
[0003] However, existing research support systems generate papers based solely on user-entered titles and fields. These papers lack independent perspectives and innovation, often summarizing and restating existing data. These papers fail to provide valuable references for scientific research. Furthermore, in complex and in-depth research fields, these generated papers lack in-depth logical discussion and rigorous reasoning, resulting in poor logic and coherence. Scientific research papers require rigorous logical reasoning and theoretical support, which existing research support systems struggle to fully grasp. Furthermore, the paragraphs generated by these systems are often monotonous, highly sequential, and heavily templated.
[0004] Therefore, there is an urgent need for a scientific research assistance system with excellent performance to generate a draft of a paper that is scientific, authentic, and in line with paper standards, so as to provide better scientific research assistance for scientific research. Summary of the Invention
[0005] In view of the above problems, the embodiments of the present application provide a scientific research assistance system and method based on a large language model to overcome the above problems or at least partially solve the above problems.
[0006] In a first aspect of an embodiment of the present application, a scientific research assistance system based on a large language model is disclosed, comprising:
[0007] The metadata layer, including the corpus and tool library, is used to provide basic data for the basic calculation layer and paper generation layer;
[0008] The basic computing layer is used to regularize and sort the basic data based on the pre-trained language model to obtain a usable knowledge base, which represents the callable knowledge meta-services and tool meta-services;
[0009] An experiment simulation layer, configured to construct and simulate experimental data generation based on a large language model according to user input information and the available knowledge base to obtain experimental data;
[0010] A paper generation layer, configured to generate a first draft of the paper based on a large language model, the user input information, the basic data, the experimental data, and the available knowledge base;
[0011] In the process of generating the first draft of the paper, the basic computing layer, the experimental simulation layer and the paper generation layer iteratively feedback their respective output results based on the principle validity measurement index and the effect evaluation index, so as to fine-tune the model used by themselves according to the output results.
[0012] Optionally, the available knowledge base includes: specialized knowledge graph construction, embedded representation of scientific data, literature material similarity retriever, callable tool retrieval coding, knowledge linking system, and logical chain preset by large model.
[0013] Optionally, the experimental simulation layer includes:
[0014] Investigative experiment module, including target data acquisition module and data visualization modeling module, used to generate basic experimental data;
[0015] Theoretical experiment module, including the inductive bias introduction module and the basic theoretical reasoning module, is used to generate theoretical experimental data;
[0016] The simulation experiment module includes virtual environment construction, meta-code module construction and code logic chain combination module, which is used to generate simulation experiment data.
[0017] Optionally, the paper generation layer includes:
[0018] A previous work analysis module, used for analyzing the basic data;
[0019] The paper logic sorting module is used to sort out the logical structure of the paper draft according to the user input;
[0020] Thesis formula deduction module, used to derive the formulas used in the first draft of the thesis;
[0021] A paper illustration generation module, used to generate illustrations for the first draft of the paper;
[0022] An experimental data normalization module, used for normalizing the experimental data;
[0023] An experimental data analysis module, used for analyzing the experimental data;
[0024] Reference regularization module, used to insert references in the first draft of the paper.
[0025] A second aspect of the embodiments of the present application discloses a research assistance method based on a large language model, which is applied to the research assistance system based on a large language model described in the first aspect of the embodiments of the present application. The method includes:
[0026] Obtaining user input information, wherein the user input information includes at least a paper title and a paper field;
[0027] The experiment simulation layer constructs and simulates experimental data based on the user input information and the available knowledge base provided by the basic computing layer to obtain experimental data;
[0028] The paper generation layer constructs a hierarchical index vector database based on the basic data provided by the metadata layer, and generates a first draft of the paper based on the user input information, the hierarchical index vector database, the available knowledge base and the experimental data.
[0029] Optionally, the paper generation layer constructs a hierarchical index vector database based on the basic data provided by the metadata layer, and generates a first draft of the paper based on the user input information, the hierarchical index vector database, the available knowledge base and the experimental data, including:
[0030] Constructing a specialized document library based on the user input information and the hierarchical index vector database;
[0031] Generate a paper abstract and keywords based on the specialized document library, and generate a paper outline based on the paper abstract and the keywords;
[0032] Sending the paper abstract, the keywords, and the paper outline to the user, so that the user can modify the paper abstract, the keywords, and the paper outline to obtain a modified paper abstract, modified keywords, and a modified paper outline;
[0033] Creating an online reference library based on the modified paper abstract, the modified keywords, and the modified paper outline;
[0034] Generate English abstracts based on the online reference library;
[0035] Based on a randomized paper paragraph generation strategy, parallel paper generation is performed according to the English abstract, the available knowledge base and the experimental data to obtain a first draft of the paper.
[0036] Optionally, generating a paper outline based on the paper abstract and the keywords includes:
[0037] Randomly searching a plurality of learning documents from the specialized document library according to the paper abstract and the keywords;
[0038] Sorting out the paper logic of the plurality of learning documents to obtain a plurality of initial paper outlines;
[0039] Imitation learning is performed based on the multiple initial paper outlines and domain-based reference materials to obtain a paper outline, where the domain-based reference materials are references provided by the user.
[0040] Optionally, the randomized paper paragraph generation strategy includes:
[0041] When the word count of the first draft of the paper is less than the first word count, the first draft of the paper includes a first-level heading, and a first paragraph generation strategy is executed. The first paragraph generation strategy calls a large language model basic unit once for each first-level heading to generate multiple paragraphs corresponding to each first-level heading;
[0042] When the word count of the first draft of the paper is greater than or equal to the first word count and less than or equal to the second word count, the first draft of the paper includes a first-level title, a second-level title, and a third-level title, and a second paragraph generation strategy is executed, wherein the second paragraph generation strategy calls the large language model basic unit once for each second-level title or each third-level title to generate multiple paragraphs corresponding to each second-level title or each third-level title, and the second word count is greater than the first word count;
[0043] When the number of words in the first draft of the paper is greater than the second number of words, the first draft of the paper includes a first-level title, a second-level title and a third-level title, and a third paragraph generation strategy is executed. The third paragraph generation strategy randomly calls the large language model basic unit multiple times in each second-level title or each third-level title to generate multiple paragraphs of content corresponding to each second-level title or each third-level title.
[0044] Optionally, the large language model basic unit is randomly called multiple times in each second-level title or each third-level title to generate multiple paragraphs of content corresponding to each second-level title or each third-level title, including:
[0045] The content of each title is divided into sections to obtain the center of each paragraph, wherein the title includes a second-level title and a third-level title;
[0046] Calling the large language model basic unit to generate the first paragraph content according to the center of the first paragraph;
[0047] The large language model basic units are called in sequence to generate the next paragraph of content according to the previous paragraph of content and the center of the next paragraph, until all the paragraph contents under the current title are generated.
[0048] Optionally, the large language model basic unit is encapsulated according to the following steps based on the reinforcement learning iterative tuning strategy:
[0049] receiving the user input information, and retrieving relevant data according to the user input information;
[0050] Constructing prompt information based on the retrieved relevant data;
[0051] Calling a large language model according to the prompt information;
[0052] Based on the large language model, a structured output result is constructed.
[0053] Optionally, the method further includes:
[0054] Screening papers according to the modified keywords to obtain screened papers;
[0055] A semantic matching library is constructed based on the content or abstract of the screened paper in sentence units;
[0056] Comparing the sentences in the first draft of the paper with the sentences in the semantic matching library to obtain a similarity comparison result;
[0057] Add literature citations to the first draft of the paper based on the similarity comparison result, the similarity threshold and the dynamic citation probability.
[0058] Optionally, the dynamic citation probability is determined by multiplying the citation probabilities corresponding to the citation threshold constraint, the citation distribution constraint, the multiple citation constraint, and the duplicate citation constraint, wherein:
[0059] The citation threshold constraint is that if the similarity comparison result is greater than the similarity threshold, the citation is performed according to the full probability; if the similarity comparison result is not greater than the similarity threshold, the citation is performed according to the first probability, where the full probability citation represents a 100% citation probability;
[0060] The citation distribution constraint is to cite the first chapter according to the full probability, cite the second chapter according to the second probability, not cite the last chapter, and cite the other chapters according to the third probability, where the third probability is less than the second probability;
[0061] The multi-citation constraint is that when multiple citation documents are retrieved in each citation bracket, the citation documents in the citation bracket are sorted by similarity, and the top N documents with the highest similarity are cited according to full probability, where N is an integer greater than 1;
[0062] The repeated citation constraint is that when the same document is cited multiple times in the same paragraph, it is cited according to the fourth probability and the number of citations of the same document in the same paragraph is recorded. The fourth probability decays exponentially according to the number of citations.
[0063] The embodiments of the present application include the following advantages:
[0064] The scientific research assistance system based on the large language model provided in the embodiment of the present application provides basic data support and available knowledge base support for the system through the metadata layer and the basic computing layer, and provides customized experimental analysis and paper draft writing services for user input information through the experimental simulation layer and the paper generation layer; since the experimental data is constructed and simulated by the experimental simulation layer according to the user input information and the available knowledge base, that is, the experimental data is obtained by customized analysis and design based on the user needs using the available knowledge base, the accuracy and authenticity of the experimental data are guaranteed; since the paper draft is generated by the paper model according to the user input information, the basic data, the experimental data and the available knowledge base, that is, the content of the paper draft is written according to user needs, with the basic data as a reference and based on real and accurate experimental data, rather than being generated in a templated manner, the scientific nature of the content of the paper draft is guaranteed; since the paper draft is directly generated based on the user input information, rather than a small text description, the system can achieve efficient content generation.
[0065] During the process of generating the first draft of the paper, the basic computing layer processes its own input data based on the pre-trained language model, ensuring the system's operating speed. The experimental simulation layer and the paper generation layer each process their own input data based on the large language model. Furthermore, the basic computing layer, the experimental simulation layer, and the paper generation layer iteratively feedback their respective output results based on principle validity metrics and effect evaluation indicators, fine-tuning the models they use based on the output results to further preserve the scientific nature of the first draft of the paper. In this way, based on the powerful information processing and tool-calling capabilities of the large language model, the system, through four functional modules, proceeds from bottom to top. Based on user needs and heuristic strategies, starting from basic data and available knowledge bases, it sequentially conducts the scientific research process of experimental construction, simulated experimental data generation, and first draft of the paper, gradually achieving in-depth discussion and rigorous reasoning of the paper, and obtaining a first draft of the paper that is scientific, authentic, and meets paper standards, providing efficient and accurate scientific research assistance for the scientific research process. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0067] Figure 1 This is a schematic diagram of the structure of a scientific research assistance system based on a large language model provided in an embodiment of the present application;
[0068] Figure 2 This is a schematic diagram of the structure of another large language model-based scientific research assistance system provided in an embodiment of the present application;
[0069] Figure 3 This is a flowchart of a scientific research assistance method based on a large language model provided in an embodiment of the present application;
[0070] Figure 4 A schematic diagram of a theoretical model for writing a first draft of a paper provided in an embodiment of the present application;
[0071] Figure 5 This is a flow chart for generating a first draft of a paper adopted by the embodiment of this application;
[0072] Figure 6 This is a flowchart of the steps of a paragraph generation method provided in an embodiment of the present application;
[0073] Figure 7 This is a method for encapsulating a large language model basic unit provided by an embodiment of the present application;
[0074] Figure 8 This is a flowchart of the steps of a method for adding references provided by the present application. DETAILED DESCRIPTION
[0075] To make the above-mentioned purposes, features, and advantages of this application more clearly understood, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of this application.
[0076] Reference Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of a scientific research assistance system based on a large language model provided in an embodiment of the present application. Figure 1 As shown in the figure, the scientific research assistance system based on the large language model includes:
[0077] The metadata layer 110 includes a corpus and a tool library, which is used to provide basic data for the basic calculation layer and the paper generation layer;
[0078] The basic computing layer 120 is used to regularize and sort the basic data based on the pre-trained language model to obtain a usable knowledge base, which represents the callable knowledge meta-services and tool meta-services;
[0079] The experiment simulation layer 130 is used to construct and simulate the generation of experimental data based on the large language model according to the user input information and the available knowledge base to obtain experimental data;
[0080] The paper generation layer 140 is used to generate a first draft of the paper based on the large language model, the user input information, the basic data, the experimental data and the available knowledge base;
[0081] In the process of generating the first draft of the paper, the basic computing layer, the experimental simulation layer and the paper generation layer iteratively feedback their respective output results based on the principle validity measurement index and the effect evaluation index, so as to fine-tune the model used by themselves according to the output results.
[0082] In the embodiment of the present application, the tool library in the metadata layer can assist researchers in providing reliable means for reproducing old experiments and designing new experiments. The tool library includes available APIs (Application Programming Interfaces), domain inductive biases, noun knowledge graphs, and other prior knowledge. In order to ensure the accuracy and stability of the content retrieved by the large language model of the basic computing layer, these heterogeneous and multi-level data in the tool library are also cleaned and vectorized, and ultimately the basic data with unified expression will be provided to the basic computing layer.
[0083] The basic computing layer is a service foundation based on the metadata layer. It organizes and organizes the basic data provided by the metadata layer to obtain a usable knowledge base, which represents the callable knowledge meta-services and tool meta-services. In other words, the usable knowledge base is a meta-service that can be better called by upper-level modules (such as the experimental simulation layer and the paper generation layer) compared to the basic data.
[0084] Specifically, the available knowledge base includes specialized knowledge graph construction, embedded representations of scientific data, a literature similarity search engine, callable tool retrieval codes, a knowledge linking system, and pre-defined logical chains for large models. These available knowledge bases serve as foundational units for experimental design in the upper-level experimental simulation layer and for paper writing in the paper generation layer. To ensure system speed, the basic computing layer and embedding-related services are primarily implemented by pretrained language models (PLMs).
[0085] The experimental simulation layer serves as the decisive basis for the generation of the first draft of the paper. It mainly implements the design and execution of investigative experiments, theoretical experiments and simulation experiments based on the basic experimental information input by the user, combined with the meta-services (i.e., available knowledge base) provided by the basic computing layer through a large language model. Specifically, the experimental simulation layer includes: an investigative experiment module, including a target data acquisition module and a data visualization modeling module, for generating basic experimental data; a theoretical experiment module, including an inductive bias introduction module and a basic theoretical reasoning module, for generating theoretical experimental data; a simulation experiment module, including a virtual environment construction, a meta-code module construction and a code logic chain combination module, for generating simulation experimental data. In this way, through the various functional modules in the experimental simulation layer, basic experimental data, theoretical experimental data and simulation experimental data can be provided to support the writing of the upper paper generation layer; because the experimental data is obtained through customized analysis and design based on user needs using the available knowledge base, the accuracy and authenticity of the experimental data are guaranteed.
[0086] The paper generation layer is the final output of the system for users. It generates a first draft of the paper based on user input information, basic data, experimental data and available knowledge base. The first draft of the paper can be an experimental result analysis report or an academic paper. Specifically, the paper generation layer includes: a previous work analysis module for analyzing the basic data; a paper logic combing module for combing out the logical structure of the first draft of the paper based on the user input; a paper formula deduction module for deriving the formulas used in the first draft of the paper; a paper illustration generation module for generating illustrations for the first draft of the paper; an experimental data regularization module for normalizing the experimental data; an experimental data analysis module for analyzing the experimental data; and a reference regularization module for inserting references into the first draft of the paper. In this way, based on the various functional modules in the paper generation layer, a compliant scientific research paper (i.e., a first draft of the paper) that is innovative, academic, feasible, and credible can be stably written according to the user input information to provide certain reference and support in the subsequent scientific research process. Since the content of the first draft of the paper is written according to user needs, with basic data as a reference and real and accurate experimental data as the basis, the scientific nature of the content of the first draft of the paper is guaranteed.
[0087] In summary, the scientific research assistance system based on the large language model provided by the embodiment of the present application provides basic data support and available knowledge base support for the system through the metadata layer and the basic computing layer, and provides customized experimental analysis and paper draft writing services for user input information through the experimental simulation layer and the paper generation layer; since the experimental data is constructed and simulated by the experimental simulation layer according to the user input information and the available knowledge base, the experimental data is obtained by customized analysis and design based on the user needs using the available knowledge base, thereby ensuring the accuracy of the experimental data; since the paper draft is generated by the paper model according to the user input information, the basic data, the experimental data and the available knowledge base, the content of the paper draft is written based on user needs, with the basic data as a reference and real and accurate experimental data as the basis, rather than being generated in a templated manner, thereby ensuring the scientific nature of the content of the paper draft; since the paper draft is directly generated based on the user input information, rather than a short text description, the system can achieve efficient content generation.
[0088] During the process of generating the first draft of the paper, the basic computing layer processes its own input data based on the pre-trained language model, ensuring the system's operating speed. The experimental simulation layer and the paper generation layer each process their own input data based on the large language model. Furthermore, the basic computing layer, the experimental simulation layer, and the paper generation layer iteratively feedback their respective output results based on principle validity metrics and effect evaluation indicators, fine-tuning the models they use based on the output results to further preserve the scientific nature of the first draft of the paper. In this way, based on the powerful information processing and tool-calling capabilities of the large language model, the system, through four functional modules, proceeds from bottom to top. Based on user needs and heuristic strategies, starting from basic data and available knowledge bases, it sequentially conducts the scientific research process of experimental construction, simulated experimental data generation, and first draft of the paper, gradually achieving in-depth discussion and rigorous reasoning of the paper, and obtaining a first draft of the paper that is scientific, authentic, and meets paper standards, providing efficient and accurate scientific research assistance for the scientific research process.
[0089] For example, Figure 2 It is a structural diagram of another scientific research assistance system based on a large language model provided in an embodiment of the present application. Specifically, the scientific research assistance system of the large language model includes a metadata layer, a basic computing layer, an experimental simulation layer, and a paper generation layer. Among them, the metadata layer includes a corpus and a tool library. The corpus includes: real-time news information, a public academic paper library, a proprietary academic document library, and other available Internet data; in order to ensure the real-time nature of the corpus, various data are connected to the metadata layer through network probes. The tool library can assist scientific researchers in providing reliable means for reproducing old experiments and designing new experiments. The tool library includes available APIs, domain induction biases, nominal knowledge graphs, and other prior knowledge.
[0090] The basic computing layer provides a usable knowledge base based on the metadata layer's service foundation. This knowledge base includes: specialized knowledge graph construction, embedded representation of scientific data, a literature material similarity search engine, callable tool retrieval encoding, a knowledge linking system, and pre-set logical chains for large models. The experimental simulation layer includes: an investigative experimental module (including a target data acquisition module and a data visualization modeling module), a theoretical experimental module (including an inductive bias introduction module and a basic theoretical reasoning module), and a simulation experimental module (including a virtual environment construction, metacode module construction, and a code logic chain combination module). The various functional modules of the experimental simulation layer provide experimental data for paper generation.
[0091] The paper generation layer includes: a prior work analysis module, a paper logic analysis module, a paper formula deduction module, a paper illustration generation module, an experimental data normalization module, an experimental data analysis module, and a reference regularization module. Through these functional modules, you can write a compliant scientific research paper that is innovative, scholarly, feasible, and credible.
[0092] Reference Figure 3 As shown, Figure 3 This is a flow chart of the steps of a scientific research assistance method based on a large language model provided in an embodiment of the present application. This method is applied to the above-mentioned scientific research assistance system based on a large language model. Figure 3 As shown, the method includes steps S310 to S330:
[0093] Step S310: Obtain user input information, where the user input information at least includes the paper title and the paper field.
[0094] Step S320: the experiment simulation layer constructs and generates simulated experiment data according to the user input information and the available knowledge base provided by the basic computing layer to obtain experimental data.
[0095] Step S330: The paper generation layer constructs a hierarchical index vector database based on the basic data provided by the metadata layer, and generates a first draft of the paper based on the user input information, the hierarchical index vector database, the available knowledge base and the experimental data.
[0096] In this embodiment of the application, the metadata layer and basic computing layer of the research support system based on the large language model are independent of the user's input information and provide basic data support and basic meta-service support for the system. After the system obtains the user's input information, it mainly calls the experimental simulation layer and the paper generation layer to perform experimental analysis and generate the first draft of the paper.
[0097] When writing the first draft of a paper, step S310 is first executed to obtain user input information, where the user input information can be entered into the system through an input text editor. Then, step S320 is executed. The experimental simulation layer customizes experimental data based on the user input information and the available knowledge base provided by the basic computing layer to construct and simulate experiments that meet the user input information and obtain experimental results. Then, step S330 is executed. The paper generation layer uses the basic data provided by the metadata layer as a reference to construct a hierarchical index vector database, thereby combining the experimental data, user input information, and available knowledge base to write the first draft of the paper.
[0098] In some embodiments, the user input information also includes: an abstract (i.e., a regularized abstract) and core references, so as to provide a more scientific reference for the generation of a draft paper by a large language model through the abstract and core references. For example, Figure 4 A schematic diagram of a theoretical model for writing a first draft of a paper provided in an embodiment of the present application, wherein the paper generation layer writes the first draft of the paper according to the theoretical model for writing a first draft of the paper. Specifically, the paper generation layer uses the basic data provided by the metadata layer as a reference, edits the metadata according to the document editor, and constructs a hierarchical index vector database; wherein the basic data may include HowNet documents, real-time news, and international public document libraries, which are accessed to the metadata layer through HowNet protocols or real-time network probes. Afterwards, the information retriever is used to retrieve domain-specific knowledge from the hierarchical index vector database based on the user input information, thereby obtaining a first draft generation prompt based on the academic style prompt design, the user-input abstract, and the core references. Finally, based on the first draft generation prompt, the large language model is called to write the paper to obtain a first draft of the paper.
[0099] Based on the above embodiment process, since the experimental data is constructed and simulated by the experimental simulation layer according to the user input information and the available knowledge base, that is, the experimental data is obtained by customized analysis and design based on the available knowledge base according to the user needs, the accuracy of the experimental data is guaranteed; since the first draft of the paper is generated by the paper model according to the user input information, the basic data, the experimental data and the available knowledge base, that is, the content of the first draft of the paper is written based on the user needs, with the basic data as a reference and real and accurate experimental data as the basis, rather than generated in a templated manner, the scientific nature of the content of the first draft of the paper is guaranteed; since the first draft of the paper is directly generated based on the user input information, rather than a short text description, the system can achieve efficient content generation. In this way, according to user needs and heuristic strategies, starting from the basic data and the available knowledge base, the scientific research process of experimental construction, simulation experimental data generation, and first draft of the paper is carried out in sequence, gradually achieving in-depth discussion and rigorous reasoning of the paper, and obtaining a first draft of the paper that is scientific and authentic and meets the paper specifications. This method provides an efficient and accurate auxiliary method for literature research, innovation discovery, experimental simulation, data analysis, and draft writing in the scientific research process.
[0100] In an optional embodiment, the paper generation layer constructs a hierarchical index vector database based on the basic data provided by the metadata layer, and generates a draft of the paper based on the user input information, the hierarchical index vector database, the available knowledge base and the experimental data, including steps A1 to A6:
[0101] Step A1: constructing a specialized document library based on the user input information and the hierarchical index vector database;
[0102] Step A2: generating a paper abstract and keywords based on the specialized literature library, and generating a paper outline based on the paper abstract and the keywords;
[0103] Step A3: sending the paper abstract, the keywords, and the paper outline to the user, so that the user can modify the paper abstract, the keywords, and the paper outline to obtain a modified paper abstract, modified keywords, and modified paper outline;
[0104] Step A4: creating an online reference library based on the modified paper abstract, the modified keywords, and the modified paper outline;
[0105] Step A5: Generate an English abstract based on the online reference library;
[0106] Step A6: Based on the randomized paper paragraph generation strategy, parallel paper generation is performed according to the English abstract, the available knowledge base and the experimental data to obtain a first draft of the paper.
[0107] In the embodiment of the present application, the proprietary document library constructed according to user input information and hierarchical index vector database is a coarse-grained reference generation database, which quickly obtains paper abstracts, keywords and paper outlines based on the coarse-grained reference generation database. In order to make the final draft of the paper more scientific, the paper abstract, keywords and paper outline are fed back to the user, and the user modifies the paper abstract, keywords and paper outline to obtain modified paper abstracts, modified keywords and modified paper outlines that are more in line with user needs. Afterwards, a fine-grained online reference library (the online reference library is more relevant to the user's research field) is constructed based on the modified paper abstract, modified keywords and modified paper outline to generate an English abstract. Ultimately, parallelized paper generation is performed according to the English abstract, available knowledge base and the experimental data to obtain a paper draft; wherein, parallelized paper generation refers to the ability to generate the content of multiple chapters at the same time, so as to achieve efficient paper draft generation.
[0108] In some embodiments, before step A1, the local document database is initialized, that is, the basic data provided by the metadata layer is loaded into the paper generation module, and after the local document database is initialized, the large language model is initialized. After step A6, the generated paper draft is written to a local Word document, and the log record is updated, and finally the cache in the system is cleared. For example, Figure 5 This is a flow chart for generating a first draft of a paper adopted in an embodiment of the present application.
[0109] In this way, by establishing two reference generation databases of different sizes, coarse-grained and fine-grained, the first draft of the paper is generated in parallel with short delay based on experimental data and available knowledge base, and the first draft of the paper is generated through diversified and accurate collaborative calls of multiple databases.
[0110] In an optional embodiment, a paper outline is generated based on the paper abstract and the keywords, including: randomly retrieving multiple learning documents from the specialized document library based on the paper abstract and the keywords; sorting out the paper logic of the multiple learning documents to obtain multiple initial paper outlines; performing imitation learning based on the multiple initial paper outlines and domain-specific reference materials to obtain a paper outline, wherein the domain-specific reference materials are reference documents provided by the user.
[0111] According to the embodiment of the present application, multiple learning documents randomly retrieved from the specialized document library based on the paper abstracts and keywords are documents related to user research. Therefore, based on the retrieved multiple learning documents, valuable references can be provided for writing the first draft of the paper, so as to obtain a more scientific paper outline based on the multiple learning documents to provide support for the subsequent writing of the first draft of the paper.
[0112] In an optional embodiment, the randomized paper paragraph generation strategy includes steps B1 to B3:
[0113] Step B1: When the word count of the first draft of the paper is less than the first word count, the first draft of the paper includes a first-level heading, and a first paragraph generation strategy is executed. The first paragraph generation strategy calls a large language model basic unit once for each first-level heading to generate multiple paragraphs corresponding to each first-level heading;
[0114] Step B2: When the word count of the first draft of the paper is greater than or equal to the first word count and less than or equal to the second word count, the first draft of the paper includes a first-level title, a second-level title, and a third-level title, executing a second paragraph generation strategy, wherein the second paragraph generation strategy calls the large language model basic unit once for each second-level title or each third-level title to generate multiple paragraphs corresponding to each second-level title or each third-level title, and the second word count is greater than the first word count;
[0115] Step B3: When the number of words in the first draft of the paper is greater than the second number of words, the first draft of the paper includes a first-level title, a second-level title, and a third-level title, and a third paragraph generation strategy is executed. The third paragraph generation strategy randomly calls the large language model basic unit multiple times in each second-level title or each third-level title to generate multiple paragraphs of content corresponding to each second-level title or each third-level title.
[0116] In the embodiment of the present application, to ensure that the content of different chapters and paragraphs in the first draft of the paper is more diverse, paragraph content is generated based on a randomized paper paragraph generation strategy. Specifically, a paper with a word count less than the first word count (for example, the first word count is 10,000 words) usually includes multiple first-level headings. In this case, the first paragraph generation strategy is executed to obtain multiple paragraphs corresponding to each first-level heading. Finally, the multiple paragraphs corresponding to each first-level heading are spliced together to obtain the first draft of the paper.
[0117] A paper with a word count greater than or equal to the first word count (for example, the first word count is 10,000 words) and less than or equal to the second word count (for example, the second word count is 30,000 words) usually includes a first-level title, a second-level title, and a third-level title. At this time, the second paragraph generation strategy is executed to obtain multiple paragraphs corresponding to each second-level title or each third-level title. Finally, the multiple paragraphs corresponding to each second-level title or each third-level title are spliced together to obtain the first draft of the paper. In addition, to prevent content confusion between multiple levels of titles, when generating the content of each level of title, all second-level titles and third-level titles contained in the first-level title are input into the basic unit of the large language model. Since multiple titles are loosely coupled individuals, this method can effectively avoid the duplication of text content between generated multiple levels of titles.
[0118] A paper with a first draft exceeding the second word count (for example, a second word count of 30,000 words) typically includes first-level, second-level, and third-level headings. In this case, the third paragraph generation strategy is executed to generate multiple paragraphs corresponding to each second-level or third-level heading. Finally, these multiple paragraphs corresponding to each second-level or third-level heading are concatenated to produce the first draft. Because multiple headings are loosely coupled and multiple paragraphs are tightly coupled, the multiple calls to the large language model's basic units ensure that the generated paragraphs are closely linked and do not repeat. Each paragraph is generated by also taking the previous paragraph into consideration.
[0119] In an alternative embodiment, referring to Figure 6 As shown, Figure 6 This is a flowchart of a paragraph generation method provided by an embodiment of the present application, which randomly calls the large language model basic unit multiple times in each second-level title or each third-level title to generate multiple paragraphs corresponding to each second-level title or each third-level title, including steps C1 to C3:
[0120] Step C1: segment the content of each title to obtain the center of each paragraph, wherein the title includes a secondary title and a tertiary title;
[0121] Step C2: calling the large language model basic unit to generate the first paragraph content according to the center of the first paragraph;
[0122] Step C3: Call the basic units of the large language model in sequence, and generate the next paragraph of content according to the previous paragraph of content and the center of the next paragraph, until all the paragraph contents under the current title are generated.
[0123] In the embodiment of the present application, the content of each title is segmented and planned to obtain the center of each paragraph, so that different paragraphs are related and the content is not repeated. Afterwards, for the first paragraph of each title, the corresponding paragraph content is directly obtained based on the center of the first paragraph. For each subsequent paragraph, not only the center of each paragraph is considered, but also the corresponding paragraph content is generated based on the previous paragraph content, so that the content of adjacent paragraphs has a better connection relationship.
[0124] In this way, in the embodiment of the present application, based on the randomized paper paragraph generation strategy, the final draft of the paper obtained by parallel paper generation based on the English abstract, available knowledge base and experimental data can better meet the needs of users, and the paragraph content of the draft of the paper is richer. It is not generated in a templated manner, which ensures the scientific nature of the content of the draft of the paper.
[0125] In an alternative embodiment, referring to Figure 7As shown, the large language model basic unit is encapsulated based on a reinforcement learning iterative tuning strategy by following these steps: receiving user input and retrieving relevant data based on the user input; constructing prompt information based on the retrieved relevant data; invoking the large language model based on the prompt information; and constructing a structured output result based on the large language model. Thus, by encapsulating the large language model as a large language model basic unit, paragraph content can be generated by directly invoking the large language model basic unit during paragraph generation.
[0126] In an optional embodiment, a method for adding references is also provided. Figure 8 As shown, Figure 8 This is a flowchart of a method for adding references provided by the present application. The method for adding references includes the following steps D1 to D4:
[0127] Step D1: Screening papers based on the modified keywords to obtain screened papers;
[0128] Step D2: constructing a semantic matching library based on sentences of the content or abstract of the screened papers;
[0129] Step D3: performing a similarity comparison between the sentences in the first draft of the paper and the sentences in the semantic matching library to obtain a similarity comparison result;
[0130] Step D4: adding literature citations to the first draft of the paper based on the similarity comparison result, the similarity threshold and the dynamic citation probability.
[0131] In the embodiment of the present application, a dual screening model is performed based on the modified keywords and the dynamic threshold of sentence semantic similarity based on user feedback, and a semantic matching library is established based on the sentences in the paper content or abstract. The greater the similarity comparison result between the sentences in the first draft of the paper and the sentences in the semantic matching library, the more similar they are. Therefore, the corresponding document can be considered for addition as a reference. By setting a similarity threshold to compare with the similarity comparison result, the similarity comparison result can be used as the basis for whether to add the reference.
[0132] Moreover, considering that the citation situation of the content in different positions in the first draft is different, for example, in the first chapter about the background technology, it may be necessary to cite a large number of references, but in the last chapter about the conclusion of the paper, it may not be necessary to cite a large number of references, so the citation situation of the references in different positions of the paper is different; in addition, the citation of references is also related to the number of citations to the same document. Therefore, the embodiment of the present application proposes to use the dynamic citation probability as the basis for whether to add references. In this way, by introducing the similarity threshold and the dynamic citation probability, the addition of references in the first draft of the paper is comprehensively considered, so that the reference citations of the final first draft of the paper are more scientific, avoiding the problems of irregular citation format, serious inconsistency between the citation content and the article, rigid citation distribution, and inaccurate data citation.
[0133] Specifically, the dynamic citation probability is determined by multiplying the citation probabilities corresponding to the citation threshold constraint, the citation distribution constraint, the multiple citation constraint, and the duplicate citation constraint, where:
[0134] Item E1: The citation threshold constraint: if the similarity comparison result is greater than the similarity threshold, the citation is performed according to the full probability; if the similarity comparison result is not greater than the similarity threshold, the citation is performed according to the first probability. The full probability citation indicates that the citation probability is 100%;
[0135] Item E2: The citation distribution constraint is to cite the first chapter according to the full probability, cite the second chapter according to the second probability, not cite the last chapter, and cite the other chapters according to the third probability, where the third probability is less than the second probability.
[0136] Item E3: The multiple citation constraint, when multiple citations are retrieved within each citation bracket, the citations within the citation bracket are sorted by similarity, and the top N citations with the highest similarity are cited according to the full probability, where N is an integer greater than 1;
[0137] Item E4: The repeated citation constraint, when the same document is cited multiple times in the same paragraph, is cited according to the fourth probability, and the number of citations of the same document in the same paragraph is recorded. The fourth probability decays exponentially according to the number of citations.
[0138] In the embodiment of the present application, for each cited reference, the four citation probabilities corresponding to the constraints of items E1 to E4 of the reference are determined in sequence, and then the dynamic citation probability is obtained according to the product of the four citation probabilities.
[0139] For the E1 multi-reference constraint, if the similarity comparison result is greater than the similarity threshold, the corresponding citation probability of the multi-reference constraint is 1. If the similarity comparison result is not greater than the similarity threshold, the corresponding citation probability of the multi-reference constraint is the first probability. The first probability is a value less than 1 and can be flexibly set according to actual conditions.
[0140] For the E2 citation distribution constraint, different citation probabilities are set based on the different positions of sentences in the paper. For the first chapter, since the content of the first chapter is usually related to technical background and often cites a large number of literature for discussion, the citation probability is set to 1. For the second chapter, the citation probability is set to the second probability (where the second probability is a value less than 1 and can be flexibly set based on actual conditions). For the last chapter, since the last chapter is usually the conclusion of the paper, the citation probability is set to 0 (i.e., no citation is performed). For the other chapters, sparse citation search is performed based on the third probability, that is, the citation probability is set to the third probability.
[0141] Regarding the E3 multi-reference constraint, when a sentence is similar to multiple sentences in the semantic matching library, it may need to cite multiple references. Set the maximum number of references (i.e., N) and use the top N most similar references as the references to be cited. Set the citation probability for these top N references to 1. Set the citation probability for the remaining references to 0, meaning they will not be cited.
[0142] For the repeated citation constraint E4, if the same paragraph cites the same reference multiple times, the corresponding citation probability (i.e., the fourth probability) is determined based on the number of citations. For example, if it is the second citation, the value of the fourth probability is smaller than the value of the first citation.
[0143] In this way, the dynamic citation probability of each reference is determined based on multiple constraints. This dynamic citation probability is then used to determine whether to add the corresponding reference citation to the first draft of the paper. This ensures the rationality and scientific nature of the reference citations in the first draft, better conforming to the standards of scientific papers, and further ensuring the authenticity and scientific nature of the first draft. Therefore, this method provides an efficient and accurate auxiliary method for literature research, innovation discovery, experimental simulation, data analysis, and first draft writing in the scientific research process.
[0144] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0145] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the systems and methods according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0146] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0148] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0149] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0150] The above is a detailed introduction to a scientific research assistance system and method based on a large language model provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A scientific research assistance system based on a large language model, characterized in that: include: The metadata layer, including the corpus and tool library, is used to provide basic data for the basic calculation layer and paper generation layer; The basic computing layer is used to regularize and sort the basic data based on the pre-trained language model to obtain a usable knowledge base, which represents the callable knowledge meta-services and tool meta-services; An experiment simulation layer, configured to construct and simulate experimental data generation based on a large language model according to user input information and the available knowledge base to obtain experimental data; A paper generation layer, configured to generate a first draft of the paper based on a large language model, the user input information, the basic data, the experimental data, and the available knowledge base; During the process of generating the first draft of the paper, the basic computing layer, the experimental simulation layer, and the paper generation layer provide iterative feedback on their respective output results based on the principle validity measurement index and the effect evaluation index, so as to fine-tune the models used by themselves according to the output results; The paper generation layer is specifically used to execute steps A1 to A6: Step A1: constructing a specialized document library based on the user input information and the hierarchical index vector database. The specialized document library is a coarse-grained reference generation database, so as to quickly obtain paper abstracts, keywords, and paper outlines based on the coarse-grained reference generation database; Step A2: Generate a paper abstract and keywords based on the specialized document library, and generate a paper outline based on the paper abstract and the keywords, including: randomly searching multiple learning documents from the specialized document library based on the paper abstract and the keywords, wherein the multiple learning documents are documents related to user research; sorting out the paper logic of the multiple learning documents to obtain multiple initial paper outlines; and performing imitation learning based on the multiple initial paper outlines and domain-specific reference materials to obtain a paper outline, wherein the domain-specific reference materials are reference materials provided by the user; Step A3: sending the paper abstract, the keywords, and the paper outline to the user, so that the user can modify the paper abstract, the keywords, and the paper outline to obtain a modified paper abstract, modified keywords, and modified paper outline that better meets the user's needs; Step A4: creating a fine-grained online reference library based on the modified paper abstract, the modified keywords, and the modified paper outline, where the online reference library is more relevant to the user's research field; Step A5: Generate an English abstract based on the online reference library; Step A6: Based on a randomized paper paragraph generation strategy, parallel paper generation is performed according to the English abstract, the available knowledge base, and the experimental data to obtain a first draft of the paper. The randomized paper paragraph generation strategy is used to ensure that the content of different chapters and paragraphs in the first draft of the paper is more diverse; The step A6 includes: when the number of words in the first draft of the paper is greater than the second number of words, the first draft of the paper includes a first-level title, a second-level title, and a third-level title, executing a third paragraph generation strategy, wherein the third paragraph generation strategy randomly calls the large language model basic unit multiple times for each second-level title or each third-level title to generate multiple paragraphs corresponding to each second-level title or each third-level title, specifically including steps C1 to C3: Step C1: Segment-planning the content of each title to obtain the center of each paragraph. The titles include secondary and tertiary titles. The center of each paragraph is obtained by segment-planning the content of each title to ensure that different paragraphs are related and the content is not repeated. Step C2: calling the large language model basic unit to generate the first paragraph content based on the center of the first paragraph, wherein the first paragraph content in each title is directly based on the center of the first paragraph to obtain the corresponding paragraph content; Step C3: Call the basic units of the large language model in sequence, generate the next paragraph of content based on the center of the previous paragraph and the next paragraph, until all the paragraph contents under the current title are generated; among them, for each subsequent paragraph of content of the first paragraph in each title, not only the center of the paragraph itself should be considered, but also the corresponding paragraph content should be generated based on the previous paragraph, so that the adjacent paragraph contents have a better connection relationship.
2. The scientific research assistance system based on a large language model according to claim 1 is characterized in that: The available knowledge base includes: specialized knowledge graph construction, embedded representation of scientific data, literature material similarity retriever, callable tool retrieval coding, knowledge linking system, and logical chain preset by large model.
3. The scientific research assistance system based on a large language model according to claim 1 is characterized in that: The experimental simulation layer includes: Investigative experiment module, including target data acquisition module and data visualization modeling module, used to generate basic experimental data; Theoretical experiment module, including the inductive bias introduction module and the basic theoretical reasoning module, is used to generate theoretical experimental data; The simulation experiment module includes virtual environment construction, meta-code module construction and code logic chain combination module, which is used to generate simulation experiment data.
4. The scientific research assistance system based on a large language model according to claim 1, characterized in that: The paper generation layer includes: A previous work analysis module, used for analyzing the basic data; The paper logic sorting module is used to sort out the logical structure of the paper draft according to the user input; Thesis formula deduction module, used to derive the formulas used in the first draft of the thesis; A paper illustration generation module, used to generate illustrations for the first draft of the paper; An experimental data normalization module, used for normalizing the experimental data; An experimental data analysis module, used for analyzing the experimental data; Reference regularization module, used to insert references in the first draft of the paper.
5. A scientific research assistance method based on a large language model, characterized in that: Applied to the large language model-based scientific research assistance system according to any one of claims 1 to 4, the method comprising: Obtaining user input information, wherein the user input information includes at least a paper title and a paper field; The experiment simulation layer constructs and simulates experimental data based on the user input information and the available knowledge base provided by the basic computing layer to obtain experimental data; The paper generation layer constructs a hierarchical index vector database based on the basic data provided by the metadata layer, and generates a first draft of the paper based on the user input information, the hierarchical index vector database, the available knowledge base and the experimental data.
6. The scientific research assistance method based on a large language model according to claim 5, characterized in that: The method further comprises: When the word count of the first draft of the paper is less than the first word count, the first draft of the paper includes a first-level heading, and a first paragraph generation strategy is executed. The first paragraph generation strategy calls a large language model basic unit once for each first-level heading to generate multiple paragraphs corresponding to each first-level heading; When the number of words in the first draft of the paper is greater than or equal to the first number of words and less than or equal to the second number of words, the first draft of the paper includes a first-level title, a second-level title, and a third-level title, and a second paragraph generation strategy is executed. The second paragraph generation strategy calls the large language model basic unit once for each second-level title or each third-level title to generate multiple paragraphs corresponding to each second-level title or each third-level title, and the second number of words is greater than the first number of words.
7. The scientific research assistance method based on a large language model according to claim 5, characterized in that: The basic unit of the large language model is encapsulated according to the following steps based on the reinforcement learning iterative tuning strategy: receiving the user input information, and retrieving relevant data according to the user input information; Constructing prompt information based on the retrieved relevant data; Calling a large language model according to the prompt information; Based on the large language model, a structured output result is constructed.
8. The scientific research assistance method based on a large language model according to claim 7, characterized in that: The method further comprises: Screen the papers based on the modified keywords to obtain the screened papers; A semantic matching library is constructed based on the content or abstract of the screened paper in sentence units; Comparing the sentences in the first draft of the paper with the sentences in the semantic matching library to obtain a similarity comparison result; Add literature citations to the first draft of the paper based on the similarity comparison result, the similarity threshold and the dynamic citation probability.
9. The scientific research assistance method based on a large language model according to claim 8, characterized in that: The dynamic citation probability is determined by multiplying the citation probabilities corresponding to the citation threshold constraint, the citation distribution constraint, the multiple citation constraint, and the duplicate citation constraint, where: The citation threshold constraint is that if the similarity comparison result is greater than the similarity threshold, the citation is performed according to the full probability; if the similarity comparison result is not greater than the similarity threshold, the citation is performed according to the first probability, where the full probability citation represents a 100% citation probability; The citation distribution constraint is to cite the first chapter according to the full probability, cite the second chapter according to the second probability, not cite the last chapter, and cite the other chapters according to the third probability, where the third probability is less than the second probability; The multi-citation constraint is that when multiple citation documents are retrieved in each citation bracket, the citation documents in the citation bracket are sorted by similarity, and the top N documents with the highest similarity are cited according to full probability, where N is an integer greater than 1; The repeated citation constraint is that when the same document is cited multiple times in the same paragraph, it is cited according to the fourth probability and the number of citations of the same document in the same paragraph is recorded. The fourth probability decays exponentially according to the number of citations.
Citation Information
Patent Citations
AI intelligent paper generation method
CN110598197A
Meta analysis generation method based on artificial intelligence
CN111552776A
System for college physics experiment teaching
CN112614033A
Literature review generation method based on large language model
CN117709306A