Automatic scientific research and review method and system
By using the researcher model to generate initial papers during the scientific research process and using the reward model for review and optimization, the problem that existing AI tools are difficult to automate the entire scientific research process is solved, and the quality and efficiency of scientific research are improved.
Patent Information
- Application Number
- CN202510233091.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI-driven research tools are difficult to achieve efficient automation of the entire scientific research process, and it is difficult to meet the high standards of peer review in terms of scientificity, innovation, reliability and presentation effects.
Provide an automated scientific research and review method, which generates initial research papers based on the researcher model and uses the reward model to review and iteratively optimize the papers until the preset standards are met.
It has realized the automation and intelligence of the entire life cycle of scientific research, improved the overall quality and efficiency of scientific research work, and ensured the scientificity and innovation of research papers.
Smart Images

Figure CN120068815A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and particularly to an automated scientific research and review method and system. Background Art
[0002] With the rapid development of artificial intelligence technology, the exploration of automation in the scientific research field has been continuously deepened. Since the 1970s and 1980s of the last century, computer science has emerged, and the idea of scientific research automation has sprouted. Researchers expect to achieve scientific research automation with the help of artificial intelligence technology. In recent years, large language models (LLMs) have brought new opportunities.
[0003] Currently, most research uses commercial LLMs to build agents to assist specific links in scientific research, such as generating ideas, assisting experiments, generating publications, etc., but fails to achieve seamless connection and efficient automation of the entire scientific research process. Existing AI-driven research tools are difficult to meet the high standards required by peer review in key aspects such as scientificity, innovation, reliability, and presentation effects. At the same time, these tools have weak integration and iterative feedback capabilities, cannot adapt to the complex requirements of each stage of scientific research, and have poor cross-stage collaborative work capabilities. In short, although the existing technology has achieved certain phased results on the road of scientific research automation, there are still many key problems to be solved in realizing an efficient, reliable, and comprehensive scientific research automation process. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an automated scientific research and review method and system, which can realize the automation and intelligence of the entire life cycle of scientific research and improve the overall quality and efficiency of scientific research work.
[0005] To achieve the above purpose, the embodiments of this application provide an automated scientific research and review method, and the method includes:
[0006] Generating an initial research paper using a researcher model based on the obtained existing research literature;
[0007] Reviewing the initial research paper using a reward model to obtain a review comment, where the review comment includes a review score and a review opinion;
[0008] The researcher model iteratively optimizes the quality of the initial research paper based on the review comment until a research paper that meets the preset criteria is generated.
[0009] Optionally, the generating an initial research paper using a researcher model based on the obtained existing research literature includes:
[0010] Input the existing research literature into the trained researcher model. The researcher model identifies research questions, derives solutions, designs experiments, and writes papers based on the existing research literature, and finally outputs an initial research paper. Among them, the researcher model is obtained by performing reinforcement learning training on the policy model using a preference pair dataset. The policy model is obtained by performing supervised learning training on the first large language model using a research paper dataset.
[0011] Optionally, performing supervised learning training on the first large language model using a research paper dataset to obtain a policy model includes:
[0012] Construct a research paper dataset. Among them, the research paper dataset includes research training samples, and the research training samples include research input features and corresponding research output labels. Among them, the research input features are the references of the existing papers obtained, and the research output labels include the paper outline and the main text. The main text includes research motivation, research methods, experimental settings, and research results.
[0013] Determine the first loss function and the first optimizer. Among them, the first loss function is used to measure the difference between the model prediction result and the research output label, and the first optimizer is used to adjust the model parameters to minimize the loss function.
[0014] Input the research input features into the first large language model to obtain a first prediction result.
[0015] Calculate the loss value between the first prediction result and the research output label, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and the first optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training rounds, completing the supervised learning training of the first large language model.
[0016] Optionally, the construction of the research paper dataset includes the process of constructing research input features and the process of constructing research output labels.
[0017] Among them, the process of constructing research input features is:
[0018] Use an academic search engine to obtain existing papers.
[0019] Use the application programming interface provided by the academic search engine to retrieve the references cited in the obtained existing papers.
[0020] Construct research input features using the retrieved references.
[0021] The process of constructing research output labels is:
[0022] Preprocess the obtained existing papers to obtain the main text.
[0023] Input the main text into an existing outline generation model to obtain the paper outline;
[0024] Construct research output labels using the paper outline and the main text.
[0025] Optionally, perform reinforcement learning training on the policy model using a preference pair dataset to obtain a researcher model, including:
[0026] Construct a preference pair dataset; wherein, the preference pair dataset includes reference documents, positive samples, and negative samples, where the positive samples are research papers with relatively high review scores generated by the policy model, and the negative samples are research papers with relatively low review scores generated by the policy model;
[0027] Input the reference documents in the preference pair dataset into the initialized policy model to generate research papers;
[0028] Input the generated research papers into a reward model to obtain review scores and feedback them to the policy model, enabling the policy model to adjust its own parameters according to the review scores to optimize the quality of the research papers generated next; wherein, the reward model is obtained by performing supervised learning training on a second large language model using a review dataset;
[0029] Input the reference documents in the preference pair dataset that have not participated in the training of the policy model into the policy model with adjusted own parameters until all reference documents have participated in the training of the policy model;
[0030] Use the policy model that generates research papers with relatively high review scores as the finally completed researcher model after reinforcement learning training.
[0031] Optionally, the construction of the preference pair dataset includes:
[0032] Input the reference documents of existing papers into the policy model to generate research papers;
[0033] Input the generated research papers into a reward model to obtain review scores;
[0034] Use the research papers with relatively high review scores as positive samples and the research papers with relatively low review scores as negative samples;
[0035] Construct a preference pair dataset using reference documents, positive samples, and negative samples.
[0036] Optionally, during the process of performing reinforcement learning training on the policy model, use a policy loss function to perform iterative training on the policy model;
[0037] The policy loss function is:
[0038]
[0039] In the formula, represents the policy loss function; λ represents a hyperparameter used to balance the SimPO loss and the negative log-likelihood loss; x represents the input of the policy model, and y w represents the positive sample, and y l represents the negative sample, represents the set of preference samples, and π θ represents the policy model, and π θ (y w |x) represents the probability that the policy model π θ outputs y w ; π θ (y l |x) represents the probability that the policy model π θ outputs y l ; π θ (y w |x) represents the sample sampled from the set of preference samples ; represents the expectation of the sample (x, y w , y l ) on the set of preference samples ; ; represents the expectation of the sample (x, y w ) on the set of preference samples , β represents the coefficient for controlling the divergence penalty intensity, σ(·) represents the sigmoid function, and γ represents the reference margin.
[0040] Optionally, the supervised learning training of the second large language model using the review data set to obtain the reward model includes:
[0041] Constructing a review data set; wherein, the review data set includes review training samples, and the review training samples include review input features and corresponding review output labels, wherein the review input features are the obtained existing papers, and the review output labels include multiple review comments with the first review weight and review comments with the second review weight;
[0042] Determining a second loss function and a second optimizer; wherein, the second loss function is used to measure the difference between the model prediction result and the review output label, and the second optimizer is used to adjust the model parameters to minimize the loss function;
[0043] Inputting the review input features into the second large language model to obtain a second prediction result;
[0044] Calculate the loss value between the second prediction result and the review output label, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and the second optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training epochs, completing the supervised learning training of the second large language model to obtain the reward model.
[0045] The embodiment of the present application also provides an automated scientific research and review system, including:
[0046] An academic search engine for obtaining existing research literature from preprint platforms;
[0047] A researcher model configured to generate an initial research paper based on the obtained existing research literature;
[0048] A reward model configured to review the initial research paper to obtain a review comment, where the review comment includes a review score and a review opinion;
[0049] The researcher model is also configured to iteratively optimize the quality of the initial research paper based on the review comment until a research paper that meets the preset criteria is obtained.
[0050] Optionally, the policy model is obtained by performing supervised learning training on the first large language model using a research paper dataset; the researcher model is obtained by performing reinforcement learning training on the policy model using a preference dataset; the reward model is obtained by performing supervised learning training on the second large language model using a review dataset.
[0051] The automated scientific research and review method provided by the embodiment of the present application can assist scientific research work in many aspects, comprehensively improving the efficiency, quality, and standardization of scientific research, and promoting knowledge integration and innovation. Using the researcher model to quickly generate an initial paper based on existing literature can significantly shorten the time from conception to the first draft, improving scientific research efficiency; and it can automatically iteratively optimize according to the review opinions, reducing the manual modification cost. The reward model conducts a comprehensive and objective review from grammar format to key scientific research indicators based on objective criteria and a large amount of data, pointing out the direction for quality improvement. The researcher model makes targeted improvements accordingly, and multiple iterations of optimization can enable the paper to reach a high scientific research level, comprehensively improving the paper quality. The present application can realize the automation and intelligence of the entire life cycle of scientific research, thereby improving the overall quality and efficiency of scientific research work. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of the automated scientific research and review method according to the embodiment of the present application;
[0053] Figure 2 It is a flowchart of step S100 in the automated scientific research and review method according to the embodiment of the present application;
[0054] Figure 3 A flowchart for supervised learning training of a first large language model using a research paper dataset in the automated scientific research and review method of an embodiment of this application;
[0055] Figure 4 A flowchart for reinforcement learning training of a policy model using a preference pair dataset in the automated scientific research and review method of an embodiment of this application;
[0056] Figure 5 A schematic structural diagram of the automated scientific research and review system of an embodiment of this application. Detailed implementation manners
[0057] Reference is made herein to the various solutions and features of this application with reference to the accompanying drawings. It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be construed as a limitation, but merely as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of this application.
[0058] An automated scientific research and review method provided by an embodiment of this application obtains review comments by reviewing an initial research paper generated based on existing research literature, and then iteratively optimizes the quality of the initial research paper based on the review comments, and finally obtains a research paper that meets the preset criteria, which can realize the full-process high-efficiency automation of scientific research work from topic selection, research, writing to review, and significantly improve the quality of scientific research results.
[0059] The following combines the accompanying drawings to provide a detailed description of this automated scientific research and review method. Figure 1 A flowchart of the automated scientific research and review method of an embodiment of this application, as Figure 1 shown, this method includes the following steps:
[0060] S100. Generate an initial research paper using a researcher model based on the obtained existing research literature.
[0061] It can be understood that, around the research topic of interest, comprehensive literature collection can be carried out from academic databases (such as CNKI, Web of Science, PubMed, etc.), professional books, academic conference proceedings, etc.
[0062] In an embodiment of this application, the required research literature can be retrieved on the preprint platform ArXiv using the academic search engine Semantic Scholar, and the research literature can be a file in LaTeX format.
[0063] A LaTeX - formatted file refers to a file created using the LaTeX typesetting system, and its source file usually has the extension.tex. In such files, users define the content, format, chapter structure, mathematical formulas, chart layout, etc. of the document through specific LaTeX markup commands and syntax structures. For example, specific commands are used to set the title, author, paragraph format, and input complex mathematical formulas. After being processed by a LaTeX compiler, documents in formats such as PDF and DVI can be generated for display and printing on different platforms. Due to its ability to achieve highly customized typesetting and professional and beautiful presentation effects, it is widely used in fields such as academic publishing and scientific writing.
[0064] By reading existing research literature, research questions are identified. Based on the research questions, corresponding solutions are determined and experiments are designed, and an initial research paper is generated according to the solutions and experimental results.
[0065] S200. Use the reward model to review the initial research paper to obtain review comments, where the review comments include review scores and review opinions.
[0066] Specifically, the review score, as a quantitative summary of the quality of the initial research paper, is the result of multi - dimensional evaluation based on a series of preset criteria. The initial research paper is reviewed from dimensions such as reliability, presentability, and contribution. The range of the review score can be set from 1 to 10 points. Different weights can be assigned to each dimension according to actual needs, and finally a comprehensive review score is calculated.
[0067] The review opinion is a more detailed and in - depth qualitative evaluation of the initial research paper. For each part of the initial research paper, a comprehensive analysis and elaboration can be carried out from two aspects: advantages and disadvantages.
[0068] S300. The researcher model iteratively optimizes the quality of the initial research paper based on the review comments until a research paper that meets the preset criteria is generated.
[0069] In the embodiments of this application, the problems existing in the initial research paper in the review opinion can be prioritized. For key problems that seriously affect the scientificity and credibility of the paper, such as fundamental errors in research methods or major defects in logical arguments, they are the primary problems to be solved. For example, if the review opinion points out that there is a sample bias in the research method, resulting in the results not being generally representative, then the first thing to do during iterative optimization is to redesign the sample selection plan and conduct data collection and analysis. For general problems, such as non - standard formats and language expression flaws, they can be optimized centrally after the key problems are solved.
[0070] In one embodiment, in the above step S100, as Figure 2As shown, based on the obtained existing research literature, an initial research paper is generated using a researcher model, specifically including:
[0071] S110. Input the existing research literature into the trained researcher model.
[0072] S120. The researcher model identifies research questions, derives solutions, designs experiments, and writes papers based on the existing research literature, and finally outputs an initial research paper.
[0073] Among them, the researcher model is obtained by performing reinforcement learning training on a policy model using a preference pair dataset. The policy model is obtained by performing supervised learning training on a first large language model using a research paper dataset.
[0074] In the embodiment of the present application, the trained researcher model uses natural language processing technology and machine learning algorithms to perform semantic understanding, information extraction, and knowledge graph construction on the text content in the existing research literature. The researcher model identifies key issues that have not been solved or are controversial in the current research field by analyzing the research background, purpose, methods, results, and discussion parts in the literature. For example, when analyzing the literature related to the battery life of new energy vehicles, the researcher model finds that although certain progress has been made in improving battery materials in existing research, the research on the stability of battery performance under different environmental conditions is still insufficient, which is identified as a potential research question. Through systematic sorting of a large amount of literature, the researcher model can keenly capture those research gaps that have not been fully explored, thereby identifying research questions and pointing the way for subsequent research.
[0075] Based on the identified research questions, the researcher model further uses the knowledge and reasoning ability it has learned to explore possible solutions. For example, for the problem of battery performance stability under different environmental conditions of the above new energy vehicle battery, the researcher model proposes to develop an adaptive battery management system to adjust the charging and discharging strategies of the battery in real time according to factors such as environmental temperature and humidity to improve the stability of battery performance. When deriving solutions, the model does not simply repeat the methods in the existing literature, but tries to propose more forward-looking and feasible ideas through in-depth integration and innovative combination of knowledge.
[0076] After determining the solution, the researcher model can proceed to design corresponding experiments to verify the effectiveness of the solution. Based on the research questions and the characteristics of the solution, the researcher model considers all the key elements of the experiment. For example, the selection and control of experimental variables, the selection of experimental samples, the determination of experimental methods, and the detailed planning of experimental procedures. For the research on the adaptive management system of new energy vehicle batteries, the researcher model can design the following experiment: Select various types of new energy vehicle batteries as experimental samples, set up different environmental simulation chambers to simulate various extreme temperature and humidity conditions, install the adaptive battery management system on the experimental batteries, and evaluate the actual effect of the adaptive battery management system by comparing the performance indicators such as the cruising range and charge-discharge efficiency of the batteries before and after installation under different environments.
[0077] Finally, based on the experimental design and the results of the previous analysis, the researcher model uses natural language generation technology to write an initial research paper according to the standard structure and format requirements of academic papers. A research paper usually includes an introduction section, a methods section, a results section, a discussion section, and a conclusion section. Among them, the introduction section is used to elaborate on the research background, purpose, and a review of existing research; the methods section is used to describe in detail the experimental design, data collection, and analysis methods; the results section is used to present the data and results obtained from the experiment; the discussion section is used to conduct an in-depth analysis of the experimental results, compare them with existing research findings, and explore the limitations of the research and future research directions; the conclusion section is used to summarize the main findings and contributions of the research. For example, when writing a research paper on new energy vehicle batteries, the introduction introduces the urgent demand for battery cruising range in the new energy vehicle industry and the current research status, the methods section details the design principle and experimental setup of the adaptive battery management system, the results section shows the change data of battery performance indicators under different environments, the discussion section analyzes the significance of these results for improving the cruising range stability of the battery and the possible deficiencies of the system, and the conclusion section emphasizes the contribution of this research in solving the problem of battery environmental adaptability.
[0078] In the above embodiment, further, as Figure 3 shown, the first large language model is trained with supervised learning using the research paper dataset to obtain a policy model, including:
[0079] S1210. Construct a research paper dataset; wherein, the research paper dataset includes research training samples, and the research training samples include research input features and corresponding research output labels. Among them, the research input features are the references of existing papers obtained, and the research output labels include the paper outline and the main text. The main text includes research motivation, research methods, experimental setup, and research results.
[0080] Specifically, the references of existing papers are stored in a bib file. A bib file is a plain text file used to store bibliographic citation information, with a file extension of.bib. It plays a crucial role in academic writing and literature management, especially when used in close conjunction with the LaTeX document preparation system.
[0081] S1220. Determine the first loss function and the first optimizer; wherein, the first loss function is used to measure the difference between the model prediction result and the research output label, and the first optimizer is used to adjust the model parameters to minimize the loss function.
[0082] It can be understood that the loss function is used to measure the difference between the model prediction result and the true label (the research output label in the embodiments of the present application). This quantitative representation of the difference can provide a clear goal for the training of the model, that is, by adjusting the model parameters, the value of the loss function is made as small as possible. Specifically, the loss function adopted in the embodiments of the present application can be cross-entropy loss or negative log-likelihood loss, etc.
[0083] For the task of generating text, it is usually desired that the probability distribution of the next word predicted by the model matches the word that appears at that position in the true text. For example, when generating a certain part of a paper outline, the model predicts the probabilities of different expressions such as "research methods" and "experimental results", and the cross-entropy loss can effectively measure the difference between these predicted probabilities and the words used in the true text. It is sensitive to changes in the probability distribution and helps the model learn the correct pattern of text generation.
[0084] Minimizing the negative log-likelihood loss is equivalent to maximizing the probability that the model generates the true text. This prompts the model to learn the parameters that can generate the correct text sequence with a higher probability. For example, when generating the body paragraphs of a paper, the model adjusts the parameters to generate content similar to the true body by minimizing the negative log-likelihood loss.
[0085] The role of the first optimizer is to continuously reduce the value of the loss function by adjusting the model parameters, so that the model gradually learns the patterns and rules in the data and improves the prediction accuracy. In the embodiments of the present application, the first optimizer uses the LoRA-GA (Low Rank Gradient Approximation) optimization algorithm to train the model.
[0086] The LoRA-GA optimization algorithm combines LoRA (Low Rank Adaptation) and GA (Gradient Accumulation). It can reduce the number of trainable parameters of the model while simulating large-batch training through gradient accumulation to further optimize the training process. This combination method can improve the training efficiency and help the model achieve better performance on specific tasks, especially suitable for fine-tuning training of large-scale models in resource-constrained environments.
[0087] S1230. Input the research input features into the first large language model to obtain the first prediction result.
[0088] In the embodiments of the present application, the first large language model can adopt widely used open - source language models such as Mistral - Nemo - 12B, Qwen2.5 - Instruct - 72B, or Mistral - Large - 2 123B. It is optimized using 8×H100 80G GPUs and a key technology ZeRO2 (a further optimized version of Zero Redundancy Optimizer) in DeepSpeed (a deep learning optimization library developed by Microsoft). Using a cluster composed of 8 H100 80G GPUs can further enhance the computing power. Through multi - GPU parallel computing, model training can be carried out simultaneously on multiple GPUs, accelerating the model training process.
[0089] Maximize the context length by setting the Mistral - Nemo - 12B model to 32K tokens, or setting the Qwen2.5 - Instruct - 72B model and Mistral - Large - 2 123B models to 24K tokens. During training, apply FP8 quantization (a technique used in deep learning and machine learning to optimize model training and inference, which involves converting data from a traditional higher - precision floating - point format to an 8 - bit floating - point format) to the model weights and use LoRA - GA for training. Considering memory limitations, samples beyond the preset context length are randomly truncated. Use a batch size of 2×8, a learning rate of 4e - 5, and train for a total of 12,000 steps. These instruction - fine - tuned models support a context window of up to 128K tokens, making them suitable for planning research projects and writing research papers.
[0090] S1240. Calculate the loss value between the first prediction result and the research output label, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and the first optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training epochs, completing the supervised learning training of the first large language model to obtain the policy model.
[0091] In supervised learning, the model generates prediction results based on the input features. For example, in a text generation task, given a text prompt as input, the model generates a subsequent text as the prediction result. The "research output label" is the true and expected output corresponding to the input features, which is the pre - labeled standard answer.
[0092] The role of the loss function is to measure the degree of difference between the prediction result and the true label. By inputting the first prediction result and the research output label into the selected loss function, a specific loss value can be calculated, which reflects the size of the prediction error of the current model in this round of training. The larger the loss value, the greater the gap between the prediction result of the model and the research output label; conversely, the smaller the loss value, the closer the prediction of the model is to the real situation.
[0093] The backpropagation algorithm is the core method for calculating gradients in deep learning. The gradient can be understood as the rate of change of a function at a certain point. In deep learning, the gradient of the loss function with respect to the model parameters represents the change in the loss function value as the model parameters change slightly. For example, if the gradient of a certain model parameter is positive, it means that increasing the value of this parameter will increase the loss function, so the value of this parameter should be decreased when updating the parameters; conversely, if the gradient is negative, the value of this parameter should be increased.
[0094] In each round of training, the first optimizer updates the parameters of the first large language model with a suitable step size according to the calculated gradients, so that the value of the loss function gradually decreases, thereby making the prediction result of the first large language model closer and closer to the research output label.
[0095] "Loss value convergence" means that after multiple rounds of training, the loss value no longer has an obvious downward trend and tends to be stable. This usually indicates that the first large language model has learned the patterns and rules in the training samples and reached a relatively good state. The "preset number of training rounds" is an upper limit of the number of training times preset before the start of training. Due to the complexity of the data or the characteristics of the model, sometimes the loss value may not converge completely, or the convergence speed is very slow. In this case, to avoid indefinite training, a preset number of training rounds is set, and when the training reaches this number of rounds, the training process stops. When the loss value converges or reaches the preset number of training rounds, it is considered that the supervised learning training of the first large language model is completed. At this time, the model has adjusted its own parameters by learning a large number of input-output data pairs and has certain prediction and generalization abilities. The trained first large language model can be used to predict new input data to complete various natural language processing tasks.
[0096] In the above embodiment, further, the construction of the research paper dataset includes the process of constructing research input features and the process of constructing research output labels.
[0097] Among them, the process of constructing research input features is as follows:
[0098] Use an academic search engine to obtain existing papers; specifically, the academic search engine can adopt SemanticScholar.
[0099] Using the application programming interface provided by an academic search engine, retrieve the references cited in the existing papers obtained; specifically, the references cited in the existing papers obtained can be retrieved from the bib file using the Semantic Scholar API.
[0100] Construct research input features using the retrieved references.
[0101] The process of constructing research output labels is as follows:
[0102] Preprocess the obtained existing papers to obtain the main text; specifically, rule-based filtering can be used to preprocess the main text of the obtained existing papers to delete irrelevant content, such as annotation content and acknowledgments marked with "%".
[0103] Input the main text into an existing outline generation model to obtain a paper outline. The existing mature technology for generating paper outline data based on the paper will not be elaborated here. Specifically, the outline generation model can adopt AIPaperGPT or ChatGPT, etc., without limitation.
[0104] Construct research output labels using the paper outline and the main text.
[0105] In the above embodiment, as Figure 4 shown, use the preference pair dataset to perform reinforcement learning training on the policy model to obtain a researcher model, including:
[0106] S1250. Construct a preference pair dataset; wherein, the preference pair dataset includes references, positive samples, and negative samples, where the positive samples are research papers with higher review scores generated by the policy model, and the negative samples are research papers with lower review scores generated by the policy model.
[0107] It can be understood that constructing the preference pair dataset aims to provide a special type of data for the reinforcement learning training of the policy model. By comparing positive samples and negative samples, it helps the policy model learn what are more preferred results and how to distinguish between good and bad. Obviously, in the embodiments of the present application, high-quality research papers are more preferred results.
[0108] S1260. Input the references in the preference pair dataset into the initialized policy model to generate research papers. The policy model is initialized before use, and the model parameters are initialized.
[0109] S1270. Input the generated research paper into the reward model to obtain the review score and feedback it to the policy model, so that the policy model adjusts its own parameters according to the review score to optimize the quality of the next generated research paper. Among them, the reward model is obtained by supervised learning training of the second large language model using the review data set.
[0110] S1280. Input the references in the preference pair data set that have not participated in the training of the policy model into the policy model with adjusted own parameters until all references have participated in the training of the policy model.
[0111] S1290. Use the policy model that generates research papers with higher review scores as the final researcher model after the reinforcement learning training is completed.
[0112] Furthermore, in the above step S1250, the construction of the preference pair data set includes:
[0113] Input the references of the existing papers into the policy model to generate more than one research paper.
[0114] Input the generated research paper into the reward model to obtain the review score corresponding to each research paper. Use the research paper with a higher review score as the positive sample and the research paper with a lower review score as the negative sample. Specifically, sort each research paper according to the review score, select a part of the research papers with higher review scores as the positive sample, and a part of the research papers with lower review scores as the negative sample.
[0115] Construct a preference pair data set using the references, positive samples, and negative samples.
[0116] It should be noted that during the process of performing reinforcement learning training on the policy model, the policy loss function is used to perform iterative training on the policy model; among them, the policy loss function includes the SimPO loss and the negative log-likelihood loss, and the SimPO loss includes the margin loss and the sparse feature loss, and the policy loss function is used to stabilize the reinforcement learning training process of the policy model.
[0117] The policy loss function is:
[0118]
[0119] In the formula, represents the policy loss function; λ represents a hyperparameter used to balance the SimPO loss and the negative log-likelihood loss; x represents the input of the policy model, y w represents the positive sample, y l represents the negative sample, represents the set of preference samples, π θ represents the policy model, π θ (yw |x) represents the probability of the policy model π θ outputting y w ; π θ (y l |x) represents the probability of the policy model π θ outputting y l ; π θ (y w |x) represents a sample sampled from the preference sample set , and represents the expectation of the sample (x, y w , y l ) on the preference sample set ; ; represents the expectation of the sample (x, y w ) on the preference sample set . β represents the coefficient controlling the divergence penalty intensity, σ(·) represents the sigmoid function, and γ represents the reference margin.
[0120] This iterative training mechanism enables the model to gradually learn a better research plan generation strategy and continuously improve with the dynamic change of the generated paper quality.
[0121] Furthermore, in the above embodiment, the second large language model is trained by supervised learning using the review data set to obtain a reward model, including:
[0122] Construct a review data set; wherein, the review data set includes review training samples, and the review training samples include review input features and corresponding review output labels. Among them, the review input features are the obtained existing papers, and the review output labels include review comments with multiple first review weights and review comments with second review weights. Specifically, the review comments involve four parts: work summary, identifying advantages and disadvantages, clarifying issues, and numerical scores for soundness, presentation, contribution, and overall rating.
[0123] Determine the second loss function and the second optimizer; wherein, the second loss function is used to measure the difference between the model prediction result and the review output label, and the second optimizer is used to adjust the model parameters to minimize the loss function.
[0124] Input the review input features into the second large language model to obtain a second prediction result.
[0125] Calculate the loss value between the second prediction result and the review output label, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and the second optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training epochs, completing the supervised learning training of the second large language model.
[0126] In the embodiments of the present application, the second large language model can adopt the Mistral-Large-2 model, configure an 8xH100 80G GPU cluster, that is, use a computing cluster composed of 8 NVIDIA H100 GPUs during training, and the video memory of each GPU is 80GB. Set training hyperparameters such as the learning rate, batch size, and number of training epochs.
[0127] The learning rate determines the step size of model parameter updates in each training iteration. A smaller learning rate means that the model parameter updates are relatively slow, the training process is more stable, but the convergence speed will be relatively slow; a larger learning rate will make the parameter update steps larger, which may cause the model to skip the optimal solution during training, unable to converge, and even the loss function value may diverge. For example, the learning rate can be set to 1e-5.
[0128] The batch size is set to 4x8. This means that in each training iteration, the model will process 4 batches of data at the same time, and each batch contains 8 samples. A larger batch size can utilize the parallel computing power of the GPU to improve the training efficiency, and at the same time can more accurately reflect the overall distribution characteristics of the data when calculating the gradient, making the model convergence more stable.
[0129] According to the specific dataset size, model complexity, and training effect, the number of training epochs can be set to 12. One epoch means that the model trains the entire review dataset once. Increasing the number of training epochs usually allows the model to better learn the features and patterns in the data, but if the number of training epochs is too large, the model may overfit, that is, perform well on the training set but degrade in performance on the test set or new data.
[0130] Based on the same concept, as Figure 5 shown, the embodiments of the present application also provide an automated scientific research and review system, including:
[0131] An academic search engine for obtaining existing research literature from preprint platforms.
[0132] A researcher model configured to generate an initial research paper based on the obtained existing research literature.
[0133] A reward model configured to review the initial research paper to obtain review comments, where the review comments include review scores and review opinions.
[0134] The researcher model is also configured to iteratively optimize the quality of the initial research paper based on review comments until a research paper that meets the preset criteria is obtained.
[0135] In the embodiment of the present application, the researcher model is obtained by performing reinforcement learning training on the policy model using a preference pair dataset; the policy model is obtained by performing supervised learning training on the first large language model using a research paper dataset; the reward model is obtained by performing supervised learning training on the second large language model using a review dataset.
[0136] In the above embodiment, the research paper dataset includes research training samples, and the research training samples include research input features and corresponding research output labels. Among them, the research input features are the references of the existing papers obtained, and the research output labels include the paper outline and the main text. The main text includes the research motivation, research methods, experimental settings, and research results.
[0137] Performing supervised learning training on the first large language model using the research paper dataset to obtain a policy model includes:
[0138] Construct a research paper dataset.
[0139] Determine a first loss function and a first optimizer; wherein, the first loss function is used to measure the difference between the model prediction result and the research output label, and the first optimizer is used to adjust the model parameters to minimize the loss function.
[0140] Input the research input features into the first large language model to obtain a first prediction result.
[0141] Calculate the loss value between the first prediction result and the research output label, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and the first optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training rounds, completing the supervised learning training of the first large language model to obtain a policy model.
[0142] In the above embodiment, the preference pair dataset includes references, positive samples, and negative samples. Among them, the positive samples are the research papers with higher review scores generated by the policy model, and the negative samples are the research papers with lower review scores generated by the policy model.
[0143] Performing reinforcement learning training on the policy model using the preference pair dataset to obtain a researcher model includes:
[0144] Construct a preference pair dataset.
[0145] Input the references in the preference pair dataset into the initialized policy model to generate research papers.
[0146] Input the generated research papers into the reward model to obtain review scores and feedback them to the policy model, so that the policy model adjusts its own parameters according to the review scores to optimize the quality of the research papers generated next time.
[0147] Input the references in the preference pair dataset that have not participated in the training of the policy model into the policy model with adjusted self-parameters until all references have participated in the training of the policy model.
[0148] Take the policy model that generates research papers with higher review scores as the final researcher model after the reinforcement learning training is completed.
[0149] In the above embodiment, the review dataset includes review training samples, and the review training samples include review input features and corresponding review output labels. Among them, the review input features are the obtained existing papers, and the review output labels include multiple review comments with the first review weight and review comments with the second review weight.
[0150] The supervised learning training of the second large language model using the review dataset to obtain the reward model includes:
[0151] Construct a review dataset.
[0152] Determine the second loss function and the second optimizer; where the second loss function is used to measure the difference between the model prediction result and the review output label, and the second optimizer is used to adjust the model parameters to minimize the loss function.
[0153] Input the review input features into the second large language model to obtain a second prediction result.
[0154] Calculate the loss value between the second prediction result and the review output label, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and the second optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training rounds, and complete the supervised learning training of the second large language model to obtain the reward model.
[0155] Based on the same concept, an embodiment of the present application also provides an electronic device, including a processor and a memory. The memory stores an executable program, and the processor executes the executable program to perform the steps of the method described above.
[0156] The storage medium in this embodiment may be included in an electronic device / system; it may also exist independently without being assembled into the electronic device / system. The above storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0157] It should be understood that the first, second, third, fourth, and various numerical numbers involved herein are only for the convenience of description and are not used to limit the scope of the present application.
[0158] It should also be understood that the term "and / or" herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0159] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor, or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0160] In various embodiments of the present application, the magnitude of the sequence numbers of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0161] Those of ordinary skill in the art can realize that the various illustrative logical blocks (ILB) and steps described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0162] In several embodiments provided by this application, it should be understood that the disclosed automated scientific research and review methods and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units (or modules) is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0163] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0164] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0165] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0166] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims described above.
Claims
1. An automated scientific research and review method, characterized in that: include: Based on the acquired existing research literature, the researcher model is used to generate initial research papers; Using the reward model to review the initial research paper and obtain review comments, wherein the review comments include review scores and review opinions; The researcher model iteratively optimizes the quality of the initial research paper based on the review comments until a research paper that meets the preset standards is generated.
2. The method according to claim 1, characterized in that Based on the existing research literature obtained, the researcher model is used to generate initial research papers, including: The existing research literature is input into the trained researcher model. The researcher model identifies research problems, derives solutions, designs experiments and writes papers based on the existing research literature, and finally outputs an initial research paper. The researcher model is obtained by using the preference pair data set to perform reinforcement learning training on the strategy model. The strategy model is obtained by using the research paper data set to perform supervised learning training on the first language model.
3. The method according to claim 2, characterized in that The first language model is trained with supervised learning using the research paper dataset to obtain a strategy model, including: Constructing a research paper dataset; wherein the research paper dataset includes research training samples, and the research training samples include research input features and corresponding research output labels, wherein the research input features are references of existing papers obtained, and the research output labels include paper outlines and texts, and the texts include research motivations, research methods, experimental settings, and research results; Determine a first loss function and a first optimizer; wherein the first loss function is used to measure the difference between the model prediction result and the research output label, and the first optimizer is used to adjust the model parameters to minimize the loss function; Input the research input features into the first language model to obtain a first prediction result; Calculate the loss value of the first prediction result and the research output label, calculate the gradient of the loss function to the model parameters through the back propagation algorithm, and the first optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training rounds, thereby completing the supervised learning training of the first language model.
4. The method according to claim 3, characterized in that The construction of the research paper dataset includes a process of constructing research input features and a process of constructing research output labels; The process of constructing the research input features is as follows: Use academic search engines to obtain existing papers; Using the application programming interface provided by academic search engines, retrieve the references cited in the existing papers; Construct research input features using retrieved references; The process of constructing the research output label is: Preprocess the acquired existing papers to obtain the main text; Input the main text into the existing outline generation model to obtain the paper outline; Use the paper outline and body to construct your research output tags.
5. The method according to claim 2, characterized in that: The policy model is trained through reinforcement learning using the preference pair dataset to obtain the researcher model, including: Constructing a preference pair data set; wherein the preference pair data set includes references, positive samples and negative samples, wherein the positive samples are research papers with higher review scores generated by the strategy model, and the negative samples are research papers with lower review scores generated by the strategy model; Inputting the references in the preference pair dataset into the initialized strategy model to generate a research paper; The generated research paper is input into the reward model, the review score is obtained and fed back to the strategy model, so that the strategy model adjusts its own parameters according to the review score to optimize the quality of the next generated research paper; wherein the reward model is obtained by using the review data set to conduct supervised learning training on the second largest language model; Input the references in the preference pair data set that have not yet participated in the strategy model training into the strategy model after its own parameter adjustment, until all references participate in the strategy model training; The strategy model that generates research papers with higher review scores is used as the researcher model that is finally trained through reinforcement learning.
6. The method according to claim 5, characterized in that The constructed preference pair data set includes: Input the references of existing papers into the strategy model to generate research papers; Input the generated research paper into the reward model to obtain the review score; Research papers with higher review scores are used as positive samples, and research papers with lower review scores are used as negative samples; The preference pair dataset is constructed using references, positive samples and negative samples.
7. The method according to claim 5, characterized in that In the process of reinforcement learning training of the policy model, the policy loss function is used to iteratively train the policy model; The strategy loss function is: In the formula, represents the policy loss function; λ represents a hyperparameter used to balance the SimPO loss and the negative log-likelihood loss; x represents the input of the policy model, y w represents a positive sample, y l represents negative samples, represents the preferred sample set, π θ represents the policy model, π θ (y w |x) means taking x as input, the strategy model π θ Output y w The probability of θ (y l |x) means taking x as input, the strategy model π θ Output y l The probability of θ (y w |x) represents the sample set from preference The samples obtained by sampling are Represents a sample (x, y w ,y l ) in the preferred sample set expectations; Represents a sample (x, y w ) in the preferred sample set , β represents the control divergence penalty strength coefficient, σ(·) represents the sigmoid function, and γ represents the reference margin.
8. The method according to claim 5, characterized in that The second largest language model is trained with supervised learning using the review dataset to obtain a reward model, including: Constructing a review data set; wherein the review data set includes review training samples, and the review training samples include review input features and corresponding review output labels, wherein the review input features are obtained existing papers, and the review output labels include multiple review comments with a first review weight and a review comment with a second review weight; Determine a second loss function and a second optimizer; wherein the second loss function is used to measure the difference between the model prediction result and the review output label, and the second optimizer is used to adjust the model parameters to minimize the loss function; Input the review input features into the second largest language model to obtain a second prediction result; Calculate the loss value of the second prediction result and the review output label, calculate the gradient of the loss function to the model parameters through the back propagation algorithm, and the second optimizer updates the model parameters according to the gradient to gradually reduce the loss value until the loss value converges or reaches the preset number of training rounds, thereby completing the supervised learning training of the second largest language model and obtaining the reward model.
9. An automated scientific research and review system, characterized in that: include: Academic search engines for obtaining existing research literature from preprint platforms; The researcher model, which is configured to generate an initial research paper based on the acquired existing research literature; The reward model is configured to review the initial research paper and obtain review comments, wherein the review comments include review scores and review opinions; The researcher model is also configured to iteratively optimize the quality of the initial research paper based on the review comments until a research paper that meets the preset standards is obtained.
10. The system according to claim 9, characterized in that The strategy model is obtained by using the research paper data set to train the first language model through supervised learning; the researcher model is obtained by using the preference pair data set to train the strategy model through reinforcement learning; The reward model is obtained by training the second largest language model through supervised learning using the review dataset.
Citation Information
Patent Citations
Scientific research auxiliary system and method based on large language model
CN118940833A
Paper generation method and device of big language model based on RAG
CN119357413A
Cited By
Scientific research review portrait construction method based on large language model and related equipment
CN121350708A
Medical automation scientific research method and device
CN122154632A