Multi-agent-driven paper intelligent reviewing method and system
By constructing a multi-agent driven intelligent paper review system, the problems of single review dimensions and distorted results in existing technologies are solved, achieving efficient and interpretable paper review and improving review quality and efficiency.
Patent Information
- Application Number
- CN202610075112.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-01
- Estimated Expiration
- 2046-01-20
AI Technical Summary
Existing intelligent paper review technologies lack in-depth design, have a single review dimension, and are difficult to meet the actual needs of paper review. They also suffer from problems such as distorted review results and inability to provide explainable review criteria.
A multi-agent driven intelligent paper review system is constructed. By building a review point list based on the basic review dimensions of paper quality review, and combining the background information of review experts and a large language model, a group of intelligent review experts is constructed, and an intelligent review workflow is designed to realize intelligent paper review.
It significantly improves the quality and consistency of reviews, provides detailed problem descriptions and targeted improvement suggestions, shortens the review cycle, reduces the cost of manual intervention, and supports large-scale batch reviews.
Smart Images

Figure CN121961793A_ABST
Abstract
Description
A Multi-Agent Driven Intelligent Paper Review Method and System Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multi-agent-driven intelligent paper review method and system. Background Technology
[0002] Thesis quality review is a core component of the higher education quality assurance system, directly impacting the quality of postgraduate training and the maintenance of academic integrity. Each thesis management department must conduct internal quality reviews to identify and rectify problems early. However, the current mainstream model of manual expert review consumes significant human, financial, and time resources in areas such as expert invitation, communication and coordination, and waiting for results, making it difficult to achieve rapid and comprehensive monitoring of thesis quality.
[0003] In recent years, breakthroughs in artificial intelligence technology, especially large language models, in natural language understanding and text analysis have provided technological possibilities for intelligent paper review. However, existing technologies still have significant limitations: First, multi-agent paper review frameworks only define the roles of agents and engineer prompt words, failing to deeply reflect the academic background, research expertise, and review style of experts. Furthermore, they lack systematic review point design, leading to a highly arbitrary review process. Moreover, this type of technology is primarily suited for academic journal paper review scenarios, fundamentally differing from the review requirements and evaluation standards of academic papers. Second, end-to-end review model training methods often focus on a single evaluation dimension, lacking comprehensive evaluation capabilities. Relying on publicly available papers and review data for training can easily lead to data leakage, resulting in distorted evaluation results. Additionally, they cannot provide explainable review criteria or specific improvement suggestions.
[0004] In summary, none of the existing technologies mentioned above can meet the core needs of paper review in order to ensure review quality, provide in-depth problem analysis, and take into account disciplinary differences and temporal characteristics. Summary of the Invention
[0005] This invention provides a multi-agent-driven intelligent paper review method and system to address the following technical problems: existing intelligent paper review technologies lack in-depth design for the construction of expert agents, have a single review dimension, and insufficient review comprehensiveness, making it difficult to meet the actual needs of intelligent paper review.
[0006] The embodiments of this invention adopt the following technical solution: On one hand, the embodiments of this invention provide a multi-agent driven intelligent paper review method, the method comprising: constructing a review point list based on the basic review dimensions of paper quality review; constructing a large review model based on historical papers and expert review data; constructing a group of review expert intelligent agents based on the background information of review experts and the large review model; constructing an intelligent review workflow based on the group of review expert intelligent agents, the review point list and review rules; inputting the paper to be reviewed into the intelligent review workflow, so as to match several suitable review expert intelligent agents in the group of review expert intelligent agents to perform intelligent review of the paper to be reviewed, and obtain intelligent review results.
[0007] In one feasible implementation, a review point list is constructed based on the fundamental review dimensions of thesis quality review. Specifically, this includes: determining the fundamental review dimensions of thesis quality review according to the thesis quality review standards to obtain a list of primary review dimensions; wherein the list of primary review dimensions includes, but is not limited to, the following review dimensions: topic selection, theory, frontier, methodology, writing, and innovation; analyzing the detectable review indicators in each review dimension to refine each review dimension into several secondary review indicators, thus obtaining the review point list; constructing a hybrid detection method for each secondary review indicator; wherein the hybrid detection method is composed of a rule-based detection method and a model-based detection method; the rule-based detection method uses preset rules to detect the reasonableness of the results of each secondary review indicator; the model-based detection method uses a large model to detect the reasonableness of the results of each secondary review indicator.
[0008] In one feasible implementation, a large-scale review model is constructed based on historical papers and expert review data. Specifically, this includes: acquiring historically reviewed papers and corresponding expert review data; collecting background information on the reviewers of the papers; constructing an initial training dataset based on the historical papers, the expert review data, the review point list, and the reviewer background information; performing sample augmentation on the initial training dataset to obtain a data-enhanced training dataset; constructing a large-scale review model based on a large language model; performing supervised training on the large-scale review model using the training dataset; if the trained large-scale review model does not meet the testing requirements, constructing a reward model for the large-scale review model based on the review point list; and performing reinforcement learning on the large-scale review model using the reward model to obtain an optimized large-scale review model.
[0009] In one feasible implementation, an initial training dataset is constructed based on the historical papers, the expert review data, the review point list, and the background information of the reviewers. Specifically, this includes: parsing the historical papers and the expert review data to obtain the paper text and review data text; automatically reviewing the paper text according to the review indicators in the review point list to obtain a multi-dimensional review result list; extracting structured information from the paper text and review data text; compressing the background information of the reviewers in a structured manner; and concatenating the compressed background information of the reviewers, the structured information, and the text content corresponding to the multi-dimensional review result list using a prompt to obtain the initial training dataset.
[0010] In one feasible implementation, the initial training dataset is subjected to sample augmentation processing to obtain a data-augmented training dataset. Specifically, this includes: performing controlled degradation operations on high-quality academic papers to generate negative samples; wherein, the controlled degradation operations include, but are not limited to: content consistency violation, format standardization violation, and academic rigor violation; and merging the negative samples into the initial training dataset to obtain the data-augmented training dataset.
[0011] In one feasible implementation, the large-scale evaluation model is subjected to supervised training using the training dataset, specifically including: constructing a training objective function based on a multi-task learning framework; using the training dataset as input to the large-scale evaluation model, and performing supervised training on the large-scale evaluation model based on the training objective function; wherein, the supervised training process includes: fine-tuning all or some parameters of the large-scale evaluation model.
[0012] In one feasible implementation, the evaluation model is subjected to reinforcement learning through the reward model to obtain an optimized evaluation model. Specifically, this includes: constructing a reward function based on the reward types included in the reward model; constructing a policy optimization model based on the GRPO reinforcement learning framework, using the parameters of the trained evaluation model as the initial weights of the policy optimization model; constructing an optimization objective function based on the reward function and the initial weights; and performing reinforcement learning on the trained evaluation model based on the optimization objective function to obtain the optimized evaluation model.
[0013] In one feasible implementation, a group of intelligent reviewer agents is constructed based on the background information of the reviewers and the large-scale review model. Specifically, this includes: collecting academic profile data of each reviewer from preset channels to obtain their background information; wherein the preset channels include, but are not limited to, academic resource databases, research project databases, and institutional websites; constructing an intelligent reviewer agent for each reviewer based on their background information, thus forming the group of intelligent reviewer agents; classifying the group of intelligent reviewer agents according to paper-related metadata, and establishing a virtual paper review committee agent for each category, used to match suitable intelligent reviewer agents for papers to be reviewed according to a matching strategy; wherein the paper-related metadata includes, but is not limited to, the subject and major; the matching strategy is related to the following factors: professional relevance, review experience, and perspective diversity.
[0014] In one feasible implementation, an intelligent review workflow is constructed based on the group of review expert agents, the list of review points, and the review rules. Specifically, this includes: constructing a data upload and parsing node, a structured information extraction node, a review expert agent selection node, a parallel review node, and a comprehensive review result output node, constituting the intelligent review workflow. The data upload and parsing node is used to parse the uploaded paper to be reviewed, obtaining the paper's main text and related metadata. The related metadata includes, but is not limited to, the subject and major. The structured information extraction node is used to extract structured information from the paper's main text. The review expert agent selection node is used to select several suitable review expert agents from the group of review expert agents based on the paper's related metadata. The parallel review node is used to perform parallel reviews using the selected review expert agents based on the list of review points and the review rules. The comprehensive review result output node is used to summarize the review results of the selected review expert agents and determine the comprehensive review result of the paper to be reviewed.
[0015] On the other hand, embodiments of the present invention also provide a multi-agent driven intelligent paper review system. The system includes: a review big model module, which constructs a review point list based on the basic review dimensions of paper quality review; and constructs a review big model based on historical papers and expert review data; an expert review intelligent agent group module, used to construct a group of review expert intelligent agents based on the background information of review experts and the review big model; and an intelligent review workflow module, used to construct an intelligent review workflow based on the group of review expert intelligent agents, the review point list, and review rules; and inputs the paper to be reviewed into the intelligent review workflow to match suitable review expert intelligent agents from the group of review expert intelligent agents to perform intelligent review of the paper, thereby obtaining intelligent review results.
[0016] Compared with existing technologies, the multi-agent-driven intelligent paper review method and system provided in this invention have the following beneficial effects: This invention significantly improves the quality and consistency of review. By integrating the paper text, structured meta-information, review checklists, and expert review data to construct a dedicated review model, and combining real expert profiles and a knowledge base to create review expert agents, the review results are more aligned with the evaluation standards and academic norms of human experts. Through a hierarchical review checklist-driven systematic diagnostic mechanism, it can provide detailed problem descriptions, supporting evidence, and targeted improvement suggestions for core dimensions such as the paper's topic value, theoretical foundation, research methods, and writing standards, breaking the limitations of traditional rating methods. It greatly improves review efficiency. Through automated document parsing and parallel review process design, it effectively shortens the review cycle of a single paper, supports large-scale batch review, and significantly reduces the cost and time consumption of manual intervention. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 is a flowchart of a multi-agent driven intelligent paper review method provided by an embodiment of the present invention; Figure 2 is a detailed architecture diagram of a multi-agent driven intelligent paper review method provided by an embodiment of the present invention; Figure 3 is a flowchart of a basic support engine construction provided by an embodiment of the present invention; Figure 4 is a schematic diagram of an intelligent review workflow provided by an embodiment of the present invention; Figure 5 is a structural schematic diagram of a multi-agent driven intelligent paper review system provided by an embodiment of the present invention. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.
[0019] This invention provides a multi-agent driven intelligent paper review method, as shown in Figure 1. The multi-agent driven intelligent paper review method specifically includes steps S101-S104: S101, constructing a review point list based on the basic review dimensions of paper quality review.
[0020] Specifically, based on the paper quality review standards, the basic review dimensions for paper quality review are obtained, resulting in a list of primary review dimensions; among them, the list of primary review dimensions includes, but is not limited to, the following review dimensions: topic selection, theory, frontier, methodology, writing, and innovation.
[0021] As a feasible implementation method, the papers mentioned in this invention include at least academic papers and dissertations. The multi-agent driven intelligent paper review method proposed in this invention can be used for intelligent review of dissertations and also for intelligent review of academic papers.
[0022] Furthermore, the detectable review indicators in each review dimension are analyzed to refine each review dimension into several secondary review indicators, resulting in a list of review points.
[0023] Furthermore, a hybrid detection method is constructed for each secondary review indicator; this hybrid method consists of a combination of rule-based detection and model-based detection. The rule-based detection method uses pre-defined rules to assess the reasonableness of the results for each secondary review indicator. The model-based detection method uses a large model to assess the reasonableness of the results for each secondary review indicator.
[0024] Figure 2 is a detailed architecture diagram of a multi-agent driven intelligent paper review method provided by an embodiment of the present invention. As shown in Figure 2, the present invention provides a basis for the construction of a large review model by constructing a review point list, and then constructs an expert intelligent agent group through the large review model. Then, together with the constructed review point list, the intelligent review workflow is realized to obtain the final review result.
[0025] As a feasible implementation method, the review checklist is the key to achieving interpretability of the review process. Figure 3 is a flowchart of the construction of a basic support engine provided by an embodiment of the present invention. As shown in the "Module B: Construction of Review Checklist" section of Figure 3, the present invention extracts the standards and academic norms requirements of the thesis review form based on the academic review forms and thesis writing norms guidance documents of graduate students' historical theses, and designs a hierarchical review system: (1) First-level review dimensions: First, based on academic evaluation theory and thesis quality standards, six core review dimensions are divided: .
[0026] (2) Secondary review indicators: Each primary dimension is further refined into specific, measurable secondary indicators. For example... For example: .
[0027] in, Used to check whether heading levels, formula numbers, table styles, etc., conform to the specifications; Used to verify the consistency between the reference citation format, in-text citations and the end-of-text list; Used to identify semantic jumps or contradictions between paragraphs or chapters; Used to evaluate the clarity of charts, the completeness of annotations, and their relevance to the text; It is used to detect grammatical errors, inappropriate terminology, and colloquial expressions.
[0028] (3) Formal Representation of Detection Rules: For each secondary review indicator, a hybrid detection method based on rules and models is designed. Taking the secondary review indicators... For example, the rule layer uses regular expressions to match citation formats and count the completeness of the mapping between citations in the paper and the reference list; the model layer calls a general large model, inputs paper excerpts and corresponding citations, and determines whether the citations support the arguments.
[0029] The detection function is formalized as follows: Where M represents the main text of the paper; For the reference list, These are parameters for a general large model. The output is a set of triples. , representing the problem severity score, problem description text, and relevant location index, respectively.
[0030] S102. Construct a large-scale review model based on historical papers and expert review data.
[0031] Specifically, we obtain historical papers and their corresponding expert review data; and we obtain background information on the reviewers of the papers from publicly available sources.
[0032] Then, based on historical papers, expert review data, review point lists, and background information of reviewers, an initial training dataset is constructed. Specifically, this includes: parsing historical papers and expert review data separately to obtain paper text and review data text; automatically reviewing the paper text according to the review indicators in the review point list to obtain a multi-dimensional review result list; extracting structured information from the paper text and review data text; compressing the background information of reviewers in a structured manner; and concatenating the compressed background information of reviewers, structured information, and the text content corresponding to the multi-dimensional review result list using prompts to obtain the initial training dataset.
[0033] Furthermore, the initial training dataset is augmented to obtain a data-enhanced training dataset. Specifically, this includes: performing controlled degradation operations on high-quality academic papers to generate negative samples; wherein, the controlled degradation operations include, but are not limited to: content consistency violation, format standardization violation, and academic rigor violation; and merging the negative samples into the initial training dataset to obtain the data-enhanced training dataset.
[0034] Furthermore, a large-scale evaluation model is constructed based on a large language model. Supervised training of this model is performed using a training dataset, specifically including: constructing a training objective function by fusing cross-entropy classification loss and negative log-likelihood loss based on a multi-task learning framework. The training dataset is used as input to the large-scale evaluation model, and supervised training is performed based on the training objective function. The supervised training process specifically includes: fine-tuning all parameters of the large-scale evaluation model, or freezing the core parameters and training only the low-rank factorization matrix. Either approach can be selected based on actual requirements.
[0035] Optionally, if the trained review model fails to meet the testing requirements, a reward model for the review model is constructed based on the review point list. The review model is then used for reinforcement learning to obtain an optimized review model. Specifically, this includes: firstly, constructing a reward model for the review model based on the review point list; wherein the reward model includes at least one or more of the following reward types: expert review level reward, comment consistency reward, review point coverage reward, compliance reward, and writing style alignment reward.
[0036] Then, a reward function is constructed based on the reward types included in the reward model. A policy optimization model is built using the GRPO reinforcement learning framework, and the parameters of the trained large-scale review model are used as the initial weights for the policy optimization model. An optimization objective function is constructed based on the reward function and the initial weights. Reinforcement learning is then performed on the trained large-scale review model based on the optimization objective function to obtain the optimized large-scale review model.
[0037] As a feasible implementation method, as shown in "Module A: Construction of the Large-Scale Review Model" in Figure 3, the construction and training process of the large-scale review model is as follows: 1. Construction and enhancement of training data: First, the existing papers in the paper database and their corresponding expert review data are systematically processed to construct a high-quality training corpus. Specifically, this includes: 1) Document parsing: OCR text parsing is performed on the PDF or scanned documents of the papers to obtain the formatted Markdown text. At the same time, the corresponding review text is also converted into Markdown based on the expert review data for subsequent unified processing.
[0038] 2) Structured Information Extraction: Automatically extract k-dimensional features from the main text of the paper, including: Meta-information dimension: title, abstract, keywords, major, discipline, year of award; Academic achievements dimension: number of published papers and journal level, participation in research projects, patent applications, awards; Document statistics dimension: number of chapters, total text, number of figures and tables, size of references; Review information dimension: review level, review comments. The extraction result is denoted as a vector. .
[0039] 3) Generation of review point annotation data: Based on the obtained list of review points, a general large model (such as GPT-4, Claude, etc.) is used to automatically analyze the problems of all papers and generate an m-dimensional review vector. ,in Indicates the first The severity of the problem at each review point is indicated by a scale of 0 (no problem), 1 (minor problem), and 2 (serious problem).
[0040] 4) Incorporating expert background information: Automatically extract experts' academic profiles, research directions, representative papers, and evaluation style characteristics from public sources (such as personal homepages, Google Scholar, CNKI, etc.). After structurally compressing this information using a large model, it is concatenated with other structured information via a prompt and used as model input.
[0041] 5) Sample augmentation strategy: To address the problem of uneven distribution of review levels, such as the fact that there are far more excellent papers than problematic papers, two types of data augmentation methods are designed: Positive augmentation: The complete paper is broken down into chapters to construct sub-samples of abstract + introduction + conclusion + key chapters, which inherit the original paper's level labels and expand the number of normal samples.
[0042] Negative amplification: Controllable degradation of high-quality papers to generate negative samples. This includes: Disruption of content consistency: Randomly deleting or replacing paragraphs to create logical breaks; Disruption of formatting standards: Modifying reference citation formats and introducing errors in figure and table labeling; Disruption of academic rigor: Reducing the frequency of key terms and increasing colloquial expressions.
[0043] 6) Data partitioning: The constructed data is divided into training set and validation / test set by year, and the data of the latest year is divided into evaluation set and validation set.
[0044] 2. Training Strategy for the Large-Scale Review Model: 1) Supervised Training: Considering the unique characteristic of paper review tasks—outputting brief comments and grade labels rather than generating long texts—a base model with a moderate parameter size (within 10 bytes) (such as Qwen3-7B or MiniCPM4-8B) is selected to construct the large-scale review model. Supervised fine-tuning is then performed on the aforementioned training dataset. Optionally, all parameters of the large-scale review model can be fine-tuned, or LoRA technology can be used for efficient parameter fine-tuning, freezing the core parameters of the base model and training only the low-rank factorization matrix to reduce computational costs.
[0045] The training objective function employs a multi-task learning framework, jointly optimizing the rating classification loss and the rating generation loss: ;in: Cross-entropy is used for classification loss; The negative log-likelihood loss generated for the comment sequence.
[0046] Input representation: Text M, expert background information B, and structured features and review vector Concatenate according to the template, for example: [Paper Text] {The first N tokens of M} [Expert Background Information] Research Direction: {b_Research Direction}, Academic Introduction: {b_Academic Introduction}, Representative Papers: {b_Representative Papers}... [Structured Information] Major: {s_Major}, Published Papers: {s_Number of Papers},... [Review Results] Topic Value: {c_1}, Theoretical Basis: {c_2},... 2) Reinforcement Learning Stage: After completing supervised training, further reinforcement learning is carried out to improve model performance. The reward model is set as follows: Overall Expert Level Reward: The final level A / B / C / D given by the experts is mapped to a relative score to guide the overall trend of the model to align with the real review standards.
[0047] Comment consistency reward: The semantic similarity between the comments generated by the model and the real expert comments is measured using a text similarity model. A positive reward is given when the semantic distance is less than a preset threshold; if the content generated by the model deviates significantly from the expert opinion, a penalty is imposed.
[0048] Review point coverage reward: Based on the aforementioned review point list, the evaluation generated by the strategy model is checked to see if key review points are involved. If more dimensions are involved and accurately expressed, the reward is increased; if important dimensions are omitted, the reward is decreased. This mechanism ensures that the model's evaluation has an "expert structure".
[0049] Compliance Rewards: The system checks whether the generated content violates academic content guidelines. If the text passes the security check, a positive reward is given; otherwise, a penalty is imposed.
[0050] Style alignment reward: The statistical distribution of expert comments in terms of sentence length, density of professional terms, and expression style is used to constrain the comments generated by the model to conform to the expression style of experts through KL divergence or probability distance.
[0051] The above rewards are weighted to obtain the reward function, and the weights can be determined through validation set tuning.
[0052] Reinforcement learning strategy model Using the large model parameters after supervised fine-tuning as initial weights ensures that the model does not deviate from its learned basic evaluation ability during the RL phase, while further improving its expert style and evaluation logic.
[0053] Policy optimization methods can employ proximal policy optimization (PPO) based approaches. Building upon traditional PPO, a reinforcement learning (Generative Reward Policy Optimization, GRPO) framework can be introduced to further enhance reward distribution modeling capabilities in generative review tasks. The objective function can be written as: ;in, To generate a complete sequence, For the reward function, The KL regularization coefficients are... This indicates the initial weights.
[0054] In practice, GRPO can be used in combination with PPO: PPO ensures the stability of policy updates, while GRPO provides stronger sequence-level reward reinforcement, thereby obtaining a generation policy that is closer to the logic of expert review.
[0055] 3. Model Evaluation and Iterative Optimization: The latest year's review data is divided into a validation set and a test set in a 5:5 ratio, with data from other years used for training. Evaluation metrics include: Grade prediction accuracy: Weighted F1 score: comprehensively considers precision and recall at each level; consistency with human reviewers: Cohen's Kappa coefficient, which measures inter-reviewer reliability.
[0056] S103. Based on the background information of the review experts and the review model, construct a group of intelligent review experts; based on the group of intelligent review experts, the review point list and review rules, construct an intelligent review workflow.
[0057] Specifically, academic profile data for each reviewer is collected through preset channels to obtain their background information. These preset channels include, but are not limited to, academic resource databases, research project databases, and institutional websites. Then, based on the reviewer background information, a reviewer agent is constructed for each reviewer, forming the reviewer agent group.
[0058] Furthermore, the group of review expert agents is categorized according to the paper-related metadata, and a virtual paper review committee agent is formed for each category. This agent is used to match suitable review expert agents to the papers to be reviewed according to the matching strategy. The paper-related metadata includes, but is not limited to, the subject and major to which the paper belongs. The matching strategy is related to the following factors: professional relevance, review experience, and diversity of perspectives.
[0059] As a feasible implementation method, as shown in the "Module C: Construction of Review Expert Intelligent Agent Group" section of Figure 3, a review expert intelligent agent group is constructed based on the review big model trained above. The construction process is as follows: (1) Multi-source data collection of expert profiles: For each expert in the historical review expert information database list We collect academic profile data from multiple channels, including but not limited to the following: academic resource databases: obtaining the expert's paper publication records from Aminer, Google Scholar, and CNKI; research project databases: collecting information on projects led or participated in; institution homepages: extracting structured information such as expert profiles, research directions, and educational backgrounds.
[0060] (2) Personalized Prompt Design for Expert Agents: Each expert agent It consists of two parts: ;in, The system-level personalized Prompt includes expert identification information and review style guidance; Define the evaluation criteria and output format for a task-level general Prompt.
[0061] System-level Prompt example: You are {Expert Name}, a professor from {Institution} {Discipline}.
[0062] Your research interests include: {List of research interests}.
[0063] You have published {number of papers} academic papers and led {number of projects} research projects.
[0064] When reviewing papers, you focus on {review characteristics, such as "methodological rigor" and "innovation"}.
[0065] The paper currently being reviewed was submitted in {year}. Please evaluate it based on the academic context at that time.
[0066] (3) Organization and management of expert intelligent agents: Taking dissertations as an example, all expert intelligent agents are organized hierarchically by discipline and major: a virtual dissertation review committee intelligent agent is set up for each discipline, which is responsible for matching suitable expert intelligent agents for specific dissertations. The matching strategy considers: professional relevance: calculating the semantic similarity between the keywords of the dissertation and the research direction of the experts; review experience: giving priority to expert intelligent agents with more historical reviews; perspective diversity: ensuring that the selected experts have certain differences in research style and institutional background.
[0067] Furthermore, a data upload and parsing node, a structured information extraction node, an expert agent selection node, a parallel review node, and a comprehensive review result output node are constructed to form an intelligent review workflow.
[0068] The data upload and parsing node is used to parse the uploaded papers to be reviewed, obtaining the main text and related metadata. The structured information extraction node is used to extract structured information from the main text. The review expert agent selection node is used to select several suitable review expert agents from the group of review expert agents based on the relevant metadata. The parallel review node is used to conduct parallel reviews by the selected review expert agents based on the review point list and review rules. The comprehensive review result output node is used to summarize the review results of several review expert agents and determine the comprehensive review result of the paper to be reviewed.
[0069] As a feasible implementation method, Figure 4 is a schematic diagram of an intelligent review workflow provided by an embodiment of the present invention. As shown in Figure 4, taking a thesis as an example, the end-to-end multi-agent driven intelligent review workflow is as follows: (1) Data upload and parsing: Input a thesis in PDF format, and obtain the main text M in Markdown format through document parsing methods such as OCR parsing.
[0070] (2) Structured information extraction: The information extraction module extracts the structured information of the main text content M. (3) Selection of review expert intelligent agents: First, identify the discipline to which the paper belongs. and professional Determine the year in which the degree is awarded. Then activate the discipline. Corresponding review committee intelligent agent The committee's intelligent agent selects data from the expert database based on paper abstracts and keywords. The top-K relevant experts are retrieved from the database and denoted as the candidate set. Then, diversity filtering is performed to ultimately select n expert agents (usually...). ), denoted as .
[0071] (4) Parallel evaluation by multiple expert agents: For each selected expert agent The following review process is executed, with each expert agent's review process running in parallel without interference: Key chapters are sampled from the main text M, controlling the input length; the expert knowledge base is dynamically loaded according to the time limit. According to the review point list The general model is called item by item to analyze the problems and generate review results. ;Will Input the evaluation model and predict the expert's evaluation rating. and brief comments .
[0072] (5) Output of comprehensive review results: The expert agent retrieves relevant academic cases, enhances the knowledge of the review results, and generates a detailed review report. ,in To provide supporting documentary evidence for the review conclusions, the review results of all expert agents were collected. The final grade of the paper is determined according to the predefined comprehensive review rules. By integrating the review results from various experts, a comprehensive list of issues for the paper is generated, along with a structured review report, including the overall rating, scores for each review dimension, specific problem descriptions, and improvement suggestions.
[0073] S104. Input the paper to be reviewed into the intelligent review workflow, so as to match a suitable intelligent reviewer agent from the group of intelligent reviewer agents to conduct intelligent review of the paper to be reviewed and obtain the intelligent review results.
[0074] The intelligent review workflow retrieves the papers to be reviewed and parses them in the data upload and parsing node, obtaining the Markdown-formatted main text content M. The structured information extraction node then extracts the structured information from the main text content M. .
[0075] Then, in the expert agent selection node, the paper-related metadata is identified. The corresponding review committee agent is then activated; based on the paper abstract and keywords, the committee agent retrieves Top-K relevant experts from the expert database, which are denoted as the candidate set; then, diversity filtering is performed, and finally, n expert agents are selected.
[0076] In the parallel review nodes, the review process is executed for each selected expert agent. The review processes of each expert agent are executed in parallel without interference. Finally, in the comprehensive review result output node, the review results of all expert agents are collected, and the final grade of the paper is determined according to the review point list and review rules. The review results of all experts are integrated to generate a comprehensive list of issues for the paper, and a structured review report is generated, including the overall grade, scores for each review dimension, specific problem descriptions, and improvement suggestions.
[0077] As a feasible implementation method, the multi-agent driven intelligent paper review system includes the following user interaction parts: (1) Paper review module: provides single paper or batch paper import function; (2) Review result module: provides review list, reason details viewing, report export and other functions.
[0078] Furthermore, this embodiment of the invention also provides a multi-agent driven intelligent paper review system, as shown in Figure 5. The multi-agent driven intelligent paper review system specifically includes: a review big model module, which constructs a review point list based on the basic review dimensions of paper quality review; and constructs a review big model based on historical papers and expert review data; an expert review intelligent agent group module, used to construct a group of review expert intelligent agents based on the background information of review experts and the review big model; and an intelligent review workflow module, used to construct an intelligent review workflow based on the group of review expert intelligent agents, the review point list, and review rules; and inputting the paper to be reviewed into the intelligent review workflow to match suitable review expert intelligent agents from the group of review expert intelligent agents to perform intelligent review of the paper to be reviewed, thereby obtaining intelligent review results.
[0079] Finally, this embodiment of the invention also provides a storage medium, which is a non-volatile computer-readable storage medium storing at least one program. Each program includes instructions, which, when executed by a terminal, cause the terminal to perform the following: constructing a review point list based on the basic review dimensions of paper quality review; constructing a large-scale review model based on historical papers and expert review data; constructing a group of intelligent review experts based on the background information of review experts and the large-scale review model; constructing an intelligent review workflow based on the group of intelligent review experts, the review point list, and review rules; and inputting the paper to be reviewed into the intelligent review workflow to match several suitable intelligent review experts from the group of intelligent review experts to perform intelligent review of the paper to be reviewed, thereby obtaining an intelligent review result.
[0080] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0081] The foregoing has described specific embodiments of the present invention. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0082] The above description is merely an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-agent-driven intelligent paper review method, characterized in that, The method includes: constructing a review point list based on the basic review dimensions of paper quality review; constructing a large-scale review model based on historical papers and expert review data; constructing a group of intelligent review experts based on the background information of review experts and the large-scale review model; constructing an intelligent review workflow based on the group of intelligent review experts, the review point list, and review rules; inputting the paper to be reviewed into the intelligent review workflow, so as to match several suitable intelligent review experts from the group of intelligent review experts to perform intelligent review of the paper to be reviewed, and obtain intelligent review results.
2. The multi-agent driven intelligent paper review method according to claim 1, characterized in that, Based on the fundamental review dimensions of thesis quality review, a review point list is constructed, specifically including: determining the fundamental review dimensions of thesis quality review according to the thesis quality review standards, resulting in a list of primary review dimensions; wherein, the list of primary review dimensions includes, but is not limited to, the following review dimensions: topic selection, theory, frontier, methodology, writing, and innovation; analyzing the detectable review indicators in each review dimension to refine each review dimension into several secondary review indicators, resulting in the review point list; constructing a hybrid detection method for each secondary review indicator; wherein, the hybrid detection method is composed of a combination of rule-based detection methods and model-based detection methods; the rule-based detection method uses preset rules to detect the reasonableness of the results of each secondary review indicator; the model-based detection method uses a large model to detect the reasonableness of the results of each secondary review indicator.
3. The multi-agent driven intelligent paper review method according to claim 1, characterized in that, The process of constructing a large-scale review model based on historical papers and expert review data includes: acquiring historically reviewed papers and corresponding expert review data; collecting background information on the reviewers of the papers; constructing an initial training dataset based on the historical papers, the expert review data, the review point list, and the reviewer background information; performing sample augmentation on the initial training dataset to obtain a data-enhanced training dataset; constructing a large-scale review model based on a large language model; performing supervised training on the large-scale review model using the training dataset; if the trained large-scale review model does not meet the testing requirements, constructing a reward model for the large-scale review model based on the review point list; and performing reinforcement learning on the large-scale review model using the reward model to obtain an optimized large-scale review model.
4. The multi-agent driven intelligent paper review method according to claim 3, characterized in that, Based on the historical papers, expert review data, review point list, and reviewer background information, an initial training dataset is constructed. Specifically, this includes: parsing the historical papers and expert review data to obtain paper text and review data text; automatically reviewing the paper text according to the review indicators in the review point list to obtain a multi-dimensional review result list; extracting structured information from the paper text and review data text; compressing the reviewer background information in a structured manner; and concatenating the compressed reviewer background information, the structured information, and the text corresponding to the multi-dimensional review result list using a prompt to obtain the initial training dataset.
5. The multi-agent driven intelligent paper review method according to claim 3, characterized in that, The initial training dataset is augmented to obtain a data-enhanced training dataset. Specifically, this includes: performing controlled degradation operations on high-quality academic papers to generate negative samples; wherein, the controlled degradation operations include, but are not limited to: content consistency violation, format standardization violation, and academic rigor violation; and merging the negative samples into the initial training dataset to obtain the data-enhanced training dataset.
6. The multi-agent driven intelligent paper review method according to claim 3, characterized in that, The supervised training of the large-scale evaluation model using the training dataset specifically includes: constructing a training objective function based on a multi-task learning framework; using the training dataset as input to the large-scale evaluation model, and performing supervised training on the large-scale evaluation model based on the training objective function; wherein, the supervised training process includes: fine-tuning all or some parameters of the large-scale evaluation model.
7. The multi-agent driven intelligent paper review method according to claim 3, characterized in that, The optimized evaluation model is obtained by performing reinforcement learning on the evaluation model using the reward model. Specifically, this includes: constructing a reward function based on the reward types included in the reward model; constructing a policy optimization model based on the GRPO reinforcement learning framework, and using the parameters of the trained evaluation model as the initial weights of the policy optimization model; constructing an optimization objective function based on the reward function and the initial weights; and performing reinforcement learning on the trained evaluation model based on the optimization objective function to obtain the optimized evaluation model.
8. The multi-agent driven intelligent paper review method according to claim 1, characterized in that, Based on the background information of the reviewers and the review model, a group of intelligent reviewer agents is constructed. Specifically, this includes: collecting academic profile data of each reviewer from preset channels to obtain their background information; wherein the preset channels include, but are not limited to, academic resource databases, research project databases, and institutional websites; constructing an intelligent reviewer agent for each reviewer based on their background information, thus forming the group of intelligent reviewer agents; classifying the group of intelligent reviewer agents according to paper-related metadata, and establishing a virtual paper review committee agent for each category, used to match suitable intelligent reviewer agents to the papers to be reviewed according to a matching strategy; wherein the paper-related metadata includes, but is not limited to, the subject and major; the matching strategy is related to the following factors: professional relevance, review experience, and perspective diversity.
9. The multi-agent driven intelligent paper review method according to claim 1, characterized in that, Based on the group of expert reviewers, the list of review points, and the review rules, an intelligent review workflow is constructed, specifically including: constructing a data upload and parsing node, a structured information extraction node, an expert reviewer selection node, a parallel review node, and a comprehensive review result output node, constituting the intelligent review workflow; wherein, the data upload and parsing node is used to parse the uploaded paper to be reviewed, obtaining the paper's main text and related metadata; the structured information extraction node is used to extract structured information from the paper's main text; the expert reviewer selection node is used to select several suitable expert reviewers from the group of expert reviewers based on the paper's related metadata; the parallel review node is used to conduct parallel reviews using the selected expert reviewers based on the list of review points and the review rules; the comprehensive review result output node is used to summarize the review results of the expert reviewers and determine the comprehensive review result of the paper to be reviewed.
10. A multi-agent driven intelligent paper review system, characterized in that, The system includes: a review big data model module, which constructs a review point list based on the basic review dimensions of paper quality review; and constructs a review big data model based on historical papers and expert review data; an expert review intelligent agent group module, which constructs an expert intelligent agent group based on the background information of review experts and the review big data model; and an intelligent review workflow module, which constructs an intelligent review workflow based on the expert intelligent agent group, the review point list, and review rules; and inputs the paper to be reviewed into the intelligent review workflow to match a suitable expert intelligent agent from the expert intelligent agent group to perform intelligent review of the paper and obtain intelligent review results.
Citation Information
Patent Citations
Paper review system
CN109658057A
Academic institution-oriented paper review auxiliary method and system
CN119088829A
Intelligent expert recommendation method for blind examination of academic papers based on expert portraits
CN119248999A
LLM-based graduation design thesis intelligent evaluation system
CN119558702A
Intelligent review method based on artificial intelligence and big data analysis
CN119579113A