Multi-agent multi-modal review generation with external knowledge
The multi-agent system addresses AI-based review system limitations by integrating pre-processing modules and specialized agents to generate reliable, context-aware reviews with reduced hallucinations and biases, ensuring accurate and efficient digital document evaluations.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE REGENTS OF THE UNIVERSITY OF COLORADO
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
Current AI-based review systems for digital documents suffer from hallucinations, factual inaccuracies, and biases, struggle with multi-modal data, and are limited by outdated training data, leading to unreliable and inefficient evaluations.
A multi-agent system with pre-processing modules and specialized agents that include a general knowledge unit, domain knowledge unit, external knowledge unit, figure module, and novelty module, along with AI models to process and generate reviews by synthesizing assessments from multiple agents, ensuring factual grounding and context-aware feedback.
The system reduces hallucinations and biases, provides accurate, context-aware reviews, and maintains up-to-date evaluations by integrating external knowledge and multi-modal data processing, enhancing reliability and efficiency.
Smart Images

Figure US2025054493_15052026_PF_FP_ABST
Abstract
Description
1 Docket No. 22151.71aMULTI-AGENT MULTI-MODAL REVIEW GENERATION WITHEXTERNAL KNOWLEDGECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to United States Provisional Patent Application Serial No. 63 / 718,281 filed on November 8, 2024, and entitled “MULTI-AGENT MULTI-MODAL SCIENTIFIC REVIEW GENERATION WITH EXTERNAL KNOWLEDGE,'’ which is expressly incorporated herein by reference in its entirety.FIELD
[0002] The present disclosure relates generally to systems, devices, and methods for generating a review of a digital document.BACKGROUND
[0003] Recent advances in artificial intelligence (Al) and natural language processing have accelerated efforts to produce automation systems. For example, Al-based systems for academic review generation have been explored, but concerns regarding their reliability, factual grounding, and potential biases have limited widespread adoption. The emergence of large language models (LLMs) with expanded context windows and multi-modal capabilities — such as processing text, figures, and tabular data — has created new opportunities to address these challenges and improve the accuracy and depth of Al-assisted review systems.
[0004] Despite these advancements, cunent Al systems face persistent challenges in the domain of digital document review. While such systems can improve efficiency, they frequently generate reviews affected by hallucinations, including factual inaccuracies, fabricated or nonexistent references, and misleading statements. These hallucinations, together with high similarity indices that elevate plagiarism risks, limit the trustworthiness of traditional automated review systems. Moreover, existing Al models struggle to produce nuanced critiques and context-aware feedback and often fail to consider multi-modal portions of the digital document.
[0005] Additionally, the exponential expansion of academic research in recent years has generated an overwhelming volume of papers requiring evaluation, exceeding the capacity of traditional, human-driven review processes. This rapid growth in cutting-edge research also challenges existing Al models, which often rely on static or incomplete training data. As a result, such models struggle to incorporate the most current findings, increasing the likelihood of outdated analyses or hallucinations and thereby reducing their reliability in the academic review context.2 Docket No. 22151.71a
[0006] Accordingly, there are a number of challenges that may be addressed.
[0007] The subject matter claimed herein is not limited to embodiments that solve any disadvantages or that operate only in environments such as those described above. Rather, this background is only provided to illustrate one exemplary’ technology area where some aspects described herein may be practiced.BRIEF SUMMARY
[0008] The present disclosure is related to systems, devices, and methods for generating a review of a digital document.
[0009] According to various aspects, for example, a system for generating a review of a digital document is provided. The system may include one or more pre-processing modules including a general knowledge unit configured to store and output one or more general instructions for constructing a review; a domain knowledge unit configured to output and dynamically generate a domain dataset from one or more references in the digital document; an external knowledge unit configured to retrieve one or more candidate prior works from one or more sources that are external to the computer system and to generate and output an external works dataset based on the one or more candidate prior works; a figure module configured to extract multi-modal data from the digital document and to generate and output textual descriptions that describe a consistency and / or clarity score associated with the figures and / or captions; and a novelty module configured to generate search queries used to retrieve the one or more candidate prior works and to process the one or more candidate prior works to generate and output a report indicating whether the digital document satisfies a novelty condition.
[0010] The system may also include one or more processing modules that are each configured to retrieve at least one output of the one or more pre-processing modules, wherein the one or more processing modules comprise a plurality of specialized agents comprising a leader agent and one or more other specialized agents, wherein: the leader agent is configured to transmit instructions to the one or more other specialized agents, to receive and synthesize outputs from the one or more other specialized agents, and to generate the review of the digital document based at least on the synthesized outputs; and the one or more other specialized agents are each configured to process the digital document based on the at least one output of the one or more pre-processing modules and to generate an output to the leader agent indicating an assessment of a respective aspect of the digital document.
[0011] According to other aspects, a method for generating a review of a digital document is provided. The method may include (a) receiving, at a computer system, a digital document comprising multi-modal data including text and one or more figures; (b) extracting non-text data3 Docket No. 22151.71a and any associated captions from the digital document and generating textual descriptions for the non-text data; (c) generating a domain dataset based on one or more references contained in the text of the digital document; (d) generating an external works dataset comprising a set of candidate prior works by generating one or more search queries based on the multi-modal data of the digital document and retrieving one or more candidate prior works from one or more sources external to the computer system using the one or more search queries; (e) processing each candidate prior work in the set of candidate prior works and determining whether the digital document satisfies a novelty condition based on each candidate prior work; (f) if the novelty condition is not satisfied, generating a report identifying one or more candidate prior works that caused the novelty condition to not be satisfied and terminating further processing of the digital document; and (g) if the novelty condition is satisfied, processing, by a plurality of specialized agents, the digital document using at least one of the domain dataset, external works dataset, textual descriptions, wherein processing comprises: (i) transmitting, by a leader agent, instructions to one or more other specialized agents to cause each specialized agent to produce one or more outputs comprising an assessment of an aspect of the digital document; (ii) receiving, by the leader agent, the one or more outputs from the one or more other specialized agents; (iii) synthesizing, by the leader agent, the one or more outputs into a synthesized review; and (iv) generating, by the leader agent, the review of the digital document based at least on the synthesized review.
[0012] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify' key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0013] Additional features and advantages will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the teachings herein. Features and advantages of the present disclosure may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. Features of the present disclosure will become more fully apparent from the following description and appended claims, or may be learned by the practice of the present disclosure as set forth hereinafter.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to describe the manner in which at least some of the advantages and features of the present disclosure may be obtained, a more particular description of aspects of the present disclosure will be rendered by reference to specific aspects thereof which are illustrated in the appended drawings. Understanding that these drawings depict only typical aspects of the present4 Docket No. 22151.71a disclosure and are not therefore to be considered to be limiting of its scope, aspects of the present disclosure will be described and explained with additional specificity and detail through the use of the accompanying drawings.
[0015] FIG. 1 illustrates a schematic representation of a system for generating a review according to some embodiments described herein.
[0016] FIG. 2 illustrates a schematic representation of a figure module according to some embodiments described herein.
[0017] FIG. 3 illustrates a block diagram of a method for generating a review according to some embodiments described herein.
[0018] FIG. 4 illustrates tables showing results of a validation study of one embodiment of the system for generating a review of a digital document described herein.DETAILED DESCRIPTION
[0019] The present disclosure is related to systems, devices, and methods for generating a review of a digital document.
[0020] Illustrated in FIG. 1 is a system for generating a review according to some embodiments described herein. As shown in FIG. 1, the system 100 may include one or more preprocessing modules 102 and one or more processing modules 104. In some embodiments, the system 100 may be configured to generate a review 5 of a received digital document 10 using a plurality of specialized agents (e.g., experiments agent 160, impact agent 170, clarity agent 180, leader agent 190, etc.). The pre-processing modules may be included in a pre-processing of the digital document 10. and the processing modules may be included in a processing of the digital document 10 that may occur after the pre-processing.
[0021] The specialized agents may include rule-based agents, general purpose and / or specially trained large-language-model (LLM) based agents, artificial-intelligence-model-based agents, and / or other computer-implemented processing agents. Each specialized agent may execute instructions stored in memory and processed by one or more processors to perform designated analytical, evaluative, or generative tasks on portions of or the entirety of the digital document 10 or on data derived therefrom. Additionally or alternatively, each specialized agent may be configured to focus on a particular aspect of the digital document 10, such as evaluating experimental methodology, assessing the significance or impact of the digital document, analyzing the clarity and organization of the presentation.
[0022] In some embodiments, the specialized agents may operate independently, in parallel, and / or under the control of one of the specialized agents that may be configured to send instructions to and synthesize outputs from the other specialized agents to generate the review 5.5 Docket No. 22151.71a
[0023] Each specialized agent may communicate through defined interfaces to the leader agent, to other specialized agents, and / or to any of the pre-processing modules 102 and may employ one or more machine-learning or rule-based techniques to generate structured outputs in response to received instructions. The specialized agents may thus include configurable processing components implemented in software, firmware, and / or hardware, which may enable the system 100 to perform review generation and related analyses in a scalable and adaptive manner.
[0024] FIG. 1 shows that the pre-processing modules 102 may include a general knowledge unit 110. a domain knowledge unit 120, an external knowledge unit 130, a figure module 140, and / or a novelty module 150. The pre-processing modules 102 may be configured to cany' out pre-processing of the digital document 10 and may be implemented in software, firmware, and / or hardware by one or more processors.
[0025] In some embodiments, the digital document 10 may include any electronically stored document containing scientific, technical, or scholarly information suitable for automated review. In various embodiments, the digital document 10 may include, for example, an academic paper, manuscript, preprint, journal submission, conference paper, grant proposal, or other research document formatted in a digital medium such as a PDF, DOCX, or XML file. The digital document 10 may contain multi-modal data including text, figures, tables, and references and may be provided to the system 100 through a network interface, storage device, or local upload. In certain embodiments, the system 100 may also process supplementary materials associated with the digital document, such as appendices or datasets.
[0026] According to some embodiments, the general knowledge unit 110 may be configured to store and output one or more instructions for generating the review 5 including standardized evaluation frameworks, terminology definitions, and procedural review guidance applicable across disciplines. In some embodiments, the general knowledge unit 110 may include one or more general instructions that, when retrieved and executed by the specialized agents, improve the performance, accuracy, and consistency of review generation across the specialized agents.
[0027] For example, the general instructions may include predefined prompt templates, reviewer checklists, and / or weighted scoring criteria configured to control how the specialized agents evaluate aspects of the digital document 10 such as a novelty aspect, clarity aspect, and / or an experimental aspect. In some embodiments, the general knowledge unit 110 may further define formatting conventions and structured output schemas such that each specialized agent produces results that the leader agent 190 may synthesize into the review 5. The inclusion of such general instructions in the general knowledge unit 110 may provide for uniform review standards and minimize divergence among specialized agents performing parallel tasks.6 Docket No. 22151.71a
[0028] The domain knowledge unit 120 may include a dynamically generated domain dataset 125 derived from a bibliography of and / or one or more works cited by and / or referenced in the digital document 10, which may enable the specialized agents to contextualize the digital document 10 and trace relationships among prior works specific to the document’s research area. The domain dataset 125 may be generated by processing the digital document 10 to extract data including author names, publication dates, titles, content such as text and / or figures, and / or source identifiers. In some embodiments, one or more of the references may be retrieved in whole or in part from an external database, while in other embodiments the references may be represented within the domain dataset 125 based solely on the extracted data from the digital document 10.
[0029] The extracted data may then be organized into a structured representation — such as a table, index, or other specialized dataset — that records relationships among the cited works and relevant topical descriptors. The domain dataset 125 may thus provide a machine-readable contextual map of the research landscape referenced by the digital document 10, which may enable the specialized agents to perform more accurate assessments of aspects of the digital document 10 such as novelty, significance, and / or relevance during the review process.
[0030] In some embodiments, because the domain dataset 125 may be dynamically generated by a pre-processing unit for each digital document 10 that is received by the system 100, the domain knowledge unit 120 may provide adaptability to the system 100. For example, the ability to generate a fresh domain dataset for each digital document 10 may allow the system 100 to perform a context-specific and / or customized analysis tailored to the content and references of that particular digital document.
[0031] FIG. 1 also shows that the external knowledge unit 130 may include a dynamically generated external works dataset 135. For example, the external works dataset 135 may include external grounding data that may include metadata and content features from one or more candidate prior works obtained during the pre-processing of the digital document 10. The metadata may include author names, publication years, journal or conference identifiers, keywords, abstracts, and citation counts, and the content features may include text, topical descriptors, or summary representations derived from the candidate prior works.
[0032] The external works dataset 135 may be obtained from one or more external sources 137 that are external to the computer system 100. In some embodiments, the external knowledge unit 130 may access one or more academic journal databases, digital libraries, or scholarly repositories that provide programmatic interfaces for retrieving bibliographic records and content. Example sources may include academic publisher databases, conference proceedings archives, preprint servers, or institutional repositories. In other embodiments, the external knowledge unit 130 may obtain external data from one or more general-purpose or domain-specific data7 Docket No. 22151.71a collections, such as technical standards databases, open-access research indices, or other verified external sources containing scientific and / or technical information.
[0033] According to some embodiments, the external data may be acquired using authorized application programming interfaces (APIs), web-based retrieval mechanisms, or other controlled access protocols. The flexibility to gather information from multiple trusted external sources may allow the external knowledge unit 130 to maintain a diverse and up-to-date repository of reference materials, which may improve the grounding and factual reliability7of analyses generated by the specialized agents and may reduce hallucinations of the specialized agents.
[0034] In some embodiments, the external works dataset 135 may further include associations between the candidate prior works and corresponding sections of the digital document 10, such as the title, abstract, or specific figures. Because the external works dataset 135 may be generated for each received digital document 10 based on dynamically retrieved external data, the external knowledge unit 130 may include up-to-date and context-relevant external grounding data.
[0035] According to some embodiments, because the external works dataset 135 may contain data obtained from actual retrieved prior works, hallucinations from the specialized agents may be reduced. For example, the inclusion of verified, externally sourced content w ithin the external works dataset 135 may provide the specialized agents with factual reference material against which generated text or analyses can be compared. Such external grounding data may limit the generation of unsupported or fabricated statements by constraining the agents’ responses to information that is verifiable within the candidate prior works.
[0036] The figure module 140 may be configured to extract, analyze, and / or evaluate multimodal contained within the digital document 10. During pre-processing, the figure module 140 may identify and separate non-text data such as figures, images, or graphical elements from the text of the digital document 10 and may associate each figure with any available captions and / or contextual metadata. The figure module 140 may then generate descriptions of the extracted nontext data and assess each figure for clarity and consistency relative to the title, abstract, and body of the digital document.
[0037] In some embodiments, the figure module 140 may include one or more artificial intelligence models trained to interpret visual content and to produce structured representations that describe a subject matter, quality, and relevance of the non-text data. The figure module 140 may thus generate a structured report including clarity scores, consistency scores, and / or descriptive summaries, which may be provided to other specialized agents for inclusion in the review generation process. The ability of the figure module 140 to transform non-text data into structured representations including textual data may allow7the system 100 to provide a multimodal analysis of the digital document 10.8 Docket No. 22151.71a
[0038] In some embodiments, the one or more artificial intelligence models included in the figure module 140 may be specially trained to enhance their ability to interpret and evaluate figures contained in scientific or technical documents. For example, training data used to train the one or more artificial intelligence models included in the figure model 140 may include pairs of figures and corresponding captions, titles, and textual descriptions drawn from diverse sources such as research articles, conference papers, and publicly available datasets. The training process may expose the model to a variety of figure types, including plots, charts, schematics, and diagrams, along with examples of clear and unclear visual presentations. The one or more artificial intelligence models may be trained to associate visual features with textual descriptions and clarity or consistency indicators, enabling the models to generate structured assessments including textual data of new figures encountered during operation.
[0039] In certain embodiments, additional fine-tuning may be performed using curated examples that link figure quality or interpretability to reviewer feedback or editorial standards. This specialized training may allow the figure module 140 to more accurately convert visual content into meaningful structured data, which may improve the reliability and usefulness of figure-based evaluations during review^ generation.
[0040] In at least such a way, the figure module 140 may be configured to receive as an input multi-modal data and produce an output containing single-modal data such as text, which may reduce system complexity and provide for a more complete review.
[0041] In some embodiments the novelty module 150 may be configured to evaluate the originality of the digital document 10 relative to existing prior works. During pre-processing, the novelty module 150 may generate one or more search queries based on the title, abstract, or other descriptive portions of the digital document 10, and may execute the queries or cause the queries to be executed across one or more external databases or repositories to obtain a set of candidate prior works.
[0042] In some embodiments, the set of candidate prior works may include any number of candidate prior w orks, for example in a range from one to one hundred. The higher the number of candidate prior works, the greater a processing load may be on the system 100. Accordingly, it may be desirable to balance the load on the system 100 with providing sufficient amount of candidate prior works so that the novelty module 150 may adequately assess a novelty of the digital document 10. For example, in some embodiments the set of candidate works includes thirtycandidate prior works.
[0043] According to some embodiments, the one or more search queries may each target a different range of scope. For example, a first search query may include narrowly defined terms based directly from the title and / or abstract of the digital document to identify closely related prior9 Docket No. 22151.71a works, and a second search query may target broader contextual or thematic terms configured to search adjacent or emerging fields. In some embodiments, a third search query may include general terminology or synonyms associated with the digital document to identify indirectly related works that may still be relevant for novelty comparison.
[0044] In some embodiments, the novelty module 150 may operate in coordination with the external knowledge unit 130. For example, the novelty module 1 0 may transmit the one or more search queries to the external knowledge unit 130, which may perform retrieval operations from external databases and populate the external works dataset 135 with metadata and content features of the retrieved prior works. The novelty module 150 may then access the populated external works dataset 135 to compute similarity or novelty measures between the digital document and the retrieved w orks. This may enable the system 100 to perform novelty assessments using verified, externally sourced information, which may reduce hallucinations within the specialized agents.
[0045] According to some embodiments, the novelty module 150 may communicate with the domain knowledge unit 120 to avoid misclassifying the digital document 10 as non-novel based on w orks that are already cited by the digital document 10. During pre-processing, the novelty module 150 may reference the domain dataset 125 generated by the domain knowledge unit 120 to identify the set of cited references extracted from the digital document 10. Before assigning a novelty condition to the digital document based on the set of candidate prior works, the novelty module 150 may compare the respective candidate prior work against the cited references contained in the domain dataset 125. If a match is detected, the novelty module 150 may exclude the corresponding work from the novelty determination, adjust the associated novelty score accordingly, and / or cause the respective candidate prior work to be removed from the set of candidate prior works. This may prevent the novelty module 150 from incorrectly assigning a novelty condition that indicates that the digital document 10 is not novel.
[0046] Additionally or alternatively, the novelty module 150 may be configured to assess the relevance of each candidate prior work in the set of candidate works. In some embodiments, the novelty module 150 may be configured to analyze the semantic similarity between the title, abstract, or keywords of the digital document and those of each candidate prior work to determine whether the respective candidate work is relevant to the digital document. The relevance assessment may employ one or more large language models or other artificial intelligence algorithms trained to identify whether the respective candidate prior work addresses the same or a sufficiently related subject matter as the digital document. For example, the novelty module 150 may assign a relevance score on a normalized scale between 0.00 and 1.00, where a higher score (e.g., a score above 0.70) may indicate that the candidate prior w ork discusses similar concepts or10 Docket No. 22151.71a methods and is thus relevant. In some embodiments, the novelty module 150 may generate a textual output that indicates a relevance of the candidate prior work to the digital document.
[0047] In some embodiments, candidate prior works with low relevance may be filtered out or removed from the set of candidate prior works, which may reduce computational load and minimize the likelihood of erroneous similarity detections.
[0048] In some embodiments, the novelty module 150 may employ one or more artificial intelligence models trained to compare textual, graphical, or topical features between two documents to determine a degree of novelty and to generate a textual output that indicates whether the digital document is novel in relation to a respective candidate prior work. For example, the novelty module 150 may include an artificial intelligence model trained to compare the similarities and / or differences between two textual passages and generate a textual output that indicates whether and how one of the textual passages is novel in relation to the other based on the similarities and / or difference. The novelty module 150 may thus be configured to receive as an input the digital document 10 and the respective candidate prior work and to output a report, (e.g., a textual report) indicating whether the digital document 10 is novel in relation to the respective candidate prior work.
[0049] In some embodiments, the novelty module 150 may generate a pairwise novelty score based on each candidate prior work in the set of candidate prior works and may generate a report summarizing whether the digital document is considered novel or non-novel with respect to the retrieved literature. In certain embodiments, if any pairwise novelty score fails to satisfy a predefined novelty condition, or if the textual output indicates that the digital document 10 is not novel in relation to the respective candidate work, the novelty module 150 may produce a structured output containing textual data that identifies the related prior work and may halt further processing for the digital document 10. The novelty module 150 may allow the system 100 to detect redundant or derivative submissions at an early stage, which may improve computational efficiency and the overall reliability of generating the review 5.
[0050] According to some embodiments, the predefined novelty condition may be based on one or more similarity thresholds that quantify the conceptual or structural overlap between the digital document and each candidate prior work. In some embodiments, textual similarity may be measured using cosine similarity between embedded vector representations of the documents, and the digital document may be deemed novel when the cosine similarity does not exceed, for example, a cosine similarity of 0.85.
[0051] As another example, if more than 60 percent of extracted technical terms or phrases overlap between the digital document and a candidate prior work, the novelty module 150 may classify the document as non-novel with respect to that prior work and thus that the novelty11 Docket No. 22151.71a condition is not satisfied. The thresholds may be adjusted dynamically based on field-specific factors such as publication density and / or common terminology. The novelty module 150 may further apply weighting factors — for example, assigning 0.6 weight to textual features, 0.3 weight to figure similarity, and 0.1 weight to citation overlap — to compute a composite novelty score. These quantitative parameters may be refined through historical benchmarking and / or user feedback.
[0052] FIG. 1 also shows that the one or more processing modules 104 may include a plurality of specialized agents, for example an experiments agent 160, an impact agent 170, a clarity agent 180, and a leader agent 190. In some embodiments, the leader agent 190 may be configured to send instructions to and to receive one or more outputs from the other specialized agents. The leader agent 190 may also be configured to facilitate communication between the specialized agents.
[0053] In some embodiments, the leader agent 190 may be configured to prompt the other specialized agents to analyze the received digital document 10 based on the respective agent’s specialty. For example, the digital document 10 may be received at the system 100. As discussed above, the system 100 may be configured to perform pre-processing using the one or more preprocessing modules 102 to prepare the digital document 10 for processing and analysis. After the pre-processing modules 102 have finished the pre-processing steps, and, in some embodiments, only if the novelty module 150 determined that the digital document 10 is sufficiently novel, the leader agent 190 may be configured to begin generating the review 5 using the one or more other specialized agents.
[0054] In some embodiments, the leader agent 190 may be pre-configured with instructions that configure the leader agent 190 to use one or more other agents to generate the review 5. The leader agent 190 may begin generating the review by sending instructions to each specialized agent indicating that the specialized agent is to begin its task.
[0055] Each specialized agent may be specialized in regard to a particular task and / or aspect of the digital document. For example, the experiments agent 160 may be configured to evaluate an experimental aspect such as experimental design, dataset selection, and / or methodological rigor within the digital document, the impact agent 170 may be configured to assess a novelty aspect such as the novelty, significance, and / or potential influence of the digital document relative to existing work, and the clarity agent 180 may be configured to evaluate a clarity aspect such as the organization, readability, and presentation quality of the digital document, including the adequacy of figures, tables, and captions.
[0056] In some embodiments, each specialized agent may retrieve and execute general instructions defined by the general knowledge unit 110, which may enable the agents to maintain12 Docket No. 22151.71a consistency in analysis and structured output generation while focusing on different aspects of the digital document. Additionally or alternatively, each specialized agent may include a large language model configured to produce a text output based on one or more text inputs.
[0057] A specialization of each specialized agent may be provided for in a number of ways. In some embodiments where the specialized agents each include instances of large language models (LLMs), specialized training sets may be employed to improve the respective model’s performance on a given task. For example, the experiments agent 160 may include an instance of an LLM that was specially trained on data that may include experimental design descriptions, methodology sections, and benchmark analyses drawn from peer-reviewed scientific papers and repositories of reproducibility studies.
[0058] Such data may include information about datasets, model architectures, evaluation metrics, ablation studies, and statistical analyses. Training the experiments agent 160 using such training data may enable the experiments agent 160 to better identify weaknesses in experimental design, assess the adequacy of baselines, and evaluate whether reported results are statistically and methodologically sound. In some embodiments, the training data may further include synthetic examples of good and poor experimental practices so that the model can classify and critique methodological rigor with greater precision.
[0059] As another example, the impact agent 170 may be trained on data that consists of abstracts, introductions, and discussion sections from published scientific papers, as well as peerreview commentary and citation-network data reflecting how research outputs influence subsequent work. Such training data may include examples of highly cited breakthrough papers and lower-impact publications, together with metadata describing citation counts, topical novelty, and cross-disciplinary relevance. By being trained on such training data, the impact agent 160 may better assess the significance, originality, and potential influence of the received digital document 10 within its research domain. In some embodiments, the data may further include reviewer feedback or acceptance-decision rationales from academic conferences to enhance the model's ability to align its assessments with real-world measures of impact.
[0060] As still another example, the clarity agent 180 may be trained on data comprising the full text of scientific manuscripts and editorial feedback focused on language quality, structure, and readability. Such data may include annotated examples highlighting issues of ambiguity, poor organization, excessive jargon, or unclear figure descriptions, as well as corresponding revisions that improved clarity and coherence. The training corpus may further include style guides, publication standards, and accepted peer-review comments addressing presentation and writing quality. Exposure to these materials may enable the clarify agent 180 to identify' passages that are confusing, poorly structured, or inconsistent with standard scientific writing conventions, and to13 Docket No. 22151.71a provide constructive recommendations for improving overall readability and presentation of the paper.
[0061] In some embodiments, each specialized agent may be specialized through customized prompts configured to instruct the agent to perform tasks or evaluations in accordance with predefined review objectives and evaluation criteria. These specialized prompts may encode detailed role instructions describing the agent’s purpose, tone, and scope of analysis, such as directing the experiments agent 160 to examine dataset adequacy and experimental controls, or instructing the clarity agent 170 to assess organization, grammar, and figure readability. The prompts may also include constraints on response format, such as requiring the specialized agent to produce structured outputs and / or confidence scores for each observation. In some embodiments, the specialized prompts may further specify collaboration protocols which may define how the leader agent 190 may delegate subtasks, request clarifications, and / or synthesizes data from the specialized agents.
[0062] Each specialized agent may be configured to selectively access the general knowledge unit 110, domain knowledge unit 120, and / or external knowledge unit 130 of the pre-processing modules 102 to support the respective specialized agent’s assigned review functions. Each agent may retrieve or query' information as required by the respective specialized agent’s role.
[0063] For example, the impact agent 170 may consult the domain dataset 125 in the domain knowledge unit 120 to assess whether a contribution found in the digital document 10 extends or diverges from prior work, and the experiments agent 160 may reference the external knowledge unit 130 to locate comparable experimental methods. The clarity agent 170 may rely on the general knowledge unit 110 to apply uniform organization and readability criteria across generated reviews and / or may retrieve the output of the figure module 140 to assess a quality of any nontext data in the digital document 10. Access and query' parameters used by7the one or more specialized agents may be provided by the leader agent 190 to ensure that each specialized agent retrieves only the data relevant to its designated function, which may improve efficiency, factual grounding, and consistency in the overall review generation process.
[0064] Each of the one or more specialized agents may be configured to output to the leader agent 190 a result of the processing performed by the respective specialized agent. For example, the experiments agent 160 may output to the leader agent 190 a block of data (e g., textual data) that includes an assessment of any' experimental designs, methods, or other details in the digital document 10. The impact agent 170 may output to the leader agent 190 a block of data (e.g., textual data) that includes an assessment of the significance, originality, and / or potential influence the digital document 10 may have. The clarity agent 190 may output to the leader agent 190 a14 Docket No. 22151.71a block of data (e.g., textual data) that includes an assessment of a clarity of the digital document 10.
[0065] In some embodiments, the leader agent 190 may be configured to receive the blocks of data from the other specialized agents. The leader agent 190 may be configured to synthesize the blocks of data and generate the review 5 based at least on the blocks of data. For example, each specialized agent may send blocks of text data to the leader agent 190. The leader agent 190 may include a large language model configured to synthesize the received blocks of text data into a cohesive output. In some embodiments, this may allow the leader agent 190 to use the one or more specialized agents to generate the review 5.
[0066] According to some embodiments, the leader agent 190 may be further configured to cause the review 5 to be displayed on a user interface to a user of the computer system 100. This may beneficially allow the user to interact with the review, assess its strengths and / or weaknesses, and in some embodiments provide feedback to the computer system 100 to improve future review generation.
[0067] Illustrated in FIG. 2 is a schematic representation of a figure module according to some embodiments described herein. The figure module 200 may be similar to or substantially the same as the figure module 140 described above. FIG. 2 shows that the figure module 200 may receive as an input a digital document 210 that includes identifying features 220 that may include textual data such as a title and / or an abstract and multi-modal data 230 such as figures, tables, charts, and associated captions. In some embodiments, the figure module 200 may be configured to extract the multi-modal data 230 from the digital document 210.
[0068] For example, the figure module 200 may include a processing module 240 configured to convert any extracted non-text data (e.g., figures, charts, graphs, tables, etc.) into text format, for example using an artificial intelligence model. For example, the processing module 240 may include a specially trained artificial intelligence model configured to process the non-text data into text-based descriptions of the non-text data.
[0069] The processing module 240 may also be configured to process textual data of the identify ing features 220 with the generated text-based descriptions and associated captions from the multi-modal data 230 to produce one or more assessments 250. For example, the processing module 240 may output a consistency assessment 252 that determines how closely each figure corresponds to the identifying features 220 of the digital document 210. Additionally or alternatively, the processing module 240 may output a clarity assessment 254 that evaluates readability, labeling accuracy, and overall visual qualify of the non-text data based on the generated text-based descriptions.15 Docket No. 22151.71a
[0070] In some embodiments, the processing module 240 may generate a descriptive summary 256 for each extracted figure based on the generated text-based descriptions, describing the respective figure’s content, variables, and / or any depicted relationships. The one or more assessments 250 may be synthesized into a unified representation by the processing module 240. For example, the processing module 240 may be configured to output a block of textual data that includes each assessment.
[0071] In embodiments that include one or more figures as non-text data, the processing module 240 may be configured to process each figure independently by providing the one or more assessments 250 for each figure. The one or more assessments 250 for each figure may then bysynthesized into a unified representation that includes each assessment for each figure.
[0072] According to some embodiments, this representation may then be provided as an output to one or more specialized agents 260, which may use the one or more assessments as contextual input during review generation. In some embodiments, the figure module 200 maystore each generated output for access by the one or more specialized agents. In some embodiments, the one or more specialized agents may include the experiments agent 160, impact agent 170, and clarity- agent 180 as described above with respect to FIG. 1.
[0073] The figure module 200 may provide for a more thorough review of a digital document by allowing a computer system that includes such a figure module 200 to process multi-modal content included in the digital document. The inclusion of multi-modal data may reduce information loss between modalities and may enable the computer system to detect inconsistencies or presentation issues that would otherw ise go unnoticed by text-only processing. As a result, a computer system that includes the figure module 200 may achieve improved processing efficiency, enhanced contextual understanding, and more reliable generation of reviews that accurately reflect both the textual and non-textual components of the digital document.
[0074] Illustrated in FIG. 3 is a block diagram of a method for generating a review according to some embodiments described herein. FIG. 3 shows that the method 300 may include step 310 of receiving a digital document. For example, step 310 may include receiving, at a computer system, a digital document including multi-modal data including text and one or more figures.
[0075] FIG. 3 also shows that the method 300 may include step 320 which may include extracting non-text data. For example, step 320 may include extracting non-text data and any- associated captions from the digital document and generating textual descriptions for the non-text data.
[0076] FIG. 3 additionally shows that the method 300 may include step 330 which may include generating a domain dataset. For example, step 330 may include generating a domain dataset based on one or more references contained in the text of the digital document.16 Docket No. 22151.71a
[0077] FIG. 3 also shows that the method 300 may include step 340 which may include generating an external works dataset. For example, step 340 may include generating an external works dataset including a set of candidate prior works by generating one or more search queries based on the multi-modal data of the digital document and retrieving one or more candidate prior works from one or more sources external to the computer system using the one or more search queries.
[0078] FIG. 3 additionally shows that the method 300 may include step 350 which may include determining whether the digital document satisfies a novelty' condition. For example, step 350 may include processing each candidate prior work in the set of candidate prior works and determining whether the digital document satisfies a novelty condition based on each candidate prior work.
[0079] FIG. 3 also shows that the method 300 may include step 355 of generating a response and ceasing operations. For example, step 355 may include, if the novelty condition is not satisfied, generating a report identifying one or more candidate prior works that caused the novelty condition to not be satisfied and terminating further processing of the digital document.
[0080] FIG. 3 additionally shows that the method 300 may include step 355 of proceeding to the next candidate in the set. For example, step 355 may be an intermediary step of step 340 of processing each candidate prior work in the set of candidate prior works. Because the method 300 may include determining the novelty condition for each candidate prior work in the set of candidate prior works, unnecessary' processing may be avoided, which may improve the efficiency of the computer system.
[0081] FIG. 3 also shows that the method 300 may include step 360 which may include processing the digital document. For example, step 360 may include, if the novelty condition is satisfied, processing, by a plurality of specialized agents, the digital document using at least one of the domain dataset, external works dataset, textual descriptions. Processing the digital document may include transmitting, by a leader agent, instructions to one or more other specialized agents to cause each specialized agent to produce one or more outputs comprising an assessment of an aspect of the digital document.
[0082] In some embodiments, processing the digital document may further include receiving, by the leader agent, the one or more outputs from the one or more other specialized agents, synthesizing, by the leader agent, the one or more outputs into a synthesized review, and generating, by the leader agent, the review of the digital document based at least on the synthesized review.
[0083] Illustrated in FIG. 4 are tables showing results of a validation study of one embodiment of the system for generating a review of a digital document described herein. In the validation17 Docket No. 22151.71a study, the inventors evaluated the performance of the embodiment against reviews authored by human reviewers and by existing Al review systems. The evaluation set included thirty’ academic papers for which at least some human-generated peer reviews were publicly available. Of these, twenty papers had corresponding human reviews. Ten of the papers were drawn from the ACL 2017 portion of the PeerRead dataset developed by Kang et al., and ten were taken from proceedings of NeurlPS 2019. For each paper with multiple human reviews, a single human review was randomly selected and treated as the human benchmark for comparison.
[0084] In the validation study, an Elo sy stem was adapted in order to determine comparative performance among the various review-generation systems and human reviewers. This adaptation of the Elo methodology allowed the inventors to quantify differences in review qualify using a dynamic, self-correcting scoring system rather than absolute scoring metrics. Because Elo updates incorporate both the expected and actual outcomes of each comparison, the framework naturally weights surprising or decisive wins more heavily, producing a stable but responsive measure of performance across multiple rounds of evaluation.
[0085] In conducting the validation study, 14 participants were instructed to evaluate pairs of reviews generated for the same digital document in order to provide comparative feedback across review sy stems. A total of 140 judgments were produced. Each participant was presented with the full text of a selected academic paper displayed alongside two anonymized reviews. After reading the paper and both reviews, the participant judged the review on for four criteria: Technical Qualify, Constructiveness, Clarify, and Overall Quality. The participant then selected whether either review was superior, if they tied, or if both were poor for each of the four criteria. Both of the reviews' Elo rating was then adjusted according to the judgments.
[0086] The Combined Score in Table 1 indicates an Elo where the Technical Qualify, Constructiveness, and Clarity are weighted at 50% and the overall qualify is rated at 50%. The Style-Adjusted in Table 1 score is an Elo rating after removing the impact of style differences between reviewers using a Bradley -Terry model.
[0087] As shown in Table 1, the embodiment ended the validation study with a higher Elo rating in every criteria than every other system tested. In particular, the embodiment was judged as superior 88% of the time in the pairwise comparisons to the other models and human-produced review.
[0088] Table 2 shows an estimated win-rate for each tested system based on the results of the validation study. As shown in Table 2, based on the final Elo of each system, the embodiment is predicted to be judged as superior to every other model tested a majority of the time, with a worstcase predicted chance of superiority being 59% over D’Arcy et al. Notably, Table 2 shows that the embodiment is predicted to be superior to a human-produced review 93% of the time.18 Docket No. 22151.71a
[0089] While certain embodiments of the present disclosure have been described in detail, with reference to specific configurations, parameters, components, elements, etcetera, the descriptions are illustrative and are not to be construed as limiting the scope of the claimed invention.
[0090] Furthermore, it should be understood that for any given element or component of a described embodiment, any of the possible alternatives listed for that element or component may generally be used individually or in combination with one another, unless implicitly or explicitly stated otherwise.
[0091] In addition, unless otherwise indicated, numbers expressing quantities, constituents, distances, or other measurements used in the specification and claims are to be understood as optionally being modified by the term “about’’ or its synonyms. When the terms “about,” “approximately,” “substantially,” or the like are used in conjunction with a stated amount, value, or condition, it may be taken to mean an amount, value or condition that deviates by less than 20%, less than 10%, less than 5%, less than 1%, less than 0.1%, or less than 0.01% of the stated amount, value, or condition. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0092] Any headings and subheadings used herein are for organizational purposes only and are not meant to be used to limit the scope of the description or the claims.
[0093] It will also be noted that, as used in this specification and the appended claims, the singular forms “a,” “an” and “the” do not exclude plural referents unless the context clearly dictates otherwise. Thus, for example, an embodiment referencing a singular referent (e.g., “widget”) may also include two or more such referents.
[0094] Embodiments described herein may also include properties and / or features (e.g., components, members, elements, modules, parts, and / or portions) described in one or more separate embodiments and are not necessarily limited strictly to the features expressly described for that particular embodiment. Accordingly, the various features of a given embodiment can be combined with and / or incorporated into other embodiments of the present disclosure. Thus, disclosure of certain features relative to a specific embodiment of the present disclosure should not be construed as limiting application or inclusion of said features to the specific embodiment. Rather, it will be appreciated that other embodiments can also include such features.
[0095] The disclosed embodiments may comprise or utilize a special-purpose or general-purpose computer including computer hardware, such as, for example, one or more and system memory units. Embodiments may also include physical and other non-transitory computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-19 Docket No. 22151.71a readable media can be any available media that can be accessed by a general-purpose or specialpurpose computer system. Computer-readable media that store computer-executable instructions in the form of data are “physical computer storage media” or a “hardware storage device.” Furthermore, computer-readable storage media, which includes physical computer storage media and hardware storage devices, exclude signals, carrier waves, and propagating signals. On the other hand, computer-readable media that cany computer-executable instructions are “transmission media” and include signals, carrier waves, and propagating signals. Thus, by way of example and not limitation, the current embodiments can comprise at least two distinctly different kinds of computer-readable media: computer storage media and transmission media.
[0096] Computer storage media (aka “hardware storage device”) are computer-readable hardware storage devices, such as RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSD”) that are based on RAM, Flash memory, phase-change memory (“PCM”), or other ty pes of memory', or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions, data, or data structures and that can be accessed by a general-purpose or specialpurpose computer.
[0097] A “network,” is defined as one or more data links and / or data switches that enable the transport of electronic data between computer systems, modules, and / or other electronic devices. When information is transferred, or provided, over a network (either hardwired, wireless, or a combination of hardwired and wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media include a network that can be used to carry data or desired program code means in the form of computer-executable instructions or in the form of data structures. Further, these computer-executable instructions can be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0098] Upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission mediate computer storage media (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a network interface card or “NIC”) and then eventually transferred to computer system RAM and / or to less volatile computer storage media at a computer system. Thus, it should be understood that computer storage media can be included in computer system components that also (or even primarily) utilize transmission media.
[0099] Computer-executable (or computer-interpretable) instructions comprise, for example, instructions that cause a general-purpose computer, special-purpose computer, or special-purpose20 Docket No. 22151.71a processing device to perform a certain function or group of functions. The computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0100] Those skilled in the art will appreciate that the embodiments may be practiced in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, pagers, routers, switches, and the like. The embodiments may also be practiced in distributed system environments where local and remote computer systems that are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network each perform tasks (e.g. cloud computing, cloud services and the like). In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0101] The present invention may be embodied in other specific forms without departing from its characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is, therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
[0102] Example Implementations
[0103] In view of the foregoing, the present disclosure relates, for example and without being limited thereto, to the following implementations.
[0104] Implementation 1. A computer system for generating a review of a digital document, comprising: one or more pre-processing modules comprising: a general knowledge unit configured to store and output one or more general instructions for constructing a review; a domain knowledge unit configured to output and dynamically generate a domain dataset from one or more references in the digital document; an external knowledge unit configured to retrieve one or more candidate prior works from one or more sources that are external to the computer system and to generate and output an external works dataset based on the one or more candidate prior works; a figure module configured to extract multi-modal data from the digital document and to generate and output textual descriptions that describe a consistency and / or clarity score associated with the figures and / or captions; and a novelty module configured to generate search queries used to21 Docket No. 22151.71a retrieve the one or more candidate prior works and to process the one or more candidate prior works to generate and output a report indicating whether the digital document satisfies a novelty condition; and one or more processing modules that are each configured to retrieve at least one output of the one or more pre-processing modules, wherein the one or more processing modules comprise a plurality of specialized agents comprising leader agent and one or more other specialized agents, wherein: the leader agent is configured to transmit instructions to the one or more other specialized agents, to receive and synthesize outputs from the one or more other specialized agents, and to generate the review of the digital document based at least on the synthesized outputs; and the one or more other specialized agents are each configured to process the digital document based on the at least one output of the one or more pre-processing modules and to generate an output to the leader agent indicating an assessment of a respective aspect of the digital document.
[0105] Implementation 2. The computer system of any one or a combination of implementations 1 and / or 3-15, wherein the plurality of specialized agents comprises an experiments agent configured to process the digital document based at least in part on the general instructions of the general knowledge unit and to generate an output to the leader agent that includes an assessment of an experimental design, dataset selection, and / or methodological rigor of the digital document.
[0106] Implementation 3. The computer system of any one or a combination of implementations 1-2 and / or 4-15, wherein the plurality of specialized agents comprises an impact agent configured to process the digital document based at least in part on the general instructions of the general knowledge unit and to generate an output to the leader agent that includes an assessment of the novelty, significance, and / or potential influence of the digital document relative to each candidate prior work in the set of candidate prior works.
[0107] Implementation 4. The computer system of any one or a combination of implementations 1-3 and / or 5-15, wherein the plurality of specialized agents comprises a clarity agent configured to process the digital document based at least in part on the general instructions of the general knowledge unit and to generate an output to the leader agent that includes an assessment of an organization, readability, and / or presentation quality7of the digital document, including figures, tables, and captions.
[0108] Implementation 5. The computer system of any one or a combination of implementations 1-4 and / or 6-15, wherein the clarity agent is further configured to process the digital document based at least in part on an output of the figure module.
[0109] Implementation 6. The computer system of any one or a combination of implementations 1-5 and / or 7-15, wherein the external knowledge unit is further configured to22 Docket No. 22151.71a access one or more academic journal databases, conference archives, or preprint repositories through application programming interfaces (APIs) or other network-based retrieval mechanisms.
[0110] Implementation 7. The computer system of any one or a combination of implementations 1-6 and / or 8-15, wherein the novelty module is further configured to remove from the set of candidate prior works a respective candidate work that is cited in the digital document.
[0111] Implementation 8. The computer system of any one or a combination of implementations 1-7 and / or 9-15, wherein the novelty module is further configured to determine if a respective candidate work in the set of candidate works is relevant and, if the respective candidate work is not, is further configured to remove from the set of candidate prior works the respective candidate work.
[0112] Implementation 9. The computer system of any one or a combination of implementations 1-8 and / or 10-15, wherein the multi-modal data comprises one or more figures and any associated captions.
[0113] Implementation 10. The computer system of any one or a combination of implementations 1-9 and / or 11-15, wherein each of the plurality of specialized agents is configured to access at least one of the general knowledge unit, domain knowledge unit, or external knowledge unit based on the respective agent's assigned task.
[0114] Implementation 11. The computer system of any one or a combination of implementations 1-10 and / or 12-15, wherein the leader agent is further configured to display the review and the digital document within a user interface.
[0115] Implementation 12. The computer system of any one or a combination of implementations 1-11 and / or 13-15, wherein the general instructions comprise a prompt template.
[0116] Implementation 13. The computer system of any one or a combination of implementations 1-12 and / or 14-15, wherein each specialized agent is configured to generate and output a block of text data comprising an assessment of an aspect of the digital document to the leader agent.
[0117] Implementation 14. The computer system of any one or a combination of implementations 1-13 and / or 15, wherein the leader agent is configured to synthesize the block of text data from each of specialized agent into the review of the digital document.
[0118] Implementation 15. The computer system of any one or a combination of implementations 1-14, wherein the leader agent is configured to receive the block of text data from each specialized agent in parallel.
[0119] Implementation 16. A method for generating a review of a digital document, comprising: (a) receiving, at a computer system, a digital document comprising multi-modal data23 Docket No. 22151.71a including text and one or more figures; (b) extracting non-text data and any associated captions from the digital document and generating textual descriptions for the non-text data; (c) generating a domain dataset based on one or more references contained in the text of the digital document; (d) generating an external works dataset comprising a set of candidate prior works by generating one or more search queries based on the multi-modal data of the digital document and retrieving one or more candidate prior works from one or more sources external to the computer system using the one or more search queries; (e) processing each candidate prior work in the set of candidate prior works and determining whether the digital document satisfies a novelty condition based on each candidate prior work; (f) if the novelty’ condition is not satisfied, generating a report identifying one or more candidate prior works that caused the novelty condition to not be satisfied and terminating further processing of the digital document; and (g) if the novelty condition is satisfied, processing, by a plurality' of specialized agents, the digital document using at least one of the domain dataset, external works dataset, textual descriptions, wherein processing comprises: (i) transmitting, by a leader agent, instructions to one or more other specialized agents to cause each specialized agent to produce one or more outputs comprising an assessment of an aspect of the digital document; (ii) receiving, by the leader agent, the one or more outputs from the one or more other specialized agents; (iii) synthesizing, by the leader agent, the one or more outputs into a synthesized review; and (iv) generating, by the leader agent, the review of the digital document based at least on the synthesized review.
[0120] Implementation 17. The method according to any one or a combination of implementations 16 and / or 18-19, wherein the method further comprises removing from the set any candidate prior work that is already cited in the digital document.
[0121] Implementation 18. The method according to any one or a combination of implementations 1 and / or 19, wherein the plurality of specialized agents further comprises an impact agent, an experiment agent, and / or a clarify agent.
[0122] Implementation 19. A non-transitory computer-readable medium storing instructions that are configured to be executed by one or more processors to cause the one or more processors to perform method of any one or a combination of implementations 16-18.
[0123] Implementation 20. A computer system for generating a review of a digital document, comprising: one or more pre-processing modules comprising: a general knowledge unit configured to store and output one or more general instructions for constructing a review; a domain knowledge unit configured to output and dynamically generate a domain dataset from one or more references in the digital document; an external knowledge unit configured to retrieve one or more candidate prior works from one or more sources that are external to the computer system and to generate and output an external works dataset based on the one or more candidate prior works; a24 Docket No. 22151.71a figure module configured to extract multi-modal data from the digital document and to generate and output textual descriptions that describe a consistency and / or clarity score associated with the figures and / or captions; and a novelty module configured to generate search queries used to retrieve the one or more candidate prior works and to process the one or more candidate prior works to generate and output a report indicating whether the digital document satisfies a novelty condition; and one or more processing modules that are each configured to retrieve at least one output of the one or more pre-processing modules, wherein the one or more processing modules comprise a plurality of specialized agents comprising a leader agent and one or more other specialized agents, wherein: the leader agent is configured to transmit instructions to the one or more other specialized agents, to receive and synthesize outputs from the one or more other specialized agents, and to generate the review of the digital document based at least on the synthesized outputs; and the one or more other specialized agents are each configured to process the digital document based on the at least one output of the one or more pre-processing modules and to generate an output to the leader agent indicating an assessment of a respective aspect of the digital document, and wherein each of the one or more processing modules comprises a large language model configured to produce a text output based on one or more text inputs.
[0124] The present disclosure may be embodied in other specific forms without departing from its spirit or characteristics. The described aspects are to be considered in all respects only as illustrative and not restrictive. The scope of the disclosure is, therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
25 Docket No. 22151.71aCLAIMSWhat is claimed is:
1. A computer system for generating a review of a digital document, comprising: one or more pre-processing modules comprising: a general knowledge unit configured to store and output one or more general instructions for constructing a review; a domain knowledge unit configured to output and dynamically generate a domain dataset from one or more references in the digital document; an external knowledge unit configured to generate a set of candidate prior works by retrieving one or more candidate prior works from one or more sources that are external to the computer system and to generate and output an external works dataset based on the one or more candidate prior works; a figure module configured to extract multi-modal data from the digital document and to generate and output textual descriptions that describe a consistency and / or clarity score associated with the figures and / or captions; and a novelty module configured to generate search queries configured to retrieve the one or more candidate prior works and to process the one or more candidate prior works to generate and output a report indicating whether the digital document satisfies a novelty condition; and one or more processing modules that are each configured to retrieve at least one output of the one or more pre-processing modules, wherein the one or more processing modules comprise a plurality of specialized agents comprising leader agent and one or more other specialized agents, wherein: the leader agent is configured to transmit instructions to the one or more other specialized agents, to receive and synthesize outputs from the one or more other specialized agents, and to generate the review of the digital document based at least on the synthesized outputs; and the one or more other specialized agents are each configured to process the digital document based on the at least one output of the one or more pre-processing modules and to generate an output to the leader agent indicating an assessment of a respective aspect of the digital document.
2. The computer system of claim 1, wherein the plurality of specialized agents comprises an experiments agent configured to process the digital document based at least in part on the general26 Docket No. 22151.71a instructions of the general knowledge unit and to generate an output to the leader agent that includes an assessment of an experimental aspect of the digital document.
3. The computer system of claim 1. wherein the plurality of specialized agents comprises an impact agent configured to process the digital document based at least in part on the general instructions of the general knowledge unit and to generate an output to the leader agent that includes an assessment of a novelty aspect of the digital document.
4. The computer system of claim 1, wherein the plurality of specialized agents comprises a clarity agent configured to process the digital document based at least in part on the general instructions of the general knowledge unit and to generate an output to the leader agent that includes an assessment of a clarity aspect of the digital document.
5. The computer system of claim 4, wherein the clarity agent is further configured to process the digital document based at least in part on an output of the figure module.
6. The computer system of claim 1, wherein the external knowledge unit is further configured to access one or more academic journal databases, conference archives, or preprint repositories through application programming interfaces (APIs) or other network-based retrieval mechanisms.
7. The computer system of claim 1, wherein the novelty module is further configured to remove from the set of candidate prior works a respective candidate work that is cited in the digital document.
8. The computer system of claim 1, wherein the novelty module is further configured to determine if a respective candidate work in the set of candidate works is relevant and, if the respective candidate work is not, is further configured to remove from the set of candidate prior works the respective candidate work.
9. The computer system of claim 1, wherein the multi-modal data comprises one or more figures and any associated captions.
10. The computer system of claim 1, wherein each of the specialized agents is configured to access at least one of the general knowledge unit, domain knowledge unit, or external knowledge unit based on a task of each of the specialized agents.27 Docket No. 22151.71a1 1. The computer system of claim 1, wherein the leader agent is further configured to display the review and the digital document within a user interface.
12. The computer system of claim 1. wherein the general instructions comprise a prompt template.
13. The computer system of claim 1, wherein each specialized agent is configured to generate and output a block of text data comprising an assessment of an aspect of the digital document to the leader agent.
14. The computer system of claim 13, wherein the leader agent is configured to synthesize the block of text data from each of specialized agent into the review of the digital document.
15. The computer system of claim 13, wherein the leader agent is configured to receive the block of text data from each specialized agent in parallel.
16. A method for generating a review of a digital document, comprising:(a) receiving, at a computer system, a digital document comprising multi-modal data including text and one or more figures;(b) extracting non-text data and any associated captions from the digital document and generating textual descriptions for the non-text data;(c) generating a domain dataset based on one or more references contained in the text of the digital document;(d) generating an external works dataset comprising a set of candidate prior w orks by generating one or more search queries based on the multi-modal data of the digital document and retrieving one or more candidate prior works from one or more sources external to the computer system using the one or more search queries;28 Docket No. 22151.71a(e) processing each candidate prior work in the set of candidate prior works and determining whether the digital document satisfies a novelty condition based on each candidate prior work;(f) if the novelty condition is not satisfied, generating a report identifying one or more candidate prior works that caused the novelty condition to not be satisfied and terminating further processing of the digital document; and(g) if the novelty condition is satisfied, processing, by a plurality' of specialized agents, the digital document using at least one of the domain dataset, external works dataset, textual descriptions, wherein processing comprises:(i) transmitting, by a leader agent, instructions to one or more other specialized agents to cause each specialized agent to produce one or more outputs comprising an assessment of an aspect of the digital document;(ii) receiving, by the leader agent, the one or more outputs from the one or more other specialized agents;(iii) synthesizing, by the leader agent, the one or more outputs into a synthesized review; and(iv) generating, by the leader agent, the review of the digital document based at least on the synthesized review.
17. The method according to claim 16, wherein the method further comprises removing from the set any candidate prior work that is already cited in the digital document.
18. The method according to claim 16, wherein the plurality' of specialized agents further comprises an impact agent, an experiment agent, and / or a clarity agent.
19. A non-transitory computer-readable medium storing instructions that are configured to be executed by one or more processors to cause the one or more processors to perform the method of claim 16.
20. A computer system for generating a review of a digital document, comprising: one or more pre-processing modules comprising: a general knowledge unit configured to store and output one or more general instructions for constructing a review;29 Docket No. 22151.71a a domain knowledge unit configured to output and dynamically generate a domain dataset from one or more references in the digital document; an external knowledge unit configured to retrieve one or more candidate prior works from one or more sources that are external to the computer system and to generate and output an external works dataset based on the one or more candidate prior works; a figure module configured to extract multi-modal data from the digital document and to generate and output textual descriptions that describe a consistency and / or clarity score associated with the figures and / or captions; and a novelty module configured to generate search queries used to retrieve the one or more candidate prior works and to process the one or more candidate prior works to generate and output a report indicating whether the digital document satisfies a novelty condition; and one or more processing modules that are each configured to retrieve at least one output of the one or more pre-processing modules, wherein the one or more processing modules comprise a plurality of specialized agents comprising leader agent and one or more other specialized agents, wherein: the leader agent is configured to transmit instructions to the one or more other specialized agents, to receive and synthesize outputs from the one or more other specialized agents, and to generate the review of the digital document based at least on the synthesized outputs; and the one or more other specialized agents are each configured to process the digital document based on the at least one output of the one or more pre-processing modules and to generate an output to the leader agent indicating an assessment of a respective aspect of the digital document, wherein each of the one or more processing modules comprises a large language model configured to produce a text output based on one or more text inputs.