Intelligent agent evaluation method and device, electronic equipment and storage medium

By cross-reviewing general large models and domain large models, the problem of unreliable assessment of report generation capabilities in professional domains in existing agent evaluation methods is solved, and multi-dimensional evaluation accuracy and reliability are achieved.

CN121859944AActive Publication Date: 2026-04-14ZHEJIANG LAB

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing intelligent agent evaluation methods cannot accurately reflect an agent's ability to generate research reports in a specific field, leading to unreliable evaluation results.

Method used

Cross-review using general large models and domain large models is employed. Literature is reviewed in multiple dimensions through general review dimensions and professional review dimensions, generating first and second review reports. The report score is determined by combining the literature and the review report, thereby evaluating the report generation capability of the agent.

Benefits of technology

It provides an effective and reliable evaluation method that can truly reflect the report generation capabilities of intelligent agents in professional fields, thereby improving the accuracy and applicability of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859944A_ABST
    Figure CN121859944A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent agent evaluation method and device, electronic equipment and a storage medium, and relates to the field of artificial intelligence. Obtaining the literature, inputting the research problem of the literature into the to-be-evaluated intelligent agent, and obtaining a to-be-evaluated research report. Based on the first cue word project, guiding the general large model to review the literature from the general review dimension to obtain a first review report; a corpus in the research field of the literature is utilized in advance to train a field large model, and the field large model is guided to review the literature from the professional review dimension related to the research field based on the second cue word project to obtain a second review report. Determining a report score of the research report on the basis of the literature and the review report; and determining an evaluation result according to the report score. The research report is subjected to cross review from different dimensions by utilizing a general large model and a field large model, and the research report is subjected to multi-dimensional scoring by taking literatures and review results as dual basis, so that the review results can reflect report generation capabilities of agents in various fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device and storage medium for evaluating intelligent agents. Background Technology

[0002] With the rapid development of artificial intelligence technology, using intelligent agents to generate research reports to assist scientific research has become a common trend. To ensure the scientific rigor, reliability, and practicality of research reports generated by intelligent agents, it is necessary to evaluate the report generation capabilities of these agents.

[0003] Current agent evaluation methods primarily assess an agent's report generation capabilities from general dimensions such as text coherence and language standardization. For agents generating research reports in general domains, these methods achieve good results and are relatively reliable. However, for agents generating research reports in specialized domains, the significant differences between specialized and general domain knowledge render the results unreliable, failing to accurately reflect the agent's report generation capabilities.

[0004] Therefore, there is an urgent need to provide a method for evaluating intelligent agents to improve the accuracy and applicability of the evaluation. Summary of the Invention

[0005] In view of this, this application provides a method for evaluating intelligent agents, the method comprising: Obtain the literature and determine the research questions and research areas of the literature; The research question is input into the agent to be evaluated, and the research report to be evaluated is generated by the agent based on the research question. Based on the first prompt word project, the document is reviewed from the preset general review dimensions using a general large model to obtain the first review report; Based on the second prompt word project, the literature is reviewed by a domain-wide model according to preset professional review dimensions to obtain a second review report; the domain-wide model is a model obtained by training a pre-trained large model using a corpus of the research field; the professional review dimensions are dimensions related to the research field. Based on the aforementioned literature, the first review report, and the second review report, a report score for the research report to be evaluated is determined; and based on the report score, an evaluation result is determined to characterize the report generation capability of the agent to be evaluated.

[0006] Optionally, the constraints in the first prompt word project include: Based on the content and structure of the document, the target document type is determined; the target document type is either a review or a research document. If the target document type is the review type, then at least one of the following will be used as the general review dimension: document structure, document content confidence, and coverage of the research field. If the target document type is the research type, then at least one of the following will be used as the general review dimension: document structure, document content confidence, and degree of difference from existing documents. The literature is reviewed and analyzed according to the general review dimensions to obtain a first review text; wherein, the first review text includes at least one of the following: the strength analysis text, the limitation analysis text, and the improvement suggestions of the literature in the corresponding general review dimensions; The first review report is obtained by summarizing the first review texts corresponding to each of the general review dimensions.

[0007] Optionally, based on the content and structure of the document, the target document type corresponding to the document is determined, including: Identify the chapter titles of the document and assemble the chapter titles into the document's structure; Determine the similarity between the document structure and the preset research structure; the chapter titles in the preset research structure include introduction, methods, results, and discussion; the similarity is used to characterize the degree of similarity between the document structure and the preset research structure; Determine the proportion of original content in the cited literature relative to the original content in other cited literature; Determine the frequency of occurrence of preset keywords in the literature; the preset keywords are words related to the literature of the research type. The document type score is obtained by weighted summing of the similarity, the originality ratio, and the frequency. Determine whether the score for the document type is greater than the preset score; If yes, then the target document type is determined to be the research type; if no, then the target document type is determined to be the review type.

[0008] Optionally, the research field is the field of Earth sciences; correspondingly, the constraints in the second prompt word project include: The professional review dimensions include at least one of the following: the degree of conformity between the reasoning thinking of the literature and the reasoning thinking in the field of earth sciences; the time scale adopted by the literature and its suitability to the research question; the spatial scale adopted by the literature and its suitability to the research question; the relevance of the literature content to the field of earth sciences; and the confidence level of the literature content. The literature is reviewed and analyzed according to the aforementioned professional review dimensions to obtain a second review text; wherein, the second review text includes at least one of the following: an analysis text of the literature's strengths, an analysis text of its limitations, and suggestions for improvement in the corresponding professional review dimensions; The second review report is obtained by summarizing the second review texts corresponding to each of the aforementioned professional review dimensions.

[0009] Optionally, the report score of the research report to be evaluated is determined based on the aforementioned literature, the first review report, and the second review report, including: Based on the third prompt word project, the literature, the first review report and the research report to be evaluated are input into the general model to obtain the first score of the general model for evaluating the research report to be evaluated from each of the general review dimensions, based on the literature and the first review report. Based on the fourth prompt word project, the literature, the second review report and the research report to be evaluated are input into the domain big model to obtain the second score of the domain big model based on the literature and the second review report, respectively, from each of the professional review dimensions. The report score is determined based on each of the first scores and each of the second scores.

[0010] Optionally, the constraints in the third prompt word project include: Extract the first review text corresponding to the general review dimension from the first review report; the first review text includes the strength analysis text, the limitation analysis text, and improvement suggestions of the literature in the general review dimension. Using the first review text as a basis for comparison, compare the degree of consistency between the research report to be evaluated and the literature in the general review dimensions; Based on the degree of consistency, the first score is determined to be positively correlated with the degree of consistency.

[0011] Optionally, the report score is determined based on each of the first scores and each of the second scores, including: Determine the first average score for each of the first scores and the second average score for each of the second scores; Obtain a pre-set mapping relationship; the mapping relationship is the correspondence between review dimensions and scoring weights; the review dimensions include the general review dimensions and the professional review dimensions. Based on the mapping relationship, determine the general scoring weights corresponding to the general review dimensions and the professional scoring weights corresponding to the professional review dimensions. The general scoring weight is used as the weight corresponding to the first average score, and the professional scoring weight is used as the weight corresponding to the second average score; The report score is obtained by weighted summing of the first average score and the second average score.

[0012] This application also provides an evaluation device for intelligent agents, the device comprising: The acquisition module is used to acquire literature and determine the research questions and research fields of the literature; The report generation module is used to input the research question into the agent to be evaluated and obtain the research report to be evaluated generated by the agent to be evaluated based on the research question; The first review module is used to review the document based on the first prompt word project and through a general model from preset general review dimensions to obtain a first review report; The second review module is used to review the literature based on the second prompt word project and according to the preset professional review dimensions using a domain-wide model, and to obtain a second review report; the domain-wide model is a model obtained by training a pre-trained large model using a corpus of the research field; the professional review dimensions are dimensions related to the research field. The evaluation module is used to determine the report score of the research report to be evaluated based on the literature, the first review report, and the second review report; and to determine the evaluation result that characterizes the report generation ability of the agent to be evaluated based on the report score.

[0013] This application also provides an electronic device, including: Memory, used to store computer programs; A processor, used to implement the steps of any of the above-described intelligent agent evaluation methods when executing the computer program.

[0014] This application also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described intelligent agent evaluation methods.

[0015] In summary, this application provides an evaluation method, apparatus, electronic device, and storage medium for intelligent agents. First, literature is acquired, and the research questions and research fields of the literature are determined. The research questions are input into the intelligent agent to be evaluated, resulting in a research report to be evaluated. Based on the first prompt word engineering, a general large model is guided to review the literature from a general review dimension, resulting in a first review report. Furthermore, a domain large model is pre-trained using a corpus of the literature's research field to obtain a domain large model. Based on the second prompt word engineering, the domain large model is guided to review the literature from a professional review dimension related to the research field, resulting in a second review report. Based on the literature, the first review report, and the second review report, a report score for the research report to be evaluated is determined. Based on the report score, an evaluation result is determined to characterize the report generation ability of the intelligent agent to be evaluated. By using the general large model and the domain large model to cross-review the research report to be evaluated from different dimensions, and using both the literature and the review results as dual bases, a multi-dimensional score is applied to the research report to be evaluated. This ensures that the evaluation results can truly reflect the report generation ability of intelligent agents in various fields, providing an effective and reliable evaluation method. Attached Figure Description

[0016] Figure 1 A first flowchart illustrating a method for evaluating an intelligent agent provided in this application; Figure 2 A schematic diagram illustrating the principle of an evaluation method for an intelligent agent provided in this application; Figure 3 A second flowchart illustrating an evaluation method for an intelligent agent provided in this application; Figure 4 A schematic diagram of the structure of an evaluation device for an intelligent agent provided in this application; Figure 5 This is a schematic diagram of the structure of a storage medium provided in this application. Detailed Implementation

[0017] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0018] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0019] Please refer to Figure 1 , Figure 1 A first flowchart illustrating an evaluation method for an intelligent agent provided in this application, the method comprising: S101. Obtain the literature and determine the research questions and research areas of the literature.

[0020] S102. Input the research question into the agent to be evaluated, and obtain the research report to be evaluated generated by the agent based on the research question.

[0021] This application does not impose any special restrictions on the rules or quantity for obtaining the aforementioned literature. For example, first determine the research field of the intelligent agent application to be evaluated, and then obtain the aforementioned literature from a literature database in that research field. Please refer to... Figure 2 , Figure 2 This is a schematic diagram illustrating the principle of an evaluation method for an intelligent agent provided in this application. Taking the intelligent agent to be evaluated as generating research reports in the field of Earth sciences as an example, it selects literature published later than a preset time from journals in the field of Earth sciences (such as *Nature Geoscience* and *Geology*) to construct a literature set in the field of Earth sciences. The number of literatures can be specified according to actual needs, for example, selecting more than 30 literatures.

[0022] After obtaining the literature, the research questions and research areas are determined. These research questions and research areas can be determined manually by experts; alternatively, existing large language models can be used to analyze the literature and automatically determine the research questions and research areas based on the literature content. This application does not impose any particular limitations on this approach.

[0023] The research questions are input into the AI ​​agent to be evaluated, which then automatically generates a research report based on the questions. The performance of the research report represents the report generation capability of the AI ​​agent. Therefore, this application subsequently determines the evaluation result of the AI ​​agent based on the report scores of the research report in each review dimension.

[0024] S103. Based on the first prompt word project, the literature is reviewed from the preset general review dimensions through a general large model to obtain the first review report.

[0025] The aforementioned general-purpose large-scale model is a model obtained by pre-training a pre-trained large-scale model using general corpora (including but not limited to news articles, encyclopedias, and literary works). This general-purpose large-scale model is capable of performing various tasks such as language understanding and text generation, and is suitable for reviewing documents from general review dimensions (such as document organization and linguistic clarity), but its performance in specific professional fields is insufficient. Therefore, this application, based on the first prompt word project, guides the general-purpose large-scale model to review documents from general review dimensions, obtaining a first review report.

[0026] The aforementioned general review dimensions are applicable to all research fields and can be customized according to specific needs. For example, general review dimensions include: document organization structure (further including logical coherence, document structure rationality, and chapter content balance), data and evidence (further including the quality of cited literature and the degree to which the document's data and viewpoints are supported by the cited literature), coverage of knowledge in the research field, timeliness, expression ability (further including language clarity, chart quality, and document layout), and soundness (further including the rigor of argumentation, academic value, and practicality).

[0027] The first prompt word project and general review dimensions will be explained in detail later, and will not be elaborated here.

[0028] S104. Based on the second prompt word project, the literature is reviewed by the domain big model according to the preset professional review dimensions to obtain the second review report; the domain big model is the model obtained by training the pre-trained big model using the corpus of the research field; the professional review dimensions are the dimensions related to the content of the research field.

[0029] Taking the field of Earth Sciences as an example, due to its large time span, strong interdisciplinary nature, and emphasis on data (such as maps, remote sensing images, geochemical diagrams, stratigraphic profiles, and chronological tables), general-purpose large-scale models struggle to accurately analyze and understand the content of this field, and cannot accurately identify professional logical errors. Therefore, this application pre-trains a large-scale model using a corpus of the research field to which the literature pertains (including but not limited to professional journals, conference papers, patent databases, and technical manuals), resulting in a domain-specific large-scale model. This domain-specific large-scale model possesses the professional knowledge of the research field, enabling it to effectively understand and analyze literature and identify professional logical errors.

[0030] Based on this, this application utilizes the second prompt word project to guide a large-scale domain model to review literature from professional review dimensions, resulting in a second review report. The aforementioned professional review dimensions are defined for the research field to which the literature belongs; different research fields require different professional review dimensions, which can be pre-set according to actual needs. Taking the Earth Science field as an example, professional review dimensions may include the compatibility of the literature's reasoning logic with geological / geophysical reasoning logic, the compatibility of the time scale (e.g., instantaneous events or long-term evolution) and spatial scale (e.g., microscopic, outcrop, regional, or global) used in the research with the geological processes studied, the appropriateness of the literature data to the Earth Science field, the research significance, and the reliability of the research results.

[0031] The first prompt word project and professional review dimensions will be explained in detail later, and will not be elaborated here.

[0032] In short, such as Figure 2 As shown, this application utilizes a general large model and a domain large model to conduct cross-reviews of literature, obtaining multi-dimensional (i.e., including general review dimensions and professional review dimensions) review results. Subsequent scoring of the research report to be evaluated based on these review results is equivalent to scoring the report from multiple dimensions.

[0033] S105. Based on the literature, the first review report, and the second review report, determine the report score of the research report to be evaluated; and based on the report score, determine the evaluation result used to characterize the report generation ability of the agent to be evaluated.

[0034] In this application, the report score of the research report to be evaluated is determined by combining the literature, the first review report, and the second review report. The first review report contains the results of the literature review and analysis from a general review perspective, while the second review report contains the results of the literature review and analysis from a professional review perspective. Determining the report score of the research report to be evaluated based on the first and second review reports is equivalent to determining the report score from both general and professional review perspectives, thus improving the reliability of the report score.

[0035] Furthermore, this application does not solely rely on literature or review reports as scoring criteria, but rather uses both. The research question is analogous to the title, the literature to the corresponding answer, and the review report to the analytical approach. When reviewing the research report to be evaluated, the analytical approach is used to assess the extent to which the report answers the question, much like the literature. In other words, the review report is equivalent to an analysis of the literature from both general and professional review dimensions. This analysis can be used to determine the degree of consistency between the research report and the literature across various review dimensions, and the report score is determined based on this consistency. Subsequent embodiments will explain the process of determining the report score; it will not be elaborated upon here.

[0036] After determining the report score of the research report to be evaluated, the evaluation result used to characterize the report generation capability of the agent under evaluation can be determined based on the report score. As an optional embodiment, the average report score of all research reports to be evaluated is used as the evaluation result. The evaluation result characterizes the comprehensive report generation capability of the agent under evaluation in both general review and professional review dimensions; the higher the evaluation result, the stronger the report generation capability of the agent under evaluation.

[0037] In summary, this application pre-trains a domain-specific large-scale model using a corpus of the research field to which the literature pertains. It then employs a general large-scale language model and a domain-specific large-scale language model to cross-review the research reports to be evaluated. The generated review reports include multi-dimensional review opinions encompassing both general and professional review dimensions. Furthermore, using both the literature and the review reports as dual criteria, the reports generated by the agent to be evaluated are scored across multiple dimensions, providing an effective and reliable evaluation method for assessing the agent's report generation capabilities within a specific professional field.

[0038] Based on the above embodiments: As an optional embodiment, the constraints in the first prompt word project include: Based on the content and structure of the literature, determine the target literature type; the target literature type is either a review or a research paper. If the target document is a review article, at least one of the following criteria will be used as a general review dimension: document structure, document content confidence, and coverage of the research field. If the target document type is research type, then at least one of the following will be used as the general review dimension: document structure, document content confidence, and degree of difference from existing documents. The literature is reviewed and analyzed according to the general review dimensions to obtain the first review text; the first review text includes at least one of the following: the strength analysis text, the limitation analysis text, and the improvement suggestions of the literature in the corresponding general review dimensions; The first review texts corresponding to each general review dimension are compiled to obtain the first review report.

[0039] Considering that different types of literature have different focuses, for example, review articles mainly summarize and comment on existing literature, lacking independent data collection, methodological design, and empirical analysis; research articles typically include a clearly defined research question, independent data or experimental design, self-collected or generated data, and analysis and discussion of the results. To ensure the accuracy of literature review and analysis based on general review dimensions, this embodiment first determines the target literature type based on the content and structure of the literature, and then selects the corresponding general review dimensions for the target literature type.

[0040] For review articles, at least one of the following criteria should be used as a general review dimension: article structure, confidence level of content, and coverage of the research field. Article structure can be determined based on the chapter titles, confidence level of content can be determined based on the quality of cited literature, and coverage of the research field can be determined by comparing existing research methods and existing literature.

[0041] The following is an example of first-click keyword engineering for review articles: "You are a rigorous yet impartial academic reviewer. Please provide a comprehensive review of the following academic review. Your review should be based on the specific content, conducting a critical analysis from the following five dimensions, and providing constructive feedback. The review dimensions include:" 1. Document Organization Logical coherence: Does the literature review present a clear narrative thread and are the transitions between sections natural? Structural rationality: Does the introduction clearly define the scope? Is the main body of the literature reasonably organized according to the theme / time / methodology? Does the conclusion of the literature effectively summarize and point out the future research direction? Chapter balance: Whether the length of each part of the literature is reasonably allocated and whether the key points are highlighted.

[0042] 2. Confidence level of the literature content, and data and evidence in the literature. Literature selection quality: Whether it cites landmark, high-impact, and latest key literature in the field; Critical analysis: Does it go beyond simple description and involve comparative, critical, and comprehensive analysis of cited literature? Supporting evidence: Whether the viewpoints in the literature are adequately supported by the cited literature, and whether there are any unfounded assertions.

[0043] 3. Coverage of the research field Comprehensiveness: Does it cover the major schools of thought, key debates, important developments, and methodologies in this research field? Representativeness: Whether different viewpoints are presented in a balanced manner, and whether there are obvious biases or omissions; Timeliness: Does it include important literature from the last 3-5 years and does it identify emerging trends?

[0044] 4. Presentation skills Clarity of language: Whether the document is accurate and concise, and whether the use of technical terms is appropriate; Visualization: Are the charts necessary, clear, informative, and consistent with the text explanation? Academic standards: Whether the citation format is consistent and complete, and whether the citation layout is professional and easy to read.

[0045] 5. Soundness Argumentation rigor: Whether the conclusions of the literature are reasonably derived from the literature being reviewed and whether the logical chain is complete; Academic value: Does it offer new insights, a comprehensive framework, or a critical perspective? Practicality: Does it provide substantial guidance for researchers in the field?

[0046] Please provide your review report strictly according to the following structure: 1. Overall evaluation, i.e., a summary of the core contributions, main strengths, and key weaknesses of the literature in one paragraph; 2. Detailed review of each of the above dimensions, specifically pointing out strengths, problems, and suggestions. The review process should adhere to the following principles: high standards and evidence-based evaluation; specific improvement suggestions for each issue; acknowledging strengths while clearly pointing out weaknesses; avoiding subjective bias and focusing on academic quality. Please review based on the provided literature. If necessary information is missing, please indicate what the authors need to supplement. We will now begin reviewing the literature. For research-type literature, at least one of the following should be used as a general review dimension: literature structure, confidence level of literature content, and degree of differentiation from existing literature. Literature structure can be determined based on the chapter titles, and confidence level of literature content can be determined based on the quality of cited literature.

[0047] The following is an example of a first-catch word project for literature of a specific research type: "You are a rigorous yet impartial academic reviewer. Please provide a comprehensive review of the following academic review. Your review should be based on the specific content, conducting a critical analysis from the following five dimensions, and providing constructive feedback. The review dimensions include:" 1. Degree of differentiation from existing literature, originality. Research question: Does it raise novel and meaningful research questions or hypotheses? Methods and Design: Are the research methods, experimental designs, or analytical frameworks innovative? Results found: Did the research findings provide new evidence, patterns, or theoretical insights? Academic positioning: Does it clearly explain the differences and advancements between this study and existing literature?

[0048] 2. Document Organization IMRaD structural integrity: Does it include the core sections such as Introduction, Methods, Results, and Discussion? Is the logical order reasonable? Internal logic within each chapter: Is the structure of each chapter clear and the arguments coherent? Transition and connection: Are there clear logical connections between the different parts? Does it form a complete academic narrative? Balance of length: Is the length of each chapter allocated reasonably (e.g., is the methodology section sufficiently detailed, and is the discussion adequate)?

[0049] 3. Contribution Theoretical contributions: Whether it has verified, extended, modified, or challenged existing theories; Practical contribution: Whether the research findings have practical application value or policy implications; Methodological contributions: Does it provide new research methods, tools, or analytical perspectives? Advancement in the field: Does it clearly indicate how this research will advance the field?

[0050] 4. Confidence level of the literature content, data and evidence. Methodological rigor: Is the research design scientifically sound and reasonable? Are the data collection / experimental processes reproducible? Data analysis: Are the analytical methods appropriate? Are the statistical tests used correctly and sufficiently? Results presentation: Are the results presented clearly and completely? Are the charts and graphs necessary and standardized? Supporting conclusions: Are the discussion and conclusions entirely based on the results? Is there any over-interpretation?

[0051] 5. Clarity Precision of expression: Are academic languages ​​accurate and unambiguous? Are professional terms used appropriately? Readability: Are the core arguments, methods, and findings easy to understand? Is the narrative fluent? The combination of charts and text: Are the charts clear and effective? Are they consistent with and complementary to the text descriptions? Formatting guidelines: Do citations, references, figure and table labels conform to academic standards?

[0052] Please provide your review report strictly according to the following structure: 1. Overall evaluation, i.e., a summary of the core contributions, main strengths, and key weaknesses of the literature in one paragraph; 2. Detailed review of each of the above dimensions, specifically pointing out strengths, problems, and suggestions. The review process should adhere to the following principles: high standards and evidence-based evaluation; specific improvement suggestions for each issue; acknowledging strengths while clearly pointing out weaknesses; avoiding subjective bias and focusing on academic quality. Please review based on the provided literature. If necessary information is missing (such as methodological details, data sources, and statistical values), please indicate what needs to be supplemented by the authors. We will now begin reviewing the literature. After determining the general review dimensions based on the target document type, the documents are reviewed and analyzed according to these dimensions to obtain the first review text. The first review text includes at least one of the following: an analysis of the document's strengths, limitations, and improvement suggestions for each general review dimension. The first review texts corresponding to each general review dimension are then compiled to obtain the first review report.

[0053] The following is an example of the first review report obtained based on the above-mentioned first prompt word project: Abstract: This paper synthesizes geological evidence from early Mars, suggesting multiple climate transitions and challenging the single paradigm of a wet-to-dry transition. It proposes seven major climate transitions supported by geomorphological observations, mineralogical and stratigraphic studies, and discusses the driving roles of factors such as dip changes, volcanic exhaust, and atmospheric escape. The paper emphasizes the significance of these findings for habitability and exoplanet research. Organization: The sections are logically coherent, but lack clear definitions of the transitions. The seven-transition framework is innovative but lacks sufficient motivation. Some key studies (such as Turbet) are omitted. The research on carbon dioxide ice clouds by [authors' name] reduces the comprehensiveness of the content. Data and Evidence: Over-reliance on images (Figures 1-5 in the literature). Claims such as wind erosion over 1 km lack quantitative support. Stratigraphic correlations (Figure 2 in the literature) are presented as facts without further verification. The interpretation of meteorite data is overly simplistic. Innovation: The multi-transformation model provides a new comprehensive analysis. It is the first time that dip chaos has been linked to habitable windows. Introducing a comparative framework for exoplanets into a geological paper is highly innovative. Contribution: The framework is valuable, but its implementation is flawed. The core concept of the "seven transformations" needs rigorous definition. The mentioned supplementary tables are not provided. If the habitability inferences are validated, it may reshape the mission planning. Clarity: Excessive use of technical terms (such as rhythmic rocks and GELs) without definition. Grammatical errors hinder understanding (e.g., Mars is dry today). Incoherent chart citations. The drivers of transformations are discussed before the evidence is presented. In summary, this embodiment selects different general review dimensions for different types of literature to ensure the accuracy of the review and analysis of literature from the general review dimensions. This provides a data foundation for the subsequent accurate scoring of the research reports to be evaluated based on the first review text, avoiding inaccurate scoring.

[0054] As an optional implementation, the target document type is determined based on the content and structure of the document, including: Identify the chapter titles of a document and assemble them into the document's structure; Determine the similarity between the document structure and the pre-defined research structure; the chapter titles in the pre-defined research structure include the introduction, methods, results, and discussion; the similarity is used to characterize the degree of similarity between the document structure and the pre-defined research structure; Determine the proportion of original content in the cited literature to the original content in other cited literature; Determine the frequency of the pre-defined keywords in the literature; the pre-defined keywords are words related to the research type of the literature. The document type score is obtained by weighting and summing the similarity, originality ratio, and frequency. Determine whether the document type score is greater than the preset score; If yes, then the target document type is determined to be a research type; if not, then the target document type is determined to be a review type.

[0055] Considering that research-type literature typically conforms to the IMRAD structure (chapter titles include Introduction, Methods, Results, and Discussion), and that the proportion of original content is high, it's worth noting that research-type literature generally follows an IMRAD structure. Furthermore, the frequently occurring words differ between research-type and review-type literature. Review-type literature often uses keywords such as "review," "retrospective," "progress," and "critique," while research-type literature typically uses keywords such as "methods," "experiment," "results," "analysis," and "validation." The proportion of original content also differs between review-type and research-type literature. Review-type literature primarily summarizes and reviews existing literature, lacking independent data collection, methodological design, and empirical analysis sections, resulting in a low proportion of original content. Research-type literature typically includes a clearly defined research question, independent data or experimental design, self-collected or generated data, and analysis and discussion of results, resulting in a high proportion of original content.

[0056] Based on this, this embodiment determines the target document type from three perspectives: the similarity between the document structure and the preset research structure, the proportion of original content in the document relative to the original content cited in other documents, and the frequency of preset keywords appearing in the document. The chapter titles in the preset research structure include Introduction, Methods, Results, and Discussion; similarity is used to characterize the degree of similarity between the document structure and the preset research structure; that is, the higher the similarity, the higher the probability that the document is the target document type. A higher originality ratio also increases the probability that the document is the target document type. Preset keywords are words related to the research type of the document and can be set according to actual needs; the frequency of preset keywords appearing in the document further increases the probability that the document is the target document type.

[0057] Based on this, the similarity, originality ratio, and frequency are weighted and summed to obtain the document type score. The weights corresponding to similarity, originality ratio, and frequency can be set according to the importance of the document type score; this embodiment does not impose any special limitations on this. If the document type score is greater than the preset score, the target document type is determined to be a research type; otherwise, the target document type is determined to be a review type.

[0058] Below are examples of prompts for determining the target document type: Please follow these criteria to determine the type of document: Characteristics of review articles: They primarily summarize and comment on existing literature, lacking independent sections on data collection, methodology design, and empirical analysis. The organizational structure is often based on themes or theoretical frameworks, without clearly defined separate chapters such as "Methods," "Results," or "Discussion." Chapter titles may contain keywords such as "review," "retrospective," "progress," and "review," and the number of references is usually large, primarily citing the work of others.

[0059] Characteristics of this research type: It follows the basic organizational structure of "Introduction, Methods, Results, and Discussion." It includes a clearly defined research question, independent data / experimental design, self-collected or generated data, and analysis and discussion of the results. The research design, data sources, experimental procedures, or analytical tools are usually described in detail in the "Methods" section.

[0060] A cautionary tale: Some papers may combine both characteristics (e.g., "review + case study"), and the document type should be determined based on the proportion of the main content. Early studies or short papers may omit parts of the IMRAD section, but they still need to be distinguished by the presence or absence of empirical analysis.

[0061] Output format requirements: Please reply strictly according to the following format: 1. Judgment result: [Review type / Research type / Difficult to judge]; 2. Key evidence: List 3-5 organizational or content features that support the judgment result; 3. Uncertainty Explanation: Indicate any parts that are questionable or require further investigation. In summary, this embodiment sets evaluation dimensions to determine the document type based on the differences in organizational structure and content between review articles and research articles, thereby accurately selecting the corresponding general review dimensions for different types of articles and improving the reliability of the first review report.

[0062] As an optional implementation, the research field is Earth Science; correspondingly, the constraints in the second prompt word project include: The professional review will consider at least one of the following dimensions: the degree of conformity between the reasoning thinking in the literature and the reasoning thinking in the field of earth sciences; the time scale used in the literature and its fit with the research question; the spatial scale used in the literature and its fit with the research question; the appropriateness of the content of the literature to the field of earth sciences; and the confidence level of the content of the literature. The literature is reviewed and analyzed according to professional review dimensions to obtain the second review text; the second review text includes at least one of the following: the strength analysis text, the limitation analysis text, and the improvement suggestions of the literature in the corresponding professional review dimension; The second review texts corresponding to each professional review dimension are compiled to obtain the second review report.

[0063] In this embodiment, given the large time and spatial spans inherent in the field of Earth sciences, the time scale adopted by the literature and its fit with the research question, as well as the fit between the spatial scale adopted by the literature and the research question, are used as professional review dimensions. Specifically, the degree of matching between the time scale (e.g., instantaneous events or long-term evolution) and spatial scale (e.g., microscopic, outcrop, regional, or global) involved in the literature's research and the geological processes studied in the literature are used as professional review dimensions. Furthermore, given the unique reasoning and research thinking, strong interdisciplinary nature, and emphasis on data within the field of Earth sciences, the degree of conformity between the literature's reasoning and the reasoning thinking within the Earth science field, the relevance of the literature's content to the Earth science field, and the confidence level of the literature's content are also used as professional review dimensions.

[0064] The above professional review dimensions are input into the domain-wide model, which then guides the model to automatically review and analyze the literature according to these dimensions. Each professional review dimension corresponds to a second review text. The second review texts corresponding to each professional review dimension are then compiled to obtain the second review report.

[0065] The following is an example of the second prompt word project for the field of Earth sciences: "You are a rigorous and impartial reviewer in the field of Earth Sciences with solid expertise in geology and geophysics. Please conduct a comprehensive and in-depth review of the following paper, focusing on the rigor and professionalism of its Earth science thinking. Your review should be based on the specific content, conducting a critical analysis from the following four core dimensions, and providing constructive and professional feedback."

[0066] 1. Scientific rigor and geological / geophysical logic Geological logical chain: From scientific questions to method selection, data acquisition, results analysis and final conclusions, does the entire reasoning process conform to geological thinking? Is the research logic rigorous and self-consistent? Handling of multiple solutions and uncertainties: Is the multiple solution nature of Earth science problems fully recognized? Have the uncertainties of data and the limitations of model assumptions been clearly discussed and evaluated? Matching of time and space scales: Whether the time scale (e.g., instantaneous events or long-term evolution) and spatial scale (e.g., microscopic, outcrop, regional or global) involved in the study match the geological processes being studied; whether the scale conversion is reasonable; Geological constraints on geophysical interpretation: For geophysical papers, are their physical property models and geological interpretations constrained by sufficient geological observations (such as core samples, outcrops, and well drilling) and is there any over-interpretation?

[0067] 2. Appropriateness of Data and Methods to Earth Sciences Data quality and representativeness: Whether the data used (field data, core data, geophysical data, geochemical data, and remote sensing data, etc.) can effectively support the research questions; whether the sampling / observation scheme avoids systematic bias; whether the influence of geological factors such as rock heterogeneity and weathering is considered; Geological rationality of method selection: Whether the selected analytical techniques, experimental methods, simulation software, or inversion algorithms are suitable for the target geological body and scientific problem; whether the parameter settings have a geological basis; The integration of methods and geological models: Does it clearly explain how specific geological processes, structures, or properties can be revealed or verified through methods such as geophysical inversion and numerical simulation? Reproducibility and openness: Are the data sources, processing methods, and key parameters described in sufficient detail so that peers can replicate or verify them? Are the ways to share the data mentioned?

[0068] 3. Geological background and regional significance Regional geological framework integration: Whether the study is adequately placed within the framework of regional geological evolution; whether key existing regional geological, geophysical, or geochemical research findings are cited and discussed; Regional significance of the scientific questions: This study aims to address which key gaps or controversies in regional geological understanding; and what specific contributions its findings make to understanding the region's geological history, tectonic evolution, resource distribution, or disaster context. Extension from point to surface: Whether the findings of local studies have been reasonably explored for their implications for similar geological backgrounds in larger regions or even globally; Map quality: Are geological maps, cross-sections, and structural outlines professional and accurate? Are all map elements complete (including scale, legend, orientation, and geological unit symbols)?

[0069] 4. Geoscientific Interpretation and Discussion of Results The depth of interpretation from data to geological processes: whether observational data, simulation results, or test data have been successfully "translated" into a deeper understanding of geological processes, mechanisms, or history; Dialogue with existing knowledge: Does the discussion section substantially compare the new results with existing theories and knowledge? Does it support revisions or challenge existing knowledge? Does it propose new geologically based models or hypotheses? Clarity of contribution: Whether the specific contribution of this study to improving our understanding of a particular geological process, event, system, or region is clearly and unexaggerated; Inspiration for future research directions: Based on the limitations or new findings of this study, have you proposed any geologically insightful future research directions and key questions?

[0070] Please provide a professional review report strictly following the structure below: I. Overall Evaluation Summarize the core geological science questions, main methods used, and key geological conclusions of the paper. Point out the most prominent strengths and most serious weaknesses of the paper in terms of geological thinking or geophysical logic.

[0071] II. Detailed review by dimension (please elaborate based on specific chapters, figures, and data) 1. Scientific rigor and geological / geophysical logic Advantages: [For example, the logical chain is clear, and multiple solutions are fully considered...]; Problems and concerns: [For example, jumping directly from result X to conclusion Y lacks intermediate geological reasoning; assumptions about method Y may seriously violate known geological conditions Z in the region…] Suggested revisions: [Provide specific and actionable suggestions for revisions, and request the author to supplement or clarify logical steps].

[0072] 2. Appropriateness of Data and Methods to Earth Sciences Advantages: [For example, unique data sources, highly targeted method selection, etc.] Problems and concerns: [For example, the sampling density is insufficient to represent the heterogeneity of the strata; the inversion algorithm used may be insensitive to parameter B under condition A, which is a key geological variable...] Suggested revisions: [It is recommended to add data explanations, method verification, or discuss limitations].

[0073] 3. Geological background and regional significance Advantages: [For example, a clear framework for regional tectonic evolution and well-defined research questions...] Problems and concerns: [For example, key comparative studies published in the region in recent years (C) were ignored; the discussion of basement properties was divorced from the results of deep exploration in the region (D)...] Suggested revision: [The author is required to integrate relevant literature and more accurately pinpoint their regional contribution].

[0074] 4. Geoscientific Interpretation and Discussion of Results Advantages: [For example, it offers insightful geological explanations for phenomenon Y...]; Problems and concerns: [For example, interpreting geophysical anomaly E solely as an ore body without fully considering the possibility of other geological bodies (such as alteration zones, lithological interfaces); the discussion fails to link the new findings with classical theory F...]; Suggested revision: [The author is required to strengthen the constraints on interpretation and deepen the dialogue with existing theories].

[0075] III. Detailed List of Revisions Major Defects (Required Revision): List 2-4 fundamental issues that directly affect the scientific credibility and geological rationality of the literature (e.g., errors in key geological assumptions, logical inference faults, neglect of the ambiguity of inversion, obvious contradictions with known geological facts, etc.). Important Improvements (Suggested Revisions): List 2-3 improvements that could significantly enhance the geological depth of the paper (e.g., supplementing the geological background of key areas, deepening the discussion of geological processes, improving the geological information on maps, etc.). Optional Detail Improvements: Point out 1-2 details that could improve professional rigor (e.g., terminology usage, citation standards, map labeling, etc.).

[0076] Review Principles: Professional geological knowledge and logic will be the primary evaluation criteria; all criticisms will aim to enhance the geological content of the paper, and suggestions should be specific, professional, and actionable; the focus will go beyond form, addressing the strength of the geological reasoning and the substantive contributions to Earth science. Please review the full text of the provided literature. The review process will begin now. The following is an example of a second review report obtained based on the above second prompt word project: "Scientific Rigor and Geological / Geophysical Logic: Climate mechanisms (e.g., HAWIC) are physically plausible but have not been rigorously tested; scale mismatches are ignored (e.g., local topography versus global claims); uncertainty propagation in paleotemperature estimates is insufficient. Other driving factors (e.g., impact) are excluded without analysis. Geoscientific Appropriateness of Data and Methods: Orbital data interpretation lacks resolution validation; meteorite analysis (e.g., ALH84001) is over-interpreted; geochronological estimates (Ga) lack error bars; climate model citations (e.g., Wordsworth) are inadequate." The analysis of the results is too superficial; paleohydrological models (e.g., runoff estimates) are not critically applied. Geological context and regional significance: The latitudinal / elevation context is well-described; the lowland and highland records are clearly distinguished; the focus on Valles Marineris is reasonable; the global significance of the volatile matter cycle is convincingly presented. Geoscientific interpretation and discussion of results: The paradoxes (e.g., limited weathering) are creatively explained, but the discussion is insufficient; the connections to exoplanets are insightful; contradictory data (e.g., carbonate deficiencies) are downplayed; the interdisciplinary integration (astrobiology / climatology) is a highlight of this paper. In summary, this embodiment sets up professional review dimensions specifically for literature in the field of earth sciences to ensure the accuracy of literature review and analysis from the perspective of professional review. This provides a data foundation for accurate scoring of the research reports to be evaluated based on the second review text, avoiding inaccurate scoring.

[0077] As an optional implementation, the report score of the research report to be evaluated is determined based on the literature, the first review report, and the second review report, including: Based on the third prompt word project, the literature, the first review report and the research report to be evaluated are input into the general big model. The general big model then evaluates the research report to be evaluated from each of the general review dimensions based on the literature and the first review report. Based on the fourth prompt word project, the literature, the second review report and the research report to be evaluated are input into the domain big model. The domain big model then evaluates the research report to be evaluated from each professional review dimension based on the literature and the second review report. The report score is determined based on the first and second scores.

[0078] In this embodiment, a general large model and a domain large model are used to jointly determine the report score of the research report to be evaluated. Given the powerful semantic understanding and analysis capabilities of the general large model for general knowledge, based on third-prompt word engineering, the general large model is guided to review the literature and the first review report from various general review dimensions, obtaining the first score corresponding to each general review dimension.

[0079] Specifically, the first review report is equivalent to an analysis of the literature from the perspective of common review dimensions. This analysis can be used to determine the degree of consistency between the research report to be evaluated and the literature across each common review dimension, and the first score of the research report to be evaluated is determined based on this consistency. The higher the degree of consistency between the research report to be evaluated and the literature, the higher the first score.

[0080] As an optional implementation, the constraints in the third prompt word project include: Extract the first review text corresponding to the general review dimensions from the first review report; the first review text includes the analysis text of the advantages of the literature in the general review dimensions, the analysis text of the limitations, and the improvement suggestions; using the first review text as the basis for comparison, compare the performance of the research report to be evaluated and the literature in the general review dimensions, and determine the degree of consistency between the research report to be evaluated and the literature; based on the degree of consistency, determine the first score that is positively correlated with the degree of consistency.

[0081] Specifically, for the strengths analysis, limitations analysis, and improvement suggestions in the first review text, the evaluation report is compared to the literature to determine if it shares the same strengths and limitations, and whether the evaluation report already contains corresponding improvement suggestions. If the evaluation report shares the same strengths and limitations as the literature, it indicates a high degree of consistency, thus resulting in a higher first score. If the evaluation report has additional limitations, it indicates a poor report generation ability, thus the first score can be appropriately lowered. If the evaluation report also has additional strengths or contains corresponding improvement suggestions, it indicates a higher report generation ability, thus the first score can be increased.

[0082] Similarly, by leveraging the domain big model's understanding and analytical capabilities of the research domain's professional knowledge, and based on the fourth prompt word project, the domain big model is guided to review the literature from various professional review dimensions based on the literature and the second review report, thereby obtaining the second scores corresponding to each professional review dimension.

[0083] Specifically, the second review report is essentially an analysis of the literature from a professional review perspective. This analysis can be used to determine the degree of consistency between the research report and the literature across various professional review dimensions, and a second score is determined based on this consistency. The higher the degree of consistency between the research report and the literature, the higher the second score.

[0084] As an optional implementation, the constraints in the fourth prompt word project include: Extract the second review text corresponding to the general review dimensions from the second review report; the second review text includes the analysis text of the advantages, the analysis text of the limitations, and the improvement suggestions of the literature in the professional review dimensions; use the second review text as a comparison basis to compare the degree of consistency between the research report to be evaluated and the literature in the professional review dimensions; based on the degree of consistency, determine the second score that is positively correlated with the degree of consistency.

[0085] Specifically, for the strengths analysis, limitations analysis, and improvement suggestions in the second review text, the evaluation report is compared to the literature to determine if it shares the same strengths and limitations, and whether the evaluation report already contains corresponding improvement suggestions. If the evaluation report shares the same strengths and limitations as the literature, it indicates a high degree of consistency, thus resulting in a higher second score. If the evaluation report has additional limitations, it indicates a poor report generation ability, thus the second score can be appropriately lowered. If the evaluation report also has additional strengths or contains corresponding improvement suggestions, it indicates a higher report generation ability, thus the second score can be increased.

[0086] This embodiment does not specify the specific rules for improving the first and second scores. For example, an initial score is set for each review dimension. Based on the consistency between the research report to be evaluated and the literature in that review dimension, the step size is adjusted according to the preset score to increase or decrease the initial score, thereby obtaining the score corresponding to that review dimension.

[0087] After obtaining the first score for each general review dimension and the second score for each professional review dimension, the report score of the report to be evaluated is determined by combining the first and second scores. The specific method for determining the report score is explained below.

[0088] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the second process of an evaluation method for an intelligent agent provided in this application. As an optional embodiment, determining the report score based on each first score and each second score includes: S301. Determine the first average score for each first score and the second average score for each second score; S302. Obtain the pre-set mapping relationship; the mapping relationship is the correspondence between the review dimensions and the scoring weights; the review dimensions include general review dimensions and professional review dimensions; S303. Based on the mapping relationship, determine the general scoring weights corresponding to the general review dimensions and the professional scoring weights corresponding to the professional review dimensions. S304. Use the general rating weight as the weight corresponding to the first average rating, and use the professional rating weight as the weight corresponding to the second average rating. S305. The first average score and the second average score are weighted and summed to obtain the report score.

[0089] In this embodiment, corresponding general scoring weights are pre-set for the general review dimensions, and corresponding professional scoring weights are set for the professional review dimensions. The general scoring weights can be set according to the importance of the report score based on the general review dimensions. Similarly, the professional scoring weights can be set according to the importance of the report score based on the professional review dimensions. This embodiment does not impose any particular limitations on this. For example, the professional scoring weight can be set to 0.6, and the general review dimension can be set to 0.4.

[0090] After obtaining the first scores for each general review dimension, the average of these first scores is determined to obtain the first average score. After obtaining the second scores for each professional review dimension, the average of these second scores is determined to obtain the second average score. Then, using the aforementioned general and professional score weights, the first and second average scores are weighted and summed to obtain the report score.

[0091] For example, general review dimensions include organizational structure (whether the structure / categorization is reasonable, useful, and sufficiently motivated), data and evidence (whether the claims are supported by sufficient data, tables, comprehensive statistical data, and reproducible results), coverage (whether the research field is comprehensively covered and whether key works are omitted), presentation (the quality of charts, organizational structure, and ease of extracting key results), and rationality (whether the analytical methods, comparisons, and any statistics are reasonable and applied prudently). Professional review dimensions include the scientific rigor and geological / geophysical logic in the aforementioned examples, the geoscientific appropriateness of the data and methods, the geological context and regional significance, and the geoscientific interpretation and discussion of the results. Each review dimension starts with a score of 5, where 5 represents excellent, 4 represents good, 3 represents acceptable, 2 represents needing improvement, and 1 represents poor.

[0092] Ultimately, the general model generates first scores of 3, 4, 4, 5, and 3 for each of the above general review dimensions; the domain model generates second scores of 4, 5, 4, and 3 for each of the above professional review dimensions. The report score can then be expressed as... Where 0.4 is the general scoring weight and 0.6 is the professional scoring weight.

[0093] In summary, this embodiment sets corresponding scoring weights based on the importance of the general review dimension and the professional review dimension in report scoring. The first average score of the general review dimension and the second average score of the professional review dimension are then weighted and summed based on these weights to obtain a report score that comprehensively and accurately represents the consistency between the research report under evaluation and the literature. Subsequently, the average report score of all research reports under evaluation can be used as the evaluation result for the agent under evaluation. The score of the evaluation result is positively correlated with the report generation capability of the agent under evaluation, providing an effective and reliable evaluation standard for assessing the agent's report generation capability in a professional field.

[0094] Please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of an evaluation device for an intelligent agent provided in this application. The device includes: The acquisition module 401 is used to acquire literature and determine the research questions and research fields of the literature; The report generation module 402 is used to input research questions into the intelligent agent to be evaluated and obtain the research report to be evaluated generated by the intelligent agent to be evaluated based on the research questions; The first review module 403 is used to review the literature from the preset general review dimensions based on the first prompt word project and through a general large model to obtain the first review report; The second review module 404 is used to review the literature based on the second prompt word project and through the domain big model according to the preset professional review dimensions to obtain the second review report; the domain big model is a model obtained by training a pre-trained big model using a corpus of the research field; the professional review dimensions are dimensions related to the research field; Evaluation module 405 is used to determine the report score of the research report to be evaluated based on the literature, the first review report, and the second review report; and to determine the evaluation result that characterizes the report generation ability of the agent to be evaluated based on the report score.

[0095] For a detailed description of the intelligent agent evaluation device provided in this application, please refer to the embodiments of the intelligent agent evaluation method described above; this application will not repeat the details here.

[0096] Based on the above embodiments: As an optional embodiment, the constraints in the first prompt word project include: Based on the content and structure of the literature, determine the target literature type; the target literature type is either a review or a research paper. If the target document is a review article, at least one of the following criteria will be used as a general review dimension: document structure, document content confidence, and coverage of the research field. If the target document type is research type, then at least one of the following will be used as the general review dimension: document structure, document content confidence, and degree of difference from existing documents. The literature is reviewed and analyzed according to the general review dimensions to obtain the first review text; the first review text includes at least one of the following: the strength analysis text, the limitation analysis text, and the improvement suggestions of the literature in the corresponding general review dimensions; The first review texts corresponding to each general review dimension are compiled to obtain the first review report.

[0097] As an optional implementation, the target document type is determined based on the content and structure of the document, including: Identify the chapter titles of a document and assemble them into the document's structure; Determine the similarity between the document structure and the pre-defined research structure; the chapter titles in the pre-defined research structure include the introduction, methods, results, and discussion; the similarity is used to characterize the degree of similarity between the document structure and the pre-defined research structure; Determine the proportion of original content in the cited literature to the original content in other cited literature; Determine the frequency of the pre-defined keywords in the literature; the pre-defined keywords are words related to the research type of the literature. The document type score is obtained by weighting and summing the similarity, originality ratio, and frequency. Determine whether the document type score is greater than the preset score; If yes, then the target document type is determined to be a research type; if not, then the target document type is determined to be a review type.

[0098] As an optional implementation, the research field is Earth Science; correspondingly, the constraints in the second prompt word project include: The professional review will consider at least one of the following dimensions: the degree of conformity between the reasoning thinking in the literature and the reasoning thinking in the field of earth sciences; the time scale used in the literature and its fit with the research question; the spatial scale used in the literature and its fit with the research question; the appropriateness of the content of the literature to the field of earth sciences; and the confidence level of the content of the literature. The literature is reviewed and analyzed according to professional review dimensions to obtain the second review text; the second review text includes at least one of the following: the strength analysis text, the limitation analysis text, and the improvement suggestions of the literature in the corresponding professional review dimension; The second review texts corresponding to each professional review dimension are compiled to obtain the second review report.

[0099] As an optional embodiment, the evaluation module 405 includes: The first scoring module is used to input the literature, the first review report and the research report to be evaluated into the general model based on the third prompt word engineering, and to obtain the first scores of the general model based on the literature and the first review report, respectively, from each general review dimension of the research report to be evaluated. The second scoring module is used to input the literature, the second review report and the research report to be evaluated into the domain big model based on the fourth prompt word project. The domain big model then evaluates the research report to be evaluated from each professional review dimension based on the literature and the second review report. The report scoring module is used to determine the report score based on each first score and each second score; The results determination module is used to determine the evaluation results, which characterize the report generation ability of the agent under evaluation, based on the report score.

[0100] As an optional implementation, the constraints in the third prompt word project include: Extract the first review text corresponding to the general review dimensions from the first review report; the first review text includes the text analyzing the strengths, limitations, and improvement suggestions of the literature in the general review dimensions. Using the first review text as a basis for comparison, the consistency of the performance of the research report to be evaluated and the literature in the general review dimensions is compared; Based on the degree of consistency, a first score that is positively correlated with the degree of consistency is determined.

[0101] As an optional embodiment, the report scoring determination module includes: The average score determination module is used to determine the first average score of each first score and the second average score of each second score. The mapping relationship acquisition module is used to acquire pre-set mapping relationships; the mapping relationship is the correspondence between review dimensions and scoring weights; review dimensions include general review dimensions and professional review dimensions; The scoring weight determination module is used to determine the general scoring weights corresponding to the general review dimensions and the professional scoring weights corresponding to the professional review dimensions based on the mapping relationship. The scoring submodule is used to use the general scoring weight as the weight corresponding to the first average score, and the professional scoring weight as the weight corresponding to the second average score; the first average score and the second average score are weighted and summed to obtain the report score.

[0102] Please refer to Figure 5 , Figure 5 A schematic diagram of a storage medium provided in this application is shown. The device includes: Memory 501 is used to store computer programs; Processor 502 is used to implement the steps of any of the above-mentioned intelligent agent evaluation methods when executing a computer program.

[0103] For a detailed description of the electronic device provided in this application, please refer to the embodiments of the evaluation method for the intelligent agent described above; this application will not repeat the details here.

[0104] This application also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described intelligent agent evaluation methods.

[0105] For a detailed description of the storage medium provided in this application, please refer to the embodiments of the above-described evaluation method for intelligent agents; this application will not repeat the details here.

[0106] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0107] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0108] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for evaluating an intelligent agent, characterized in that, The method includes: Obtain the literature and determine the research questions and research areas of the literature; The research question is input into the agent to be evaluated, and the research report to be evaluated is generated by the agent based on the research question. Based on the first prompt word project, the document is reviewed from the preset general review dimensions using a general large model to obtain the first review report; Based on the second prompt word project, the literature is reviewed by a domain-wide model according to preset professional review dimensions to obtain a second review report; the domain-wide model is a model obtained by training a pre-trained large model using a corpus of the research field; the professional review dimensions are dimensions related to the research field. Based on the aforementioned literature, the first review report, and the second review report, a report score for the research report to be evaluated is determined; and based on the report score, an evaluation result is determined to characterize the report generation capability of the agent to be evaluated.

2. The method for evaluating an intelligent agent as described in claim 1, characterized in that, The constraints in the first prompt word project include: Based on the content and structure of the document, the target document type is determined; the target document type is either a review or a research document. If the target document type is the review type, then at least one of the following will be used as the general review dimension: document structure, document content confidence, and coverage of the research field. If the target document type is the research type, then at least one of the following will be used as the general review dimension: document structure, document content confidence, and degree of difference from existing documents. The literature is reviewed and analyzed according to the general review dimensions to obtain a first review text; wherein, the first review text includes at least one of the following: the strength analysis text, the limitation analysis text, and the improvement suggestions of the literature in the corresponding general review dimensions; The first review report is obtained by summarizing the first review texts corresponding to each of the general review dimensions.

3. The method for evaluating intelligent agents as described in claim 2, characterized in that, Based on the content and structure of the document, the target document type corresponding to the document is determined, including: Identify the chapter titles of the document and assemble the chapter titles into the document's structure; Determine the similarity between the document structure and the preset research structure; the chapter titles in the preset research structure include introduction, methods, results, and discussion; the similarity is used to characterize the degree of similarity between the document structure and the preset research structure; Determine the proportion of original content in the cited literature relative to the original content in other cited literature; Determine the frequency of occurrence of preset keywords in the literature; the preset keywords are words related to the literature of the research type. The document type score is obtained by weighted summing of the similarity, the originality ratio, and the frequency. Determine whether the score for the document type is greater than the preset score; If yes, then the target document type is determined to be the research type; if no, then the target document type is determined to be the review type.

4. The method for evaluating an intelligent agent as described in claim 1, characterized in that, The research field is Earth science; Correspondingly, the constraints in the second prompt word project include: The professional review dimensions include at least one of the following: the degree of conformity between the reasoning thinking of the literature and the reasoning thinking in the field of earth sciences; the time scale adopted by the literature and its suitability to the research question; the spatial scale adopted by the literature and its suitability to the research question; the relevance of the literature content to the field of earth sciences; and the confidence level of the literature content. The literature is reviewed and analyzed according to the aforementioned professional review dimensions to obtain a second review text; wherein, the second review text includes at least one of the following: an analysis text of the literature's strengths, an analysis text of its limitations, and suggestions for improvement in the corresponding professional review dimensions; The second review report is obtained by summarizing the second review texts corresponding to each of the aforementioned professional review dimensions.

5. The method for evaluating an intelligent agent as described in claim 1, characterized in that, Based on the aforementioned literature, the first review report, and the second review report, the report score of the research report to be evaluated is determined, including: Based on the third prompt word project, the literature, the first review report and the research report to be evaluated are input into the general model to obtain the first score of the general model for evaluating the research report to be evaluated from each of the general review dimensions, based on the literature and the first review report. Based on the fourth prompt word project, the literature, the second review report and the research report to be evaluated are input into the domain big model to obtain the second score of the domain big model based on the literature and the second review report, respectively, from each of the professional review dimensions. The report score is determined based on each of the first scores and each of the second scores.

6. The method for evaluating intelligent agents as described in claim 5, characterized in that, The constraints in the third prompt word project include: Extract the first review text corresponding to the general review dimension from the first review report; the first review text includes the strength analysis text, the limitation analysis text, and improvement suggestions of the literature in the general review dimension. Using the first review text as a basis for comparison, compare the degree of consistency between the research report to be evaluated and the literature in the general review dimensions; Based on the degree of consistency, the first score is determined to be positively correlated with the degree of consistency.

7. The method for evaluating intelligent agents as described in claim 5, characterized in that, The report score is determined based on each of the first scores and each of the second scores, including: Determine the first average score for each of the first scores and the second average score for each of the second scores; Obtain a pre-set mapping relationship; the mapping relationship is the correspondence between review dimensions and scoring weights; the review dimensions include the general review dimensions and the professional review dimensions. Based on the mapping relationship, determine the general scoring weights corresponding to the general review dimensions and the professional scoring weights corresponding to the professional review dimensions. The general scoring weight is used as the weight corresponding to the first average score, and the professional scoring weight is used as the weight corresponding to the second average score; The report score is obtained by weighted summing of the first average score and the second average score.

8. An evaluation device for an intelligent agent, characterized in that, The device includes: The acquisition module is used to acquire literature and determine the research questions and research fields of the literature; The report generation module is used to input the research question into the agent to be evaluated and obtain the research report to be evaluated generated by the agent to be evaluated based on the research question; The first review module is used to review the document based on the first prompt word project and through a general model from preset general review dimensions to obtain a first review report; The second review module is used to review the literature based on the second prompt word project and according to the preset professional review dimensions using a domain-wide model, and to obtain a second review report; the domain-wide model is a model obtained by training a pre-trained large model using a corpus of the research field; the professional review dimensions are dimensions related to the research field. The evaluation module is used to determine the report score of the research report to be evaluated based on the literature, the first review report, and the second review report; and to determine the evaluation result that characterizes the report generation ability of the agent to be evaluated based on the report score.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the evaluation method for an intelligent agent as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the evaluation method for an intelligent agent as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Scientific field-oriented multi-modal corpus data construction method and device

    CN118170933A

  • Intelligent agent evaluation method and device, electronic equipment, storage medium and program

    CN119621588A

  • Metro field large language model evaluation method and system

    CN120163142A

  • Intelligent agent-based automatic verification and rule matching system for insurance-waiting evaluation report

    CN121304091A

  • Dynamic evaluation of language model prompts for model selection and output validation and methods and systems of the same

    US12147513B1

Cited By

  • Geoscience big data oriented agent evaluation method, device and product

    CN122220476A