Computing device, method, computer-readable storage medium, and computer program product for literature meta-analysis

By designing a computing device for multi-agents to work collaboratively in literature meta-analysis, using large language models and multi-modal large models to automate literature evaluation, data extraction, integration and inspection, the problem of lack of full-process automation in the existing technology is solved, and efficiency and reliability of results are improved.

CN119830891BActive Publication Date: 2025-06-13SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510308312.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing literature meta-analysis methods based on deep learning lack the full process automation design, resulting in manual intervention in each link, and the efficiency and result reliability are limited.

Method used

By designing a computing device that includes multiple agents, they use them to call large language models and multimodal large models respectively to realize the automated process of literature evaluation, data extraction, data integration and data inspection.

Benefits of technology

The literature meta-analysis process is automated, the overall efficiency and accuracy of results are improved, manual intervention is reduced, and the analysis is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830891B_ABST
    Figure CN119830891B_ABST
Patent Text Reader

Abstract

The present invention relates to a computer system using a computer model, and discloses a computing device, a method, a computer-readable storage medium, and a computer program product for literature meta-analysis. A computing system for document meta-analysis includes computing resources and a plurality of agents, which are executed by the computing resources. The plurality of agents includes a first agent, a second agent, a third agent, and a fourth agent. The first agent is configured to perform literature evaluation. The second agent is configured to perform data extraction. The third agent is configured to perform data integration. The fourth agent is configured to perform data checking. Among them, a report of the literature meta-analysis is generated based on a single set of data that has passed the checking. The computing device according to the present invention overcomes the limitation that the literature meta-analysis technology based on deep learning processes each link of the literature meta-analysis dispersedly, and improves the automation degree and efficiency of the entire processing flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to computer systems utilizing computational models, and more particularly to computational devices, methods, computer-readable storage media, and computer program products for literature meta-analysis. Background Art

[0002] Literature meta-analysis is the re-statistical analysis of a large number of existing literatures to extract the relationships between variables in the literatures, so as to discover new relationships or knowledge. Literature meta-analysis is widely applied in scientific research fields such as medicine, agriculture, ecological environment, etc. Since the analyzed literatures are all independent, the relationships obtained through literature meta-analysis are more scientific and accurate, thus being able to make up for the deficiencies of single studies and discover potential laws that cannot be found in single studies.

[0003] The implementation of literature meta-analysis generally includes the following key steps: First, clarify the research question and set clear research objectives and hypotheses; Second, perform a systematic literature search to ensure that all relevant studies are covered; Subsequently, screen out eligible literatures according to preset criteria; Then, extract the key inputs in the studies, such as sample size and effect size, etc.; On this basis, select an appropriate statistical model, usually determined by the heterogeneity among studies to use a fixed model or a random effects model; Finally, conduct a sensitivity analysis to test the robustness of the results and write a detailed result interpretation and report.

[0004] Commonly used literature meta-analysis methods can be divided into two categories: one is the traditional quantitative meta-analysis method, and the other is modern data integration technology. Traditional methods (such as fixed effect models and random effects models) evaluate the overall effect by calculating the weighted average effect size of each study. Modern methods are more flexible and usually adopt Bayesian methods or machine learning techniques to handle more complex consistencies and data structures.

[0005] Traditional literature meta-analysis methods are very time-consuming and laborious in the process of literature collection and screening, and a large amount of manpower needs to be invested to sort out and evaluate literature data one by one. For the already screened literatures, the process of extracting and sorting out data is equally cumbersome, thus further increasing the workload. Secondly, there may be significant differences in the experimental methods and standards adopted in different literatures, and simply integrating literatures from different sources may lead to wrong conclusions. The lack of a unified evaluation standard makes the comparability of different research results low, thus affecting the reliability of meta-analysis. Finally, if the quality of the analyzed literatures is low, then the final analysis conclusions will inevitably be affected, resulting in inaccurate or misleading results. These problems make the traditional literature meta-analysis technology need to be improved in terms of efficiency and result reliability, and there is an urgent need for more advanced automated and standardized solutions to optimize this process.

[0006] In recent years, the application of deep learning technology in literature meta-analysis has gradually attracted attention, especially the use of large language models (LLMs) to analyze and summarize research literature. By systematically reviewing existing literature and designing corresponding analysis models, the analysis efficiency and the accuracy of results can be improved, thus achieving a comprehensive evaluation and data extraction of existing literature.

[0007] Literature meta-analysis based on deep learning technology mainly relies on the text analysis method of large language models. Its typical applications include using simple prompts and application programming interface (API) calls to perform specific steps in literature meta-analysis, such as literature screening and data extraction. These methods usually focus on a certain step in literature meta-analysis, such as automatically evaluating the quality of literature or extracting tabular data through large language models, so as to improve work efficiency.

[0008] Existing literature meta-analysis methods based on deep learning technology have at least two levels of limitations.

[0009] First, each step of literature meta-analysis still requires separate manual intervention to be completed separately, lacking an automated design for the entire process of literature meta-analysis. For example, some studies focus on using large language models to accelerate the literature screening process to cope with the high labor cost in the literature screening step; while other studies focus on the extraction of tabular data, trying to simplify the data collation steps through large language models. This fragmented processing method cannot achieve full automation and still relies on manual intervention in multiple steps, restricting the overall efficiency and accuracy of literature meta-analysis.

[0010] Second, there are still some limitations in each step involved in the literature meta-analysis process itself.

[0011] For example, the literature meta-analysis process involves a literature evaluation step. Deep learning-based methods have limitations in literature evaluation. First, due to the context window length limitation of large language models, these models often cannot process the complete content of the literature. Traditional large language model evaluation methods usually only read the abstract part of the literature, and this incomplete reading makes the evaluation of the literature one-sided, which may miss key experimental results and important conclusions. Second, there is a "hallucination" phenomenon in large language models, that is, the scores and evaluations output by the models may not be real. This situation may lead to misjudgment of the literature quality, thus affecting subsequent research decisions and the quality of literature screening.

[0012] For example, the process of literature meta-analysis involves literature data mining and processing. Literature data mining and processing techniques also face multiple challenges, especially in the application of converting documents in a specific format (e.g., Portable Document Format (PDF)) to text based on Optical Character Recognition (OCR). Some OCR models cannot effectively adapt to the diverse data formats in different literatures, resulting in poor recognition effects. For complex tables, especially those containing multiple sub-tables, OCR technology often makes misidentifications, affecting the accuracy of data. In addition, when simply extracting table data, information ambiguity is likely to occur. This is mainly because the information contained in the table may be incomplete, especially when different variables use abbreviations that are not clearly explained, leading to misunderstandings and incorrect analyses in subsequent data mining processes.

[0013] There is a need in the art for a literature meta-analysis method that improves in at least one of the above aspects. Summary of the Invention

[0014] The present invention is provided to further improve the literature meta-analysis technology based on deep learning.

[0015] One aspect of the present invention provides a computing device for literature meta-analysis, including: computing resources; and a plurality of agents, which are executed by the computing resources, and the plurality of agents include: a first agent configured to call a first large language model to: receive a plurality of documents and a first user input, the plurality of documents including literatures, and the first user input including user requirements; score the literatures corresponding to each document based on the user requirements; screen the plurality of documents based on the scoring results to determine a subset of documents; a second agent configured to call a multimodal large model to: receive the subset of documents; convert the multimodal data in each document in the subset of documents into multiple sets of data in tabular form, each set of data including entries and data for the entries; a third agent configured to call a second large language model to: receive the multiple sets of data and a second user input, the second user input including entries that the user is concerned about; integrate the multiple sets of data into a single set of data in tabular form based on the entries that the user is concerned about; and a fourth agent configured to call a third large language model to: receive the multiple sets of data and the single set of data; check the single set of data against the multiple sets of data based on at least one of user interest relevance and accuracy; and iteratively provide an indication to update the single set of data based on the check results until the updated single set of data passes the check, wherein a report of literature meta-analysis is generated based on the single set of data that passes the check.

[0016] The computing device as described above, wherein the multiple documents include a first number of documents, and the first agent is configured to invoke the first large language model to: determine an independent score for each of the first number of documents based on the user requirement, and determine a relative score for each document separately for the first number of documents as a whole; determine a comprehensive score for each document based on the independent score and the relative score of each document in the first number of documents; and screen the multiple documents based on the comprehensive score of each document in the first number of documents.

[0017] The computing device as described in any of the above, wherein the user requirement includes at least one of topic relevance, innovativeness, and feasibility, and / or is characterized in that the first user input requires the first large language model to output the reasons for the score.

[0018] The computing device as described in any of the above, wherein the first agent is configured to invoke the first large language model to: read a part at a predetermined position of each document to score each document, wherein the part at the predetermined position includes an abstract of the literature and a part of the body of the literature, and the amount of text in the part at the predetermined position conforms to the context window length limit of the first large language model.

[0019] The computing device as described in any of the above, wherein the second agent is configured to invoke the multi-modal large model to: generate a hierarchical data description based on the multi-modal data in each document, wherein the data descriptions at different levels describe the multi-modal data with different degrees of structuring; and generate the multiple sets of data based on the hierarchical data description.

[0020] The computing device as described in any of the above, wherein the hierarchical data description includes a description of the overall content of the multi-modal data.

[0021] The computing device as described in any of the above, wherein the hierarchical data description includes a dictionary-form data description, and the dictionary-form data description includes a description of row entries and / or column entries associated with a table form.

[0022] The computing device as described in any of the above, wherein the third agent is configured to invoke the second large language model to: analyze the semantic similarity of the description of the row entries and / or the description of the column entries to perform a consistency check on the multiple sets of data, wherein the description of the row entries and / or the description of the column entries includes the full name and abbreviation of variables, and the multiple sets of data are integrated based on the result of the consistency check.

[0023] The computing device according to any one of the above, wherein the multimodal data includes pictures, and the second agent is configured to call the multimodal large model to: output the description content of the picture point by point.

[0024] The computing device according to any one of the above, wherein the third agent is configured to call the second large language model to: determine whether each set of data in the multiple sets of data includes target data associated with the item that the user is concerned about; and based on the result of the determination, integrate the sets of data in the multiple sets of data that include the target data into the single set of data.

[0025] The computing device according to any one of the above, wherein the third agent is configured to call the second large language model to: output the source of each data in the single set of data, and the source is used to indicate which set of data in the multiple sets of data and which document in the multiple documents each data comes from.

[0026] The computing device according to any one of the above, wherein the fourth agent is configured to call the third large language model to: compare each data in the single set of data with the corresponding data in the multiple sets of data; and based on the result of the comparison, determine the accuracy of the data.

[0027] The computing device according to any one of the above, wherein the fourth agent is configured to call the third large language model to: determine an overall score based on at least one of the user interest relevance and the accuracy; compare the overall score with a preset threshold; and based on the result of the comparison, iteratively provide the indication until the overall score reaches the threshold.

[0028] The computing device according to any one of the above, wherein the fourth agent is configured to call the third large language model to: provide the indication in response to the result of the comparison indicating that the overall score is lower than the threshold; and not provide the indication in response to the result of the comparison indicating that the overall score is greater than or equal to the threshold.

[0029] The computing device according to any one of the above, wherein the multiple agents further include: a fifth agent configured to: receive multiple original documents including literature; perform page parsing on each of the multiple original documents to determine multiple parsing elements; and merge the multiple parsing elements to form each parsed document, wherein each formed parsed document constitutes the multiple documents.

[0030] The computing device according to any one of the above, wherein the multiple original documents include charts, and the fifth agent is configured to: call a multimodal language model to convert the charts into corresponding parsing elements in the multiple parsing elements.

[0031] The computing device according to any one of the above, wherein the chart includes a statistical chart and a table, the multimodal language model includes a first multimodal language model and a second multimodal language model, and the fifth agent is configured to: call the first multimodal language model to convert the statistical chart into a corresponding parsing element among the plurality of parsing elements; and call the second multimodal language model to convert the table into a corresponding parsing element among the plurality of parsing elements.

[0032] The computing device according to any one of the above, wherein the plurality of original documents include text and formulas, and the fifth agent is configured to: use optical character recognition technology to convert the text into a corresponding parsing element among the plurality of parsing elements; call a fourth large language model to convert the formulas into a corresponding parsing element among the plurality of parsing elements.

[0033] The computing device according to any one of the above, wherein the fifth agent is configured for each document to be formed: input the plurality of parsing elements into the page of the document in a plurality of candidate layouts respectively; record the size and position in the page of each parsing element for each candidate layout, and calculate the fitness under the current candidate layout based on the area of the region occupied by the plurality of parsing elements in the page and the total area of the page; and merge the plurality of parsing elements with the candidate layout having the maximum fitness.

[0034] Another aspect of the present invention provides a method for literature meta-analysis, comprising the following steps:

[0035] S1: The first agent invokes the first large language model to perform a literature review. The step S1 includes: S11: Receiving multiple documents and a first user input, where the multiple documents include literature and the first user input includes user requirements; S12: Scoring the literature corresponding to each document based on the user requirements; S13: Screening the multiple documents based on the scoring results to determine a subset of documents; S2: The second agent invokes a multimodal large model to perform data extraction. The step S2 includes: S21: Receiving the subset of documents; S22: Converting the multimodal data in each document in the subset of documents into multiple sets of data in tabular form, where each set of data includes entries and data for the entries; S3: The third agent invokes the second large language model to perform data integration. The step S3 includes: S31: Receiving the multiple sets of data and a second user input, where the second user input includes entries that the user is concerned about; S32: Integrating the multiple sets of data into a single set of data in tabular form based on the entries that the user is concerned about; S4: The fourth agent invokes the third large language model to perform data checking. The step S4 includes: S41: Receiving the multiple sets of data and the single set of data; S42: Checking the single set of data against the multiple sets of data based on at least one of user interest relevance and accuracy; S43: Based on the checking results, iteratively providing an indication to update the single set of data until the updated single set of data passes the check. Among them, a report on literature meta-analysis is generated based on the single set of data that has passed the check.

[0036] For the method as described above, the step S22 includes the following steps: S221: Generating a hierarchical data description based on the multimodal data in each document, where the data description at each level describes the multimodal data with different degrees of structural hierarchy; and S222: Generating the multiple sets of data based on the hierarchical data description.

[0037] For the method as described in any of the above, the step S32 includes the following steps: S321: Determining whether each set of data in the multiple sets of data includes target data associated with the entries that the user is concerned about; and S322: Integrating the sets of data in the multiple sets of data that include the target data into the single set of data.

[0038] For the method as described in any of the above, the step S42 includes the following steps: S421: Determining an overall score based on at least one of the user interest relevance and accuracy; S422: Comparing the overall score with a preset threshold; and S423: Based on the comparison results, iteratively providing the indication until the overall score reaches the threshold.

[0039] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0040] Another aspect of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0041] The computing device and method according to the present invention overcome the limitations of processing each link separately in the literature meta-analysis based on deep learning technology, and use multiple agents to form a workflow capable of automatically executing a complete document meta-analysis, increasing the degree of automation of the entire process and improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The embodiments of the present invention are described in conjunction with the accompanying drawings.

[0043] Figure 1 is a block diagram of a computing device for literature meta-analysis according to some embodiments of the present invention.

[0044] Figure 2 shows a first example chart of data extraction processed using the literature meta-analysis technology according to some embodiments of the present invention.

[0045] Figure 3 shows a second example chart of data extraction processed using the literature meta-analysis technology according to some embodiments of the present invention.

[0046] Figure 4 is a flowchart of a processing pipeline for performing literature meta-analysis according to some embodiments of the present invention.

[0047] Figure 5 is a flowchart of a method for literature meta-analysis according to some embodiments of the present invention.

[0048] Figure 6 is a flowchart of a first process associated with the method for literature meta-analysis according to some embodiments of the present invention.

[0049] Figure 7 is a flowchart of a second process associated with the method for literature meta-analysis according to some embodiments of the present invention.

[0050] Figure 8 is a flowchart of a third process associated with the method for literature meta-analysis according to some embodiments of the present invention.

[0051] Figure 9 is a flowchart of a fourth process associated with the method for literature meta-analysis according to some embodiments of the present invention.

[0052] Figure 10 is a flowchart of a fifth process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0053] Figure 11 is a flowchart of a sixth process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0054] Figure 12 is a flowchart of a seventh process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0055] Figure 13 is a flowchart of an eighth process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0056] Figure 14 is a flowchart of a ninth process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0057] Figure 15 is a flowchart of a tenth process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0058] Figure 16 is a flowchart of an eleventh process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0059] Figure 17 is a flowchart of a twelfth process associated with a method for literature meta - analysis according to some embodiments of the present invention.

[0060] Figure 18 is a block diagram of a computer - readable storage medium according to some embodiments of the present invention.

[0061] Figure 19 is a block diagram of a computer program product according to some embodiments of the present invention. Detailed Description of the Invention

[0062] In the present application, the term "agent" refers to an entity that can perceive the environment and take actions to execute specific goals. An agent mainly refers to software code. The agent can be executed by the computing resources of a computing device. The agent can call a corresponding model through an API interface and call corresponding tools (such as a PDF reader, a Python interpreter, a calculator, etc.) to interact with various forms of input or implement corresponding functions.

[0063] In this application, ordinal numbers such as "first", "second", "third", etc. are used to distinguish different instances of objects with the same name. The ordinal numbers "first", "second", "third", etc. do not represent the relative order of the indicated objects in terms of time, space, sorting, and other aspects.

[0064] According to one aspect of the present invention, there is provided a computing device for literature meta-analysis.

[0065] Figure 1 It is a block diagram of a computing device 100 for literature meta-analysis according to some embodiments of the present invention.

[0066] The computing device 100 can be a local or remote computer, server, etc. The computing device 100 may include computing resources 110 and multiple agents. The multiple agents can be executed by the computing resources 110 to cause the computing resources to perform corresponding operations.

[0067] In some embodiments, the computing resources may include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and various other processing units or cores (e.g., arithmetic logic unit, integer unit, floating-point unit, tensor unit, ray tracing core, etc.).

[0068] In some embodiments, each of the multiple agents can call a corresponding model through an API. As an example, the model may include a large language model, a multi-modal model, a multi-modal language model, etc. In some embodiments, multiple models that can be called by the multiple agents may be deployed locally on the computing device 100. In some embodiments, multiple models that can be called by the multiple agents may be deployed remotely from the computing device 100, for example, deployed in the cloud. In some embodiments, some of the multiple models that can be called by the multiple agents may be deployed locally on the computing device 100, and some other models may be deployed remotely from the computing device 100. In some embodiments, each of the multiple agents can call various tools to interact with various forms of input or implement corresponding functions. For example, the agent can call a PDF reader to read a PDF document. The agent can call a python interpreter to interact with python code. The agent can call a calculator to implement a calculation function.

[0069] The multiple agents may include a first agent 121. The first agent 121 may be configured to call a first large language model to perform operations related to literature evaluation.

[0070] The first agent 121 can be configured to receive multiple documents and a first user input, where the multiple documents can include literature, and the first user input can include user requirements. The first agent 121 can be configured to score the literature corresponding to each document based on the user requirements. The first agent 121 can be configured to screen the multiple documents based on the scoring results to determine a subset of the documents. Through the literature evaluation performed by the first agent 121, the literature that meets the user requirements is screened out for subsequent further analysis, which improves the quality of the literature to be analyzed and helps to improve the reliability of the literature meta-analysis.

[0071] The multiple agents can include a second agent 122. The second agent 122 can be configured to call a multimodal large model to perform data extraction from the documents.

[0072] The second agent 122 can be configured to receive the subset of the documents (determined by the first agent 121). The second agent 122 can be configured to convert the multimodal data in the subset of the documents into multiple sets of data in tabular form, where each set of data can include entries and data for the entries. As an example, each set of data can correspond to the data of a table, which can include a table header (which contains the entries) and a table body (which contains the data). The second agent 122 can extract the multimodal data in the documents screened by the first agent 121 into multiple sets of data in tabular form, so as to provide structured data for the downstream agents to facilitate further processing.

[0073] The multiple agents can include a third agent 123. The third agent 123 can be configured to call a second large language model to perform data integration of the multiple sets of data in tabular form.

[0074] The third agent 123 can be configured to receive the multiple sets of data (in tabular form converted by the second agent 122) and a second user input, where the second user input can include the entries that the user is concerned about. The third agent 123 can be configured to integrate the multiple sets of data into a single set of data in tabular form based on the entries that the user is concerned about. The third agent integrates the multiple sets of data in tabular form provided by the upstream second agent 122 based on the second user input to provide structured data in a single table for downstream analysis.

[0075] The multiple agents can include a fourth agent 124. The fourth agent 124 can be configured to call a fourth large language model to perform data checking on the single set of data in tabular form.

[0076] In some embodiments, the fourth agent 124 may be configured to receive multiple sets of data (in tabular form provided by the second agent 122) and a single set of data (in tabular form synthesized by the third agent 123). The fourth agent 124 may be configured to check the single set of data against the multiple sets of data based on at least one of user interest relevance and accuracy. As an example, user interest relevance may indicate whether the relevant content belongs to the fields, aspects, etc. that the user is concerned about and / or the degree of association between the relevant content and the fields, aspects, etc. that the user is concerned about. The fourth agent 124 may be configured to iteratively provide an indication to update the single set of data based on the detected result until the updated single set of data passes the check. Based on the single set of data that has passed the check, a report of the literature meta-analysis may be generated, for example, by a downstream agent. The fourth agent 124 checks the single set of data (synthesized data) against the multiple sets of data (the source of the synthesized data) and iteratively provides an indication to update the single set of data based on the detected result. This indication may be fed back to the third agent 123 located upstream. The third agent 123 may modify one or more data based on the received indication to generate an updated single set of data. Through the above iterative and feedback process, the accuracy and reliability of the literature meta-analysis result are further improved.

[0077] Some embodiments integrate the first agent 121, the second agent 122, the third agent 123, and the fourth agent 124 to integrate the literature evaluation phase, data extraction phase, data integration phase, and data checking phase in the literature meta-analysis process, providing a solution for efficient automated execution of the entire process of literature meta-analysis. During the execution of the literature meta-analysis, the results generated or output by the upstream agents all contribute to improving the accuracy and reliability of the data to be processed or analyzed by the downstream agents, and the downstream agents can also iteratively check the synthesized data to iteratively improve the synthesized data through feedback, thereby increasing the accuracy and reliability of the literature meta-analysis.

[0078] Regarding the first agent 121, its implementation details are further described.

[0079] In some embodiments, the multiple documents may be multiple documents downloaded by searching a literature database. In some embodiments, the multiple documents may be documents obtained after document processing of the downloaded multiple documents (for example, as described below in connection with the fifth agent 125).

[0080] In some embodiments, the user requirements may include at least one of topic relevance, innovativeness, and feasibility. As an example, topic relevance can be judged based on factors such as whether the topic of the literature belongs to the same field as the topic of interest to the user, whether there is a direct connection, whether it is relevant, and how many requirements of the user are met. As an example, innovativeness can be judged based on whether the literature proposes new methods or new discoveries, and the significance or value of the proposed new methods or new discoveries. As an example, feasibility can be determined based on whether the literature contains experimental or data support, the integrity, sufficiency, and reproducibility of the experiments and data.

[0081] In the prior art, traditional corpus analysis techniques are often used to evaluate the quality of literature, which may have several limitations. On the one hand, traditional corpus analysis relies on word frequency analysis, and word frequency analysis cannot accurately reflect the novelty and innovation of a paper because even if many names of new methods and new problems appear in the paper, these words are likely to be over-packaging of existing names, and the literature itself does not propose real new technologies. In addition, word frequency analysis may be accidental for shorter literatures. On the other hand, traditional corpus analysis relies on the number of literature citations. The number of literature citations is an important factor in measuring the influence of literature, but the number of citations is affected by the publication time and the field of the literature. Therefore, the literature scores obtained by analyzing citations cannot be compared horizontally. For example, a literature in a popular field published in 2010 and a paper in a less popular field published in 2020 may have a much larger number of citations for the former than the latter, but this does not directly mean that the former is more useful than the latter.

[0082] Some embodiments can overcome the above limitations of traditional corpus analysis to a certain extent by selecting at least one of topic relevance, innovativeness, and feasibility as the criteria for literature evaluation.

[0083] In addition, different evaluation dimensions and / or evaluation criteria for literature review can be selected according to different fields. For example, in the medical field, the rigor of experimental design may be more emphasized. Accordingly, feasibility is more important for literature evaluation and may have a higher weight in literature evaluation; optionally, the weights of other evaluation dimensions can be reduced or other evaluation dimensions can be ignored. For example, in the social science field, theoretical innovation may be more emphasized. Accordingly, it may have other weights in literature evaluation; optionally, the weights of other evaluation dimensions can be reduced or other evaluation dimensions can be ignored.

[0084] In some embodiments, scores can be separately given based on two or three of topic relevance, innovativeness, and feasibility, and a comprehensive score can be determined based on the individual scores.

[0085] In some embodiments, the first user input may require the first large language model to output the reasons for the score.

[0086] In some embodiments, the first user input may include a prompt input. The following is an example of a prompt input that can serve as the first user input:

[0087] ==========;

[0088] Prompt Example 1:

[0089] You are a professional reviewer, and your professional field is atmospheric science. Please review the following paper and assign it a score.

[0090] You need to evaluate the response from the following dimensions: topic relevance, innovation, and feasibility.

[0091] When replying, please strictly follow the following rules:

[0092] 1. Evaluate the paper from different dimensions, point out its advantages or disadvantages in each dimension, and assign a score from 1 to 10 for each dimension.

[0093] 2. Finally, based on the evaluations of each dimension, provide an overall score from 1 to 10 for the paper.

[0094] Generally speaking, the higher the quality of the paper and the more it meets the user's requirements, the higher the score. Papers that do not meet the user's requirements will receive a lower score.

[0095] Scoring rules:

[0096] Topic relevance:

[0097] When the paper is not relevant to the user's needs, the score is 1 - 2.

[0098] The paper is in the same field as the topic the user is interested in but has no direct connection, and the score is 3 - 4.

[0099] The paper is relevant to the topic the user is concerned about but does not meet specific requirements (such as time, location, method), and the score is 5 - 6.

[0100] The paper is closely related to the topic the user is interested in and meets most requirements (such as time, location, method), and the score is 7 - 8.

[0101] The paper is closely related to the topic the user is interested in and meets all requirements (such as time, location, method), and the score is 9 - 10.

[0102] Innovation:

[0103] The paper does not propose new methods or discoveries, and the score is 1 - 2.

[0104] The innovation proposed by the paper is very small or incremental, and the score is 3 - 4.

[0105] The paper presents a new method or discovery, but its significance is limited, with a score of 5-6.

[0106] The new method or discovery presented in the paper is valuable, with a score of 7-8.

[0107] A score of 9-10 indicates that the paper presents a very important new method or discovery, having a great impact on the relevant field.

[0108] Feasibility:

[0109] A score of 1-2 indicates that the paper has no experimental or data support.

[0110] A score of 3-4 indicates that the paper has little experimental or data support.

[0111] A score of 5-6 indicates that there are some experiments and data in the paper, but they are incomplete and insufficient.

[0112] A score of 7-8 indicates that the paper provides sufficient experiments and data, and they are reproducible.

[0113] A score of 9-10 indicates that the experiments and data in the paper are very complete, including experimental details and data descriptions, and can be used as the basis for subsequent work.

[0114] Total score:

[0115] A score of 1-2 indicates that the paper is not worth reading.

[0116] A score of 3-4 indicates that only the abstract of the paper needs to be read, without the need to read it carefully.

[0117] A score of 5-6 indicates that the paper is worth a cursory reading.

[0118] A score of 7-8 indicates that the paper needs to be read carefully.

[0119] A score of 9-10 indicates that the paper is very valuable and worth reading repeatedly.

[0120] ==========;

[0121] The following is another example of a prompt that can be used as the first user input:

[0122] ==========;

[0123] Prompt example 2:

[0124] You are an expert in the agricultural field, specializing in literature analysis.

[0125] Please determine whether the following paper is relevant to the topic of interest to the user.

[0126] When replying, please strictly follow the following rules:

[0127] 1. For each paper, use 1 or 0 to indicate whether it is relevant to the topic of interest, where 1 means yes and 0 means no.

[0128] Please do not give other responses.

[0129] 2. When judging relevance, strictly follow each requirement put forward by the user in the topic of interest.

[0130] Including location, time, etc. If not met, judge it as not relevant, that is, 0.

[0131] 3. Only papers that strictly meet all the user's requirements can be judged as relevant, that is, 1.

[0132] 4. Please express the final answer in list form and do not reply with other words.

[0133] Example:

[0134] [1, 0, 1, 0, 1,...];

[0135] ==========;

[0136] In some embodiments, the first agent 121 may be configured to read a portion of a predetermined location of each document to screen a plurality of documents, wherein the portion of the predetermined location includes an abstract of the literature and includes a portion of the body of the literature, and the amount of text in the portion of the predetermined location meets the context window length limit of the first large language model. Limited by the context length limit of the large language model, too much text cannot be input into the window of the large language model for evaluation. To meet the requirements of the context window length limit, some existing literature analysis methods only input the abstract portion of the literature into the dialogue window of the large language model for analysis. However, the analysis based only on the abstract of the literature may miss key information in the body part, resulting in an incomplete evaluation result.

[0137] In some embodiments, the first agent 121 may input the abstract of each document and the first paragraph of each chapter into the first large language model. In some embodiments, when inputting the abstract of the document and the first paragraph of each chapter into the first large language model, if it is prompted that the length limit of the context window of the first large language model has been reached, the first agent 121 may instruct the first large language model to stop receiving further input. In some embodiments, after inputting the abstract of the document and the first paragraph of each chapter into the first large language model, if the length limit of the context window of the first large language model has not been reached, the first agent 121 may instruct to continue inputting another part of each chapter (e.g., the second paragraph, the third paragraph,..., the last paragraph, etc.) into the first large language model.

[0138] In some embodiments, by inputting the abstract of a document and a part of the text into a large language model simultaneously, the comprehensiveness of the analyzed corpus and the length limit of the context window of the large language model are taken into account, enabling the large language model to understand the overall content of the document while also understanding the details of the document as much as possible, thereby giving a more accurate and comprehensive evaluation result.

[0139] In some embodiments, the multiple documents may include a first number of documents. The first agent 121 may be configured to determine an independent score for each of the first number of documents based on a user request, and determine a relative score for each document as a whole for the first number of documents. The first agent 121 may be configured to determine a comprehensive score for each document based on the independent score and the relative score of each document in the first number of documents. The first agent may be configured to screen the multiple documents based on the comprehensive score of each document in the first number of documents.

[0140] In some embodiments, the multiple documents may include a first document and a second document. The first agent 121 may be configured to determine a first independent score for the first document, a second independent score for the second document, and a first relative score for the first document and a second relative score for the second document for both the first document and the second document based on a user request. The first agent 121 may be configured to determine a first comprehensive score based on the first independent score and the second relative score. The first agent 121 may be configured to determine a second comprehensive score based on the second independent score and the second relative score. The first agent 121 may be configured to screen the multiple documents based on the first comprehensive score and the second comprehensive score.

[0141] As an example, assume that the first agent 121 may call a first large language model to determine the independent scores of document i and document j respectively according to the following equations (1) and (2) 、 :

[0142] (1);

[0143] (2);

[0144] Where, R i 、R j are the topic relevance scores, I i 、I j are the innovation scores, F i 、F j are the feasibility scores.

[0145] The first intelligent agent 121 can call the first large language model to take document i and document j as inputs together, evaluate them as a whole, and score each document with a number between 0 and 1 to determine their relative scores 、 。

[0146] The first intelligent agent 121 can call the first large language model to determine the comprehensive scores of document i and document j based on their respective independent scores 、 and relative scores 、 respectively according to the following equations (3) and (4): 、 :

[0147] (3);

[0148] (4);

[0149] As an example, when evaluating document i alone, its independent score can be determined as ={8, 9, 6}, and when evaluating document j alone, its independent score can be determined as ={6, 6, 6}. When inputting document i and document j into the first large language model simultaneously, their respective relative scores can be determined as =0.6, =0.9. Thus, the comprehensive score of document i can be determined as ={8×0.6 = 4.8, 9×0.6 = 5.4, 6×0.6 = 3.6}, ={6×0.9 = 5.4, 6×0.9 = 5.4, 6×0.9 = 5.4}. From this example, it can be seen that if evaluating a single document alone, the score of document i is better than that of document j. However, through the mixed evaluation, the comprehensive score of document j is higher than that of document i. This shows that the addition of the mixed evaluation helps to unify the scoring criteria, thereby enabling a more accurate horizontal comparison of each document.

[0150] In some embodiments, the multiple documents may include three or more documents. The first intelligent agent 121 can be configured to determine the comprehensive score of each document in a similar manner and screen the multiple documents based on the comprehensive score of each document.

[0151] In some embodiments, the independent scores of some of the multiple documents can be determined, and the comprehensive scores of another part of the multiple documents can be determined. Finally, the multiple documents can be screened based on the determined independent scores and comprehensive scores.

[0152] For the second intelligent agent 122, its implementation details will be further described.

[0153] In some embodiments, the second agent 122 may be configured to generate a hierarchical data description based on the multimodal data in each document, wherein the data descriptions at different levels describe the multimodal data with different degrees of structuring. As an example, the multimodal data may include text, pictures, and tables. The second agent 122 may be configured to generate multiple sets of data based on the hierarchical data description. By means of the hierarchical data description with data descriptions of different degrees of structuring, it is possible to prevent the extracted data from being ambiguous during the data extraction process, thereby further avoiding errors in the subsequent integration process of the extracted data.

[0154] In some embodiments, the hierarchical data description includes a description of the overall content of the multimodal data. In some embodiments, the hierarchical data description includes a data description in dictionary form, wherein the data description in dictionary form may include a description of row indexes and row entries associated with a table form, and / or the data description in dictionary form may include a description of column indexes and column entries associated with a table form. By providing a data description in dictionary form, it helps downstream agents (e.g., the third agent 123) to integrate tabular data by semantic similarity rather than string matching during the data integration phase, so that more data can be integrated while ensuring accuracy.

[0155] In some embodiments, the second agent 122 may also be configured to generate more hierarchical data descriptions. In some embodiments, the second agent 122 may also be configured to have data descriptions with a higher degree of structuring. For example, the data description in dictionary form may include more data elements. In some embodiments, the second agent 122 may also be configured to have data descriptions with a lower degree of structuring (e.g., between the description of the overall content and the data description in dictionary form).

[0156] The following Table 1 is an example table. As an example, Table 1 may be converted to markdown format by the second agent 122, but the scope of the present invention is not limited thereto.

[0157]

[0158] For the above Table 1, the second agent 122 may be configured to output the following two-level data description (e.g., caption description):

[0159] ==========;

[0160] [Start of the caption for level 1]:

[0161] This table provides the clay mineralogy results and inferred thermal offsets of different gouge samples and host rock samples, which helps to understand their thermal history and mineral transformation.

[0162] [End of Level 1 caption];

[0163] [Start of Level 2 caption]:

[0164] Sample: Identification number of each sample in the gouge or host rock.

[0165] Type: Indicates whether the sample is gouge or host rock, which helps to classify the source of the sample.

[0166] Illite humidity (relative): Relative humidity of illite, which can indicate the water-bearing conditions.

[0167] Chlorite (peak ratio*): Peak ratio indicated by the 001 reflection and 002 reflection from chlorite measurements, which is necessary for evaluating the mineral structure.

[0168] Inferred thermal offset (illite): Estimated temperature range experienced by the illite sample, which indicates historical thermal events.

[0169] Inferred thermal offset (chlorite): Estimated temperature range experienced by the chlorite sample, which indicates historical thermal events.

[0170] [End of Level 2 caption];

[0171] ==========;

[0172] In some embodiments, in order to enable the multi-modal large model invoked by the second agent 122 to generate hierarchical data descriptions for each table, the multi-modal answer model can be fine-tuned. For example, a large number of documents containing literature can be collected and the tables therein can be divided into different parts. In some embodiments, each table can be divided into a triple containing three parts. For example, a table can be divided into a table body, a table title, and a table footnote. After collecting a large number of table triples, they can be used as a fine-tuning data set for the multi-modal large model to fine-tune the multi-modal large model.

[0173] When fine-tuning the multi-modal large model, the input of the multi-modal large model can be the table body, and the targets output by the multi-modal large model can be the table title and table footnote in the triple. As an example, the multi-modal large model can be optimized by the gradient descent algorithm, whereby the multi-modal large model can learn to output including the table title and table footnote.

[0174] In some embodiments, to accelerate the fine-tuning training of a multi-modal large model, the Low-Rank Adaptation (LoRA) algorithm can be used. LoRA is a technique for fine-tuning large-scale pre-trained models, which aims to reduce the scale of parameter updates by introducing low-rank matrices, thereby improving training efficiency and reducing computational costs. The basic idea of LoRA is to add a set of low-rank adaptive adjustments to the original weights of the model, allowing for effective fine-tuning of the model while only updating a small number of parameters.

[0175] To implement the LoRA algorithm, in some embodiments, parameter decomposition of the model (e.g., a multi-modal large model) can be performed first. For the weight matrix W of certain layers, assuming the size of W is m×n, it can be decomposed into two low-rank matrices A and B, as shown in the following equation (5), where the matrix A can be m×r, the size of matrix B can be r×n, and r is the rank of the low-rank.

[0176] (5);

[0177] Subsequently, during the fine-tuning process, the objective of optimization is to minimize the loss function In some embodiments, the loss function can be the cross-entropy loss, which can be calculated by the following equation (6):

[0178] (6);

[0179] where, are the parameters of the original model, f() is the forward propagation function of the model, x i is the input sample, and y i is the corresponding label.

[0180] During the optimization training process, only the parameters A and B can be updated, while keeping the parameters unchanged. The parameters can be updated by the gradient descent method, as shown in the following equations (7) and (8):

[0181] (7);

[0182] (8);

[0183] The following is another example of data description in dictionary form:

[0184] ==========;

[0185] "Row 1": The content of various pm2.5 pollutants in the Beijing area,

[0186] "The second row": The contents of various PM2.5 pollutants in the Shanghai area,

[0187] "The first column": The content of sulfides,

[0188] "The second column": The content of black carbon;

[0189] ==========;

[0190] For each number in the table in the original document of the literature or the table obtained from the original document, its row entry description (e.g., row attribute) and column entry description (e.g., column attribute) can be obtained according to its row and column (e.g., based on the row index and column index of the table) as shown in the above example. For example, for the number in the first row and second column of the table, through the data description in the form of a dictionary in the above example, its row attribute is "The contents of various PM2.5 pollutants in the Beijing area" and the column attribute is "The content of black carbon". This structured data description in the form of a dictionary can be used by downstream agents (e.g., the fourth agent 124) to check for errors during the downstream data integration phase (e.g., executed by the third agent 123). For example, if the first column of the integrated table is "The content of black carbon", then it is only necessary to verify whether the column attribute of each number in the first column is "The content of black carbon".

[0191] The following will be combined with Figure 2 Describe an example in which the second agent 122 calls a multimodal large model to convert the charts in the document into multiple sets of data in chart form (e.g., multiple tables).

[0192] In some embodiments, for pictures in the document that are difficult to convert into tables, the second agent 122 can be configured to call a multimodal large model to output the description content of the picture point by point. In some embodiments, the multimodal large model can be instructed to output the description content of the picture point by point through prompt design. The following will be combined with Figure 3 Describe an example in which the second agent 122 calls a multimodal large model to output the description of the main content in the picture point by point.

[0193] For the third agent 123, its implementation details will be further described.

[0194] In some embodiments, the user can provide a table template to the second large language model called by the third agent 123 to provide the items that the user cares about. The following Table 2 is an example table template.

[0195]

[0196] Through the above table template, the second large language model can understand that the items the user is concerned about include "country", "province", "city", "including record number", "sulfate", "nitrate", "ammonium salt", "organic carbon", and "black carbon". By providing the existing table as a template, there is no need for the user to additionally input the items the user is concerned about.

[0197] In some embodiments, a table template in markdown format may be provided, but the scope of the present invention is not limited thereto.

[0198] In some embodiments, the user can also provide the items the user is concerned about in other ways. For example, the items the user is interested in can be provided to the second large language model by means of text input.

[0199] In some embodiments, after the second large language model receives the second user input including the items the user is concerned about, the third intelligent agent 123 can be configured to determine whether each set of data (in tabular form) in the multiple sets of data (in tabular form converted by the second intelligent agent 122) contains the target data associated with the items the user is concerned about. The third intelligent agent 123 can be configured to integrate the sets of data containing the target data in the multiple sets of data into a single set of data (in tabular form) based on the determination result.

[0200] In some embodiments, the second large language model is made to determine whether each set of data includes the target data associated with the items the user is concerned about through a prompt. The following Prompt Example 3 shows an example of such a prompt.

[0201] ==========;

[0202] Prompt Example 3:

[0203] You are <input1>Experts in the field, proficient in analyzing whether the tables in papers contain data of interest.

[0204] Please determine whether multiple tables in the paper contain data on the topic of interest.

[0205] When replying, please strictly follow the following rules:

[0206] 1. The user will provide multiple tables (first level), and each table can contain sub-tables (second level).

[0207] You only need to judge whether the first-level tables are relevant to the topic.

[0208] In other words, the number of answers you give should be equal to the number of first-level tables.

[0209] 2. For each first-level table, use 1 or 0 to indicate whether it contains data on the topic of interest, where 1 means yes and 0 means no.

[0210] Please do not give any other replies.

[0211] 3. Please express the final answers in a list form and do not use other text for replies.

[0212] Example:

[0213] [1, 0, 1, 0, 1,...];

[0214] ==========;

[0215] By inputting the above-mentioned prompt example 3, the second largest language model will give a judgment of 0 or 1 for each set of data in table form (which can correspond to each table), where 0 means the table does not include the entries of interest to the user, and 1 means the table includes the entries of interest to the user. The tables marked as 1 can be input into the second largest language model for data integration.

[0216] In some embodiments, the third intelligent agent 123 can be configured to output the source of each data in a single set of data, where the source indicates which set of data among multiple sets of data and which document among multiple documents. By outputting the source of each data in the single set of data while outputting the integrated single set of data, it can prevent the second largest language model from having "hallucinations" and outputting non-existent data, and the source is input to the downstream intelligent agent for its inspection. At the same time, in order to verify the accuracy of data extraction and integration, a part of the tables will be randomly inspected manually. During the inspection process, the source can be used to help the inspectors quickly verify the authenticity of the data.

[0217] Table 3 below shows a single set of data in tabular form output by the third agent 123 with the data source. The entries in Table 3 are associated with the entries in Table 2 above.

[0218]

[0219] As described above, in some embodiments, the second agent 122 may generate multiple sets of data based on a hierarchical data description. In some embodiments, the hierarchical data description may include a data description in dictionary form, which may include descriptions of row entries and / or column entries associated with the tabular form. For these embodiments, in some embodiments, semantic matching may be used, and the second large language model invoked by the third agent 123 may be fine-tuned such that the second large language model can better utilize the hierarchical data description provided by the upstream second agent 122 for data integration.

[0220] In some embodiments, the third agent 123 may be configured to analyze the semantic similarity of the descriptions of row entries and / or column entries to perform a consistency check on multiple sets of data. In some embodiments, the descriptions of row entries and / or column entries may include the full names and abbreviations of variables. Multiple sets of data are integrated based on the results of this consistency check.

[0221] The semantic matching and instruction fine-tuning process of the second large language model as an example is described below.

[0222] First, the hierarchical data description output by the second agent 122 is converted into a semantic vector representation according to the following equation (9):

[0223] (9);

[0224] where E() represents a pre-trained language model, T represents the hierarchical data description (which is in text form), and v T represents the semantic vector.

[0225] Second, for the keywords "table title" and "table footnote", the same processing is performed according to the following equations (10) and (11) to obtain the corresponding semantic vector representations:

[0226] (10);

[0227] (11);

[0228] Subsequently, the cosine similarity is used to calculate the similarity between the hierarchical data description and the keywords "table title" and "table footnote" 、 , as shown in the following equations (12) and (13):

[0229] (12);

[0230] (13);

[0231] Subsequently, select the hierarchical data description with the highest similarity as the part corresponding to the keywords "table title" and "table footnote", as shown in the following formulas (14) and (15):

[0232] (14);

[0233] (15);

[0234] Finally, through the instruction fine-tuning technology, make the second large language model focus on the hierarchical data description, so as to avoid disagreements during the data integration stage.

[0235] The following Prompt Example 4 is an example of a prompt that can achieve the above process.

[0236] ==========;

[0237] Prompt Example 4:

[0238] Please extract the data related to the theme from multiple tables in the paper and organize them into a new table.

[0239] When replying, please strictly follow the following rules:

[0240] 1. Ensure that the converted data matches the original data in terms of numerical value and unit, and clearly indicate the unit.

[0241] 2. Please only provide an integrated table in markdown format, discarding unnecessary and unmergeable data.

[0242] 3. The format of the integrated output table should follow the template provided by the user. But do not output the data in the template.

[0243] 4. If only one piece of data can be integrated, output a single-row table.

[0244] If multiple pieces of data can be integrated, output a multi-row table.

[0245] If one piece of data is missing in the integrated data, use "NaN" to replace it.

[0246] 5. After integrating the table, explain which table each data comes from, with the coordinates being the i-th row and the j-th column.

[0247] 6. During the process of integrating tables, pay attention to the first-level description and second-level description of each table. The first-level description reflects the main content of the table. The second-level description reflects the specific meaning of each row and column in the table. When integrating data from different tables, ensure that the meanings of rows and columns are consistent and the units are the same.

[0248] ==========;

[0249] By using semantic matching instead of string matching in the data merging stage with the second large language model, data of entries with the same semantics but not exactly the same string expressions can be effectively integrated, thereby increasing the amount of effective and accurate integrated data. For example, "Beijing area" and "Beijing" have the same semantics, but the strings are not the same. The data corresponding to these two entries can be effectively integrated through semantic matching.

[0250] For the fourth intelligent agent 124, its implementation details are further described.

[0251] In some embodiments, the fourth intelligent agent 124 may be configured to determine an overall score based on at least one of user interest relevance and accuracy. The fourth intelligent agent 124 may be configured to compare the overall score with a preset threshold. The fourth intelligent agent 124 may be configured to iteratively provide an indication to update a single set of data based on the result of the comparison until the overall score reaches the threshold.

[0252] In some embodiments, the fourth intelligent agent 124 may be configured not to provide an indication to update a single set of data in response to the result of the comparison indicating that the overall score is lower than the threshold. As an example, the fourth intelligent agent 124 may provide an indication to the third intelligent agent 123 in response to the result of the comparison indicating that the overall score is lower than the threshold. Based on this indication, the third intelligent agent 123 may reject some data during the data integration stage. In some embodiments, the indication may include details of one or more data that need to be updated, so that the third intelligent agent 123 can locate these data faster to re-integrate the data. As an example, the fourth intelligent agent 124 may check that one or more data in the single set of data in table form output by the third intelligent agent 123 do not come from the multiple sets of data in table form output by the second intelligent agent 122, and thus determine that these one or more data are inaccurate. Accordingly, the fourth intelligent agent 124 provides an indication to the third intelligent agent 123. This indication may cause the third intelligent agent 123 to continue to reject these one or more data during the data integration stage and re-integrate the data. Thereafter, the fourth intelligent agent 124 may check the single set of data generated after the third intelligent agent 123 re-integrates. This process may be iterated until the overall score of the check of the single set of data integrated by the third intelligent agent 123 reaches the threshold.

[0253] In some embodiments, the fourth agent 124 may be configured to not provide an indication to update a single set of data in response to the result of the comparison indicating that the overall score is greater than or equal to a threshold. This indicates that the current single set of data has passed the inspection and can be used for subsequent processes to generate a literature meta-analysis report based on it.

[0254] The following shows the process in which the fourth agent 124 iteratively improves the single set of data output by the third agent 123 through inspection, and the fourth agent 124 scores the inspection results and provides outputs of suggestions and decisions.

[0255] ==========;

[0256] {"Subject Relevance": 6, "Accuracy": 4, "Overall Score": 5, "Suggestion": "Incorporate geographical data, avoid data mismatches, enhance explanatory details, and ensure clear data source references.", "Decision": "Reject"};

[0257] ==========;

[0258] {"Subject Relevance": 7, "Accuracy": 4, "Overall Score": 5, "Suggestion": "Ensure that all necessary details are included, such as the exact year and geographical coordinates. Improve the clarity and justifiability of data usage, and ensure high accuracy when interpreting and integrating data from multiple sources", "Decision": "Reject"};

[0259] ==========;

[0260] {"Subject Relevance": 7, "Accuracy": 8, "Overall Score": 7, "Suggestion": "Even if speculative, students should explore potential ways to infer or approximate missing geographical and temporal data to enhance the integrity of the analysis. In addition, discussion sessions on how this limitation affects the applicability of the data provide deeper insights into data utilization and reliability.", "Decision": "Accept"};

[0261] ==========;

[0262] The following Table 4 shows an example table after being inspected by the fourth agent 124.

[0263]

[0264] As described above, in some embodiments, the third agent 123 is configured to output the source of each data in a single set of data. For such embodiments, in some embodiments, the fourth agent 124 may be configured to compare each data in the single set of data with the corresponding data in multiple sets of data. The fourth agent 124 may be configured to determine the accuracy based on the result of the comparison.

[0265] The following shows an example process in which the third agent 123 updates a single set of data based on the inspection results included in the indication provided by the fourth agent 124 to improve the overall score.

[0266] The third agent 123 may output an example integration table with data source explanations, as shown in Table 5 below:

[0267]

[0268] The fourth agent 124 checks the output result of Table 5 above as follows:

[0269] ==========;

[0270] ‘Subject Relevance’: 8, ‘Accuracy’: 4, ‘Overall Score’: 5, ‘Suggestion’: ‘The source of number 15 is incorrect. The data recorded in Table 4 is nitrogen dioxide, not carbon dioxide’, ‘Decision’: ‘Reject’;

[0271] ==========;

[0272] The third agent 123 updates Table 5 based on the above output of the fourth agent 124 to output an example updated integration table, as shown in Table 6 below:

[0273]

[0274] The fourth agent 124 checks the output result of Table 6 above as follows:

[0275] ==========;

[0276] ‘Subject Relevance’: 8, ‘Accuracy’: 8, ‘Overall Score’: 8, ‘Suggestion’: ‘No obvious problems’, ‘Decision’: ‘Accept’;

[0277] ==========;

[0278] The above output result indicates that Table 6 has passed the inspection by the fourth agent 124, and the data in this table can be used to generate a report on the literature meta-analysis.

[0279] In some embodiments, the multiple agents may further include an optional fifth agent 125. The fifth agent 125 may be located upstream of the first agent 121.

[0280] In some embodiments, the fifth agent 125 may be configured to receive a plurality of original documents including literature. The plurality of original documents may include documents downloaded from a database or obtained in any other way. The fifth agent 125 may be configured to perform page parsing on each of the plurality of original documents to determine a plurality of parsing elements. The fifth agent 125 may be configured to merge the plurality of parsing elements to form each parsed document, wherein each formed parsed document constitutes the plurality of documents received by the first agent 121. Some embodiments utilize the fifth agent 125 to perform page parsing on the original documents containing literature, and the formed parsed documents can be better analyzed and processed by downstream agents.

[0281] In some embodiments, the plurality of original documents may include charts. The fifth agent 125 may be configured to call a multimodal language model to convert the charts into corresponding parsing elements among the plurality of parsing elements.

[0282] In some embodiments, the charts may include statistical charts and tables, and the multimodal language model may include a first multimodal language model and a second multimodal language model. The fifth agent 125 may be configured to call the first multimodal language model to convert the statistical chart into a corresponding parsing element among the plurality of parsing elements. The fifth agent 125 may be configured to call the second multimodal language model to convert the table into a corresponding parsing element among the plurality of parsing elements. As an example, the first multimodal language model may be a Vision-Language Model (VLM). As an example, the second multimodal language model may be a Vision-Language Model. Some embodiments further divide the charts into statistical charts and tables, and according to the different characteristics of the two, use separate multimodal language models to parse the statistical charts and tables into corresponding parsing elements respectively, improving the accuracy and effect of the parsing of the original documents.

[0283] During the document parsing process, for complex tables, especially those containing multiple sub-tables, OCR technology often has recognition errors, thus affecting the accuracy of the parsed data.

[0284] In some embodiments, through prompt design, the first multimodal language model is made to convert each table (or sub-table) separately when parsing complex tables, thereby improving the parsing accuracy. The following Prompt Example 5 is an example of the prompt input to the first multimodal language model to make it parse the table.

[0285] ==========;

[0286] Prompt Example 5:

[0287] You are <input1>Experts in the field and good at converting tables in papers from picture format to markdown text.

[0288] Please convert the following jpg table into markdown text.

[0289] When replying, please strictly follow the following rules:

[0290] 1. Ensure that the converted data matches the original data in terms of numerical values and units, and clearly indicate the units.

[0291] 2. Preserve the original format of the table.

[0292] 3. If there are multiple tables in the picture, convert each table separately.

[0293] 4. For each table, give its title to reflect the content of the table.

[0294] 5. For each table, provide footnotes to describe the full names of each row and column.

[0295] ==========;

[0296] In some embodiments, for statistical charts, through prompt design, the second multimodal language model can be made to convert multiple statistical charts into text respectively, and for each data among the multiple data in each statistical chart, list them separately in the corresponding text. The following Prompt Example 6 is an example of the prompt for inputting to the second multimodal language model to make it parse the statistical chart.

[0297] ==========;

[0298] Prompt Example 6:

[0299] You are <input1>Experts in the field, proficient in converting statistical graphs in papers from image format to text.

[0300] Please convert the following jpg statistical graph into text.

[0301] When replying, please strictly follow the following rules:

[0302] 1. Ensure that the converted data is consistent with the original data in terms of numerical value and unit, and clearly indicate the unit.

[0303] 2. If there are multiple statistical graphs in the picture, convert them into multiple texts respectively.

[0304] 3. If there are multiple data in a statistical graph, list them separately in the corresponding text.

[0305] 4. Please analyze the data in each statistical graph and draw conclusions. For example: Figure 1 It shows that the carbon dioxide emissions of plants are positively correlated with the soil water content.

[0306] ==========;

[0307] In some embodiments, multiple original documents may include text. The fifth intelligent agent 125 may be configured to use OCR technology to convert the text into corresponding parsing elements in the multiple parsing elements.

[0308] In some embodiments, multiple original documents may include formulas. The fifth intelligent agent 125 may be configured to call the fourth large language model to convert the formulas into corresponding parsing elements in the multiple parsing elements.

[0309] In some embodiments, a large number of formula datasets can be used to fine-tune and train the fourth large language model, so as to greatly improve the ability of the fourth large language model in formula conversion.

[0310] In addition to the accurate parsing of each element of the document, the layout of each parsing element in the parsed document is also very important for subsequent analysis and processing.

[0311] In some embodiments, the fifth intelligent agent 125 is configured to input multiple parsing elements into the page of the document in multiple candidate layouts respectively for each document to be formed. The fifth intelligent agent 125 is configured to record the size and position of each parsing element in the page for each candidate layout for each document to be formed, and calculate the adaptability under the current candidate layout based on the area of the region occupied by the multiple parsing elements in the page and the total area of the page. The fifth intelligent agent 125 is configured to merge the multiple parsing elements with the candidate layout having the maximum fitness for each document to be formed.

[0312] As an example, DocLayout-YOLO can be used to implement the parsing of document layouts. DocLayout-YOLO uses YOLO-v10 as its basic architecture, and its efficient real-time detection ability is suitable for diverse document layout detection.

[0313] For this model, in the pre-training stage, the model regards document synthesis as a two-dimensional bin-packing problem by introducing the Mesh-candidate BestFit method. Through this innovative perspective, the model can better understand and identify the layout characteristics of documents.

[0314] The Mesh-candidate BestFit algorithm aims to achieve and utilize layout effects by efficiently organizing and arranging document elements. It operates according to the following process.

[0315] First, input document elements. Specifically, collect the document elements to be synthesized, such as text boxes, images, tables, etc., and record their sizes and positions.

[0316] Second, create a grid structure. Specifically, divide the document area into two-dimensional grid (Mesh) cells, where each grid cell represents a possible layout position. Set the size of the grid cells. Usually, the layout can be carried out according to the overall size of the document and the size of the elements.

[0317] Subsequently, generate candidate layouts. Specifically, for each document element to be synthesized, generate multiple candidate layouts. These candidate layouts are arrangements of different positions and combinations in the grid cells.

[0318] Subsequently, conduct fitness evaluation. Specifically, calculate the fitness of each subsequent layout to evaluate its space utilization rate and layout rationality. As an example, the fitness can be calculated according to the following formula (16):

[0319] (16);

[0320] Among them, "used area" refers to the area of the region occupied by the document elements, and "total area" refers to the total area of the document region.

[0321] Subsequently, select the best candidate layout. Specifically, select the layout with the highest fitness from the generated candidate layouts as the best layout.

[0322] Subsequently, implement the layout. Specifically, according to the selected best layout, place the document elements in the corresponding grid cells to complete the synthesis of the document.

[0323] Finally, output the synthesized document. Specifically, output the synthesized document to ensure that all elements are arranged according to the best layout.

[0324] In some embodiments, a large-scale diverse synthetic document dataset can be created to train the above-mentioned model for document layout. This synthetic document dataset can cover the structures of various documents, thereby enhancing the generalization ability of the model.

[0325] For multiple pages of a document, when the corresponding agent invokes the model for analysis and processing, both the overall space utilization rate needs to be considered and the page overlap needs to be reduced.

[0326] In some embodiments, in order to balance the space utilization rate and the page overlap problem, when the fifth agent 125 merges multiple pages of a document, the comprehensive page score Score can be calculated according to the following formula (17):

[0327] Score = Used space - Overlapped space (17);

[0328] Wherein, "Used space" represents the intersection of the page layouts, and "Overlapped space" represents the union of the page layouts.

[0329] By maximizing the comprehensive page score when merging pages, the model can make full use of the overall space when merging pages and reduce page number overlap, so that the final page integration effect has better visual perception and layout rationality.

[0330] In some embodiments, the processing pipeline formed by multiple agents can also extend upstream.

[0331] In some embodiments, multiple agents can also include an optional sixth agent 126. The sixth agent 126 can be configured to search for relevant literature based on the research field of interest input by the user to obtain a literature list.

[0332] In some embodiments, multiple agents can also include an optional seventh agent 127. The seventh agent 127 can be configured to receive the literature list output by the sixth agent 126 and automatically download the documents of each literature in the literature list. The multiple documents downloaded by the seventh agent 127 can be used as the input of the fifth agent 125 for the fifth agent 125 to process the documents to parse the page numbers, or can be directly used as the input of the first agent 121 for the first agent 121 to evaluate the literature.

[0333] In some embodiments, the processing pipeline formed by multiple agents can also extend downstream.

[0334] In some embodiments, multiple agents can also include an optional eighth agent 128. The eighth agent 128 can be configured to receive a single set of data in tabular form that has passed the inspection by the fourth agent 124 to perform data analysis on this single set of data. As an example, regression analysis, clustering analysis, classification analysis, etc. can be performed on the single set of data.

[0335] In some embodiments, the eighth agent 128 may be configured to visualize the results of the analysis. In some embodiments, the visualization may include an automatic mode. In the automatic mode, the eighth agent 128 may call the corresponding model to automatically generate the code for visualizing the analysis according to the type, content, etc. of the data. Subsequently, by running the code, the visualization results of the data analysis can be directly obtained. In some embodiments, the visualization may include manual visualization. As an example, for more complex data, such as data that needs to be visualized in combination with a map, etc., manual visualization may be selected.

[0336] In some embodiments, the plurality of agents may further include an optional ninth agent 129. The ninth agent 129 may be configured to receive the results of the data analysis output by the eighth agent 128 to generate a report on the file meta-analysis.

[0337] Figure 2 The first example chart showing the data extraction process using the literature meta-analysis technique according to some embodiments of the present invention.

[0338] Figure 2 Three statistical charts, upper, middle and lower, are shown. Tables 7, 8, and 9 below show the tables generated by the computing device based on Figure 2 the three statistical charts in accordance with some embodiments of the present invention:

[0339]

[0340]

[0341]

[0342] Figure 3 The second example chart showing the data extraction process using the literature meta-analysis technique according to some embodiments of the present invention.

[0343] Figure 3 The second example chart in cannot be converted into a table. In some embodiments, the description content of the picture can be output point by point (e.g., by the second agent 122 in Figure 1 ). In some embodiments, through prompt design, the multi-modal large model called by the second agent 122 in Figure 1 can describe the content of the picture point by point and output it when the picture cannot be converted into a table. The following Prompt Example 6 is an example of such a prompt.

[0344] ==========;

[0345] Prompt Example 6:

[0346] If possible, convert the picture into a Markdown table. If the picture cannot be converted into a Markdown table, describe the picture content point by point.

[0347] Start directly with "1. xxx" without other words.

[0348] Example 2:

[0349] 1. The Nainital area is marked on the map, which may be an important geological sampling point.

[0350] 2. There is a small inset in the upper right corner of the map, showing the larger geographical location relative to Nainital.

[0351] 3. xxx;

[0352] ==========;

[0353] Figure 4 It is a flowchart of a processing pipeline for performing a literature meta-analysis according to some embodiments of the present invention.

[0354] The processing pipeline may include a literature evaluation stage 410, a data extraction stage 420, a data integration stage 430, and a data inspection stage 440. The data integration stage 430 and the data inspection stage 440 may be executed iteratively until the integrated data meets the requirements. The literature evaluation stage 410 may correspond to the operations performed by the first agent 121 in Figure 1 The data extraction stage 420 may correspond to the operations performed by the second agent 122 in Figure 1 The data integration stage 430 may correspond to the operations performed by the third agent 123 in Figure 1 The data inspection stage 440 may correspond to the operations performed by the fourth agent 124 in Figure 1

[0355] In some embodiments, the processing pipeline may further include an optional document processing stage 450. The document processing stage 450 may correspond to the operations performed by the fifth agent 125 in Figure 1

[0356] In some embodiments, the processing pipeline may further include an optional literature search stage 460. The literature search stage 460 may correspond to the operations performed by the sixth agent 126 in Figure 1

[0357] In some embodiments, the processing pipeline may further include an optional document download stage 470. The document download stage 470 may correspond to the operations performed by the seventh agent 127 in Figure 1

[0358] ​​​​In some embodiments, the processing pipeline may further include an optional data analysis stage 480. The data analysis stage 480 may correspond to the operations performed by the eighth agent 128 in Figure 1 .

[0359] In some embodiments, the processing pipeline may further include an optional report generation stage 490. The report generation stage 490 may correspond to the operations performed by the ninth agent 129 in Figure 1 .

[0360] According to another aspect of the present invention, there is provided a method for literature meta-analysis.

[0361] Figure 5 is a flowchart of a method for literature meta-analysis according to some embodiments of the present invention. The method may be executed by a computing device 100 in Figure 1 .

[0362] The method may include step S1: invoking a first large language model by a first agent to perform literature evaluation. Step S1 may be executed by the first agent 121 in Figure 1 .

[0363] The method may include step S2: invoking a multimodal large model by a second agent to perform data extraction. Step S2 may be executed by the second agent 122 in Figure 1 .

[0364] The method may include step S3: invoking a second large language model by a third agent to perform data integration. Step S3 may be executed by the third agent 123 in Figure 1 .

[0365] The method may include step S4: invoking a third large language model by a fourth agent to perform data checking. Step S4 may be executed by the fourth agent 124 in Figure 1 .

[0366] Optionally, the method may further include step S5: performing document processing by a fifth agent. Step S5 may be executed by the fifth agent 125 in Figure 1 .

[0367] Optionally, the method may further include step S6: performing literature search by a sixth agent. Step S6 may be executed by the sixth agent 126 in Figure 1 .

[0368] Optionally, the method may further include step S7: performing document download by a seventh agent. Step S7 may be executed by the seventh agent 127 in Figure 1 .

[0369] Optionally, the method may further include step S8: performing data analysis by an eighth agent. Step S8 may be performed by the eighth agent 128 in Figure 1 .

[0370] Optionally, the method may further include step S9: performing report generation by a ninth agent. Step S9 may be performed by the ninth agent 129 in Figure 1 .

[0371] Figure 6 is a flowchart of a first process associated with a method for literature meta - analysis according to some embodiments of the present invention. The first process may be performed by the first agent 121 in Figure 1 , and may be a specific implementation of step S1 in the method in Figure 5 , but the scope of the present invention is not limited thereto.

[0372] The first process may include step S11: receiving a plurality of documents and a first user input, where the plurality of documents include literature and the first user input includes user requirements.

[0373] The first process may include step S12: scoring the literature corresponding to each document based on the user requirements.

[0374] The first process may include step S13: screening the plurality of documents based on the scoring results to determine a subset of documents.

[0375] Figure 7 is a flowchart of a second process associated with a method for literature meta - analysis according to some embodiments of the present invention. The second process may be performed by the second agent 122 in Figure 1 , and may be a specific implementation of step S2 in the method in Figure 5 , but the scope of the present invention is not limited thereto.

[0376] The second process may include step S21: receiving a subset of documents.

[0377] The second process may include step S22: converting the multimodal data of each document in the subset of documents into multiple sets of data in tabular form.

[0378] Figure 8 is a flowchart of a third process associated with a method for literature meta - analysis according to some embodiments of the present invention. The third process may be performed by the third agent 123 in Figure 1 , and may be a specific implementation of step S3 in the method in Figure 5 , but the scope of the present invention is not limited thereto.

[0379] The third process may include step S31: receiving multiple sets of data and a second user input, where the second user input includes items that the user is concerned about.

[0380] The third process may include step S32: integrating the multiple sets of data into a single set of data in tabular form based on the items that the user is concerned about.

[0381] Figure 9 is a flowchart of a fourth process associated with a method for literature meta-analysis according to some embodiments of the present invention. This fourth process may be executed by Figure 1 the fourth agent 124 in Figure 5 and may be a specific implementation of step S4 in the method in

[0382] The fourth process may include step S41: receiving multiple sets of data and a single set of data.

[0383] The fourth process may include step S42: checking the single set of data against the multiple sets of data based on at least one of user interest relevance and accuracy.

[0384] The fourth process may include step S43: iteratively providing an indication to update the single set of data based on the result of the check until the updated single set of data passes the check, where a report on the literature meta-analysis is generated based on the single set of data that has passed the check.

[0385] Figure 10 is a flowchart of a fifth process associated with a method for literature meta-analysis according to some embodiments of the present invention. This fifth process may be a specific implementation of step S12 in Figure 6 the first process in

[0386] The fifth process may include step S121: determining an independent score for each of a first number of documents based on user requirements and determining a relative score for each document as a whole.

[0387] The fifth process may include step S122: determining a comprehensive score for each document based on the independent score of each of the first number of documents and the pair of scores.

[0388] The fifth process may include step S123: reading a portion at a predetermined position of each document to score each document, where the portion at the predetermined position includes an abstract of the literature and includes a part of the body of the literature, and the amount of text in the portion at the predetermined position conforms to the context window length limit of the first large language model.

[0389] In some embodiments, step S123 can be executed independently of steps S121 and S122. In some embodiments, step S123 can be executed in combination with steps S121 and S122.

[0390] Figure 11 is a flowchart of a sixth process associated with a method for literature meta-analysis according to some embodiments of the present invention. The sixth process can be Figure 7 a specific implementation of step S22 in the second process in , but the scope of the present invention is not limited thereto.

[0391] The sixth process may include step S221: generating a hierarchical data description based on multimodal data in each document, wherein the data descriptions at different levels describe the multimodal data with different degrees of structuring.

[0392] The sixth process may include step S222: generating multiple sets of data based on the hierarchical data description.

[0393] The sixth process may include step S223: outputting the description content of the picture point by point.

[0394] In some embodiments, step S223 can be executed independently of steps S221 and S222. In some embodiments, step S223 can be executed in combination with steps S221 and S222.

[0395] Figure 12 is a flowchart of a seventh process associated with a method for literature meta-analysis according to some embodiments of the present invention. The seventh process can be Figure 8 a specific implementation of step S32 in the third process in , but the scope of the present invention is not limited thereto.

[0396] The seventh process may include step S321: determining whether each set of data in multiple sets of data includes target data associated with an item of interest to the user.

[0397] The seventh process may include step S322: integrating the sets of data that include the target data in the multiple sets of data into a single set of data based on the judgment result.

[0398] The seventh process may include step S323: outputting the source of each data in the single set of data, where the source indicates which set of data in the multiple sets of data and which document among multiple documents the data comes from.

[0399] In some embodiments, step S323 can be executed independently of steps S321 and S322. In some embodiments, step S323 can be executed in combination with steps S321 and S322.

[0400] Figure 13 It is a flowchart of an eighth process associated with a method for literature meta-analysis according to some embodiments of the present invention. The eighth process may be Figure 9 a specific implementation of step S42 in the fourth process in

[0401] but the scope of the present invention is not limited thereto. The eighth process may include step S421: determining an overall score based on at least one of user interest relevance and accuracy.

[0402] The eighth process may include step S422: comparing the overall score with a preset threshold.

[0403] The eighth process may include step S423: based on the result of the comparison, iteratively providing an indication until the overall score reaches the threshold. The indication may be Figure 9 an indication to update a single set of data in step S43 in

[0404] Figure 14 It is a flowchart of a ninth process associated with a method for literature meta-analysis according to some embodiments of the present invention. The ninth process may be Figure 13 a specific implementation of step S423 in the eighth process in

[0405] The ninth process may include step S4321: determining whether the overall score is greater than or equal to the threshold.

[0406] The ninth process may include step S4322: in response to the result of the above comparison indicating that the overall score is greater than or equal to the threshold, not providing an indication. The indication may be Figure 9 an indication to update a single set of data in step S43 in

[0407] The ninth process may include step S4323: in response to the result of the above comparison indicating that the overall score is less than the threshold, providing an indication. The indication may be Figure 9 an indication to update a single set of data in step S43 in

[0408] Figure 15 It is a flowchart of a tenth process associated with a method for literature meta-analysis according to some embodiments of the present invention. The tenth process may be executed by Figure 1 the fifth agent 125 in Figure 5 and may be a specific implementation of step S5 in the method in

[0409] but the scope of the present invention is not limited thereto. The tenth process may include step S51: receiving a plurality of original documents including literature.

[0410] The tenth process may include step S52: performing page parsing on each of the multiple original documents to determine multiple parsing elements.

[0411] The tenth process may include step S53: merging the multiple parsing elements to form each parsed document, where each of the formed parsed documents constitutes multiple documents.

[0412] Figure 16 is a flowchart of an eleventh process associated with a method for literature meta - analysis according to some embodiments of the present invention. The eleventh process may be Figure 15 a specific implementation of step S52 in the tenth process in

[0413] The eleventh process may include step S521: invoking a multimodal language model to convert a chart into a corresponding parsing element among the multiple parsing elements.

[0414] In some embodiments, the chart may include a statistical chart. Step S521 may include step S5211: invoking a first multimodal language model to convert the statistical chart into a corresponding parsing element among the multiple parsing elements.

[0415] In some embodiments, the chart may include a table. Step S521 may include step S5212: invoking a second multimodal language model to convert the table into a corresponding parsing element among the multiple parsing elements.

[0416] The eleventh process may include step S522: using optical character recognition to convert text into a corresponding parsing element among the multiple parsing elements.

[0417] The eleventh process may include step S523: using a fourth large - language model to convert a formula into a corresponding parsing element among the multiple parsing elements.

[0418] Step S521, step S522, and step S523 may be executed independently of each other or in combination with each other.

[0419] Figure 17 is a flowchart of a twelfth process associated with a method for literature meta - analysis according to some embodiments of the present invention. The twelfth process may be Figure 15 a specific implementation of step S53 in the tenth process in

[0420] The twelfth process may include step S531: inputting the multiple parsing elements into the pages of the document in multiple candidate layouts respectively.

[0421] The twelfth process may include step S532: recording the size and position in the page of each parsing element for each candidate layout, and calculating the fitness under the current candidate layout based on the area of the region occupied by multiple parsing elements in the page and the total area of the page.

[0422] The twelfth process may include step S533: merging the multiple parsing elements with the candidate layout having the maximum fitness.

[0423] According to another aspect of the present invention, there is provided a computer-readable storage medium.

[0424] Figure 18 is a block diagram of a computer-readable storage medium 1800 according to some embodiments of the present invention.

[0425] A computer program 1850 is stored on the computer-readable storage medium 1800. When the computer program 1850 is executed by a processor, it implements the steps of the various methods or processes described above in conjunction with Figures 5 - 17 description.

[0426] According to another aspect of the present invention, there is provided a computer program product.

[0427] Figure 19 is a block diagram of a computer program product 1900 according to some embodiments of the present invention.

[0428] The computer program product 1900 may include the computer program 1850. When the computer program 1850 is executed by a processor, it implements the steps of the various methods or processes described above in conjunction with Figures 5 - 17 description.

Claims

1. A computing device for literature meta-analysis, characterized in that: include: Computing resources; as well as A plurality of agents, the plurality of agents being executed by the computing resources, the plurality of agents comprising: The first agent is configured to call the first language model to: receiving a plurality of documents and a first user input, the plurality of documents comprising documents, the first user input comprising a user requirement; Scoring the literature corresponding to each document based on the user's requirements; Filtering the plurality of documents based on the results of the scoring to determine a subset of documents; The second agent is configured to call the multimodal large model to: receiving a subset of the documents; converting the multimodal data in each document in the subset of documents into a plurality of sets of data in a tabular form, each set of data including an entry and data for the entry; Wherein, the second agent is configured to call the multimodal large model to: Based on the first part of the multimodal data in each document, generate a description of the overall content; generating a data description in a dictionary form based on the first part, wherein the data description in the dictionary form includes descriptions of row items and / or descriptions of column items associated with a table form, wherein the plurality of sets of data are generated based on the description of the overall content and the data description in the dictionary form; The third agent is configured to call the second largest language model to: receiving the plurality of sets of data and a second user input, wherein the second user input includes items of interest to the user; Integrating the multiple sets of data into a single set of data in a table format based on the items that the user is concerned about; and The fourth agent is configured to call the third language model to: receiving the multiple sets of data and the single set of data; checking the single set of data against the multiple sets of data based on at least one of relevance to user interests and accuracy; and Based on the result of the check, an instruction to update the single set of data is iteratively provided until the updated single set of data passes the check, wherein a report of the literature meta-analysis is generated based on the single set of data that passes the check.

2. The computing device of claim 1, wherein: the plurality of documents comprises a first number of documents, The first agent is configured to call the first large language model to: Based on the user requirement, determining an independent score for each document in the first number of documents, and determining a relative score for each document in the first number of documents as a whole; determining a composite score for each document based on the individual scores and the relative scores of each document in the first number of documents; as well as The plurality of documents are filtered based on a composite score of each document in the first number of documents.

3. The computing device of claim 1, wherein: The user requirements include at least one of subject relevance, innovation, feasibility, and / or It is characterized in that the first user inputs the reason for requiring the first language model to output a score.

4. The computing device of claim 1, wherein: The first agent is configured to call the first large language model to: A portion at a predetermined position of each document is read to score each document, wherein the portion at the predetermined position includes an abstract of the document and a portion of the main text of the document, and the amount of text in the portion at the predetermined position complies with a context window length limit of the first largest language model.

5. The computing device of claim 1, wherein: The third agent is configured to call the second language model to: Analyze the semantic similarity of the description of the row item and / or the description of the column item to perform a consistency check on the multiple sets of data, wherein the description of the row item and / or the description of the column item includes the full name and the abbreviation of the variable, and the multiple sets of data are integrated based on the result of the consistency check.

6. The computing device of claim 1, wherein: The multimodal data includes pictures, The second agent is configured to call the multimodal large model to: Output the description content of the picture point by point.

7. The computing device of claim 1, wherein: The third agent is configured to call the second language model to: Determining whether each of the plurality of data sets includes target data associated with the item of interest to the user; and Based on the result of the determination, each set of data including the target data in the multiple sets of data is integrated into the single set of data.

8. The computing device of claim 1, wherein: The third agent is configured to call the second language model to: The source of each data in the single set of data is output, the source indicating which set of data in the multiple sets of data and which document in the multiple documents the data comes from.

9. The computing device of claim 8, wherein: The fourth agent is configured to call the third language model to: Comparing each data in the single set of data with corresponding data in the multiple sets of data; as well as Based on the results of the comparison, the accuracy of the data is determined.

10. The computing device of claim 1, wherein: The fourth agent is configured to call the third language model to: determining an overall score based on at least one of the user interest relevance and accuracy; comparing the overall score with a preset threshold; as well as Based on a result of the comparing, the indication is iteratively provided until the overall score reaches the threshold.

11. The computing device of claim 10, wherein: The fourth agent is configured to call the third language model to: In response to a result of the comparing indicating that the overall score is below the threshold, providing the indication; In response to a result of the comparing indicating that the overall score is greater than or equal to the threshold, not providing the indication.

12. The computing device of claim 1, wherein: The plurality of agents also include: The fifth agent is configured as: receiving a plurality of original documents including documents; performing page parsing on each of the plurality of original documents to determine a plurality of parsed elements; The plurality of parsed elements are merged to form each parsed document, wherein each parsed document formed constitutes the plurality of documents.

13. The computing device of claim 12, wherein: The plurality of original documents include diagrams, The fifth agent is configured to: The multimodal language model is invoked to convert the graph into corresponding parsed elements of the plurality of parsed elements.

14. The computing device of claim 13, wherein: The chart includes a statistical chart and a table, the multimodal language model includes a first multimodal language model and a second multimodal language model, The fifth agent is configured to: calling the first multimodal language model to convert the statistical graph into corresponding parsed elements among the plurality of parsed elements; and The second multimodal language model is called to convert the table into corresponding parsed elements among the plurality of parsed elements.

15. The computing device of claim 13, wherein: The multiple original documents include text and formulas, The fifth agent is configured to: converting the text into corresponding parsed elements of the plurality of parsed elements using optical character recognition technology; A fourth language model is called to convert the formula into corresponding parsed elements among the multiple parsed elements.

16. The computing device of claim 12, wherein: The fifth agent is configured to, for each document to be formed: Inputting the plurality of parsed elements into pages of the document in a plurality of candidate layouts respectively; Recording the size and position of each parsed element in the page for each candidate layout, and calculating the fitness under the current candidate layout based on the area of ​​the region occupied by the multiple parsed elements in the page and the total area of ​​the page; as well as The multiple parsed elements are merged using the candidate layout with the maximum fitness.

17. A method for literature meta-analysis, characterized in that: The following steps are involved: S1: The first agent calls the first language model to perform literature review, and the step S1 includes: S11: receiving a plurality of documents and a first user input, wherein the plurality of documents include documents, and the first user input includes a user requirement; S12: Score the literature corresponding to each document based on the user's requirements; S13: Filtering the plurality of documents based on the scoring result to determine a subset of documents; S2: The second agent calls the multimodal large model to perform data extraction, and the step S2 includes: S21: receiving a subset of the document; S22: converting the multimodal data in each document in the subset of documents into multiple sets of data in a table form, each set of data including an entry and data for the entry; The step S22 comprises the following steps: Based on the first part of the multimodal data in each document, generate a description of the overall content; generating a data description in a dictionary form based on the first part, wherein the data description in the dictionary form includes descriptions of row items and / or descriptions of column items associated with a table form, wherein the plurality of sets of data are generated based on the description of the overall content and the data description in the dictionary form, S3: The third agent calls the second largest language model to perform data integration, and the step S3 includes: S31: receiving the multiple groups of data and a second user input, where the second user input includes items that the user is concerned about; S32: integrating the multiple sets of data into a single set of data in a table format based on the items that the user is concerned about; S4: The fourth agent calls the third language model to perform data checking, and the step S4 includes: S41: receiving the multiple groups of data and the single group of data; S42: Checking the single set of data against the multiple sets of data based on at least one of user interest relevance and accuracy; S43: Based on the result of the check, iteratively provide an instruction to update the single set of data until the updated single set of data passes the check, wherein a report of the literature meta-analysis is generated based on the single set of data that passes the check.

18. The method according to claim 17, characterized in that The step S32 comprises the following steps: S321: Determine whether each of the plurality of data groups includes target data associated with the item that the user is concerned about; and S322: Integrate the groups of data including the target data in the multiple groups of data into the single group of data.

19. The method according to claim 17, characterized in that The step S42 comprises the following steps: S421: Determine an overall score based on at least one of the user interest relevance and accuracy; S422: Compare the overall score with a preset threshold; and S423: Based on the result of the comparison, iteratively provide the indication until the overall score reaches the threshold.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 17 to 19 are implemented.

21. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 17 to 19 are implemented.

Citation Information

Patent Citations

  • Construction method and device of image-text search database, database and storage medium

    CN119293270A

  • System and method for automatically generating review based on artificial intelligence

    CN119322842A

  • Paper generation method and device of big language model based on RAG

    CN119357413A