Multi-agent real-world clinical curative effect evaluation and accurate decision-making system

By integrating causal inference technology with a large language model through a multi-agent system, the entire process of clinical efficacy evaluation is automated, solving the problems of low efficiency, high professional knowledge requirements and data security of traditional methods, and achieving rapid and reliable support for clinical research.

CN121306388APending Publication Date: 2026-01-09WEST CHINA HOSPITAL SICHUAN UNIV

Patent Information

Application Number
CN202511882362.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Traditional clinical efficacy evaluation methods are inefficient, require a high level of professional knowledge, lack intelligence and data security, and cannot quickly respond to the timeliness requirements of clinical research. Furthermore, existing AI-assisted tools have failed to achieve full-chain causal inference.

Method used

Construct a multi-agent system, including a research needs analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and an outcome report generation agent. Integrate causal inference technology with a large language model to automate the entire evaluation process from research needs input to clinical evaluation report.

Benefits of technology

It significantly lowers the technical threshold, allowing non-professionals to participate in research, significantly improves research efficiency, breaks through the bottleneck of causal inference, ensures data security, supports accurate clinical decision-making, shortens the research cycle, and improves the reliability of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306388A_ABST
    Figure CN121306388A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent real-world clinical curative effect evaluation and accurate decision-making system, and relates to the technical field of medical data processing. According to the invention, a research demand analysis agent generates a research scheme after understanding the input content of a user and transmits the research scheme to a data management agent, a statistical analysis modeling agent, a result report generation agent, a data security agent and the data management agent privacy and encrypt data uploaded by the user; meanwhile, additional feature construction is carried out according to the research requirement of a user to form brand new analysis data used by a subsequent statistical analysis modeling agent, and after the statistical analysis modeling agent obtains the data, an analysis result is obtained by calling external software and is transmitted to a result report generation agent; and generating a clinical evaluation report together with the research scheme transmitted by the research demand analysis agent. According to the invention, full-process automatic, specialized and safe research support is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing technology, particularly to the field of clinical efficacy evaluation technology, and more specifically to a multi-agent real-world clinical efficacy evaluation and precision decision-making system. Background Technology

[0002] Clinical efficacy evaluation is a crucial part of medical research, especially real-world data-based evaluation. Real-world studies involve observing and analyzing patients' treatment processes in natural environments to assess the actual clinical effects of drugs, treatments, etc. These studies reflect real-world medical conditions and provide important evidence for clinical decision-making. However, traditional observational clinical efficacy evaluations often face numerous challenges, such as the large and complex volume of data, the significant time and manpower required for data processing and analysis, and the high level of specialized knowledge required.

[0003] Traditional methods face three major technical bottlenecks: (1) Data integration challenges: Real-world medical data exhibits multimodal characteristics (electronic medical record text, image data, laboratory indicators, etc.), and existing ETL tools have a semantic parsing accuracy of less than 60% for unstructured text; (2) Low analytical efficiency: Traditional statistical methods require manual design of covariate balancing strategies, and a single study takes an average of 6-8 months; (3) Limitations of causal inference: PSM (propensity score matching) methods suffer from the curse of dimensionality when dealing with high-dimensional confounding variables. When the number of variables exceeds 50, the matching failure rate exceeds 40%. Moreover, complex causal inference techniques, such as target trial simulation frameworks involving dozens of indicator definitions, require researchers to have deep expertise in clinical, statistical, epidemiological and even artificial intelligence fields.

[0004] In existing clinical efficacy evaluation processes, collaboration among multidisciplinary professionals (such as clinicians, statisticians, and epidemiologists) is typically required. First, researchers need to develop a detailed research protocol based on the research objectives, clearly defining inclusion and exclusion criteria, observation indicators, and other key elements. Then, the collected data undergoes rigorous quality control to ensure its completeness and accuracy. Next, various statistical analysis methods are used to process and model the data to assess treatment effectiveness, and the results are professionally interpreted, ultimately forming a clinical evaluation report.

[0005] This traditional process has significant limitations. On the one hand, the data processing and statistical analysis are complex, requiring extensive manual coding and prone to errors. On the other hand, for non-professionals, understanding and applying complex statistical analysis methods and accurately interpreting results is quite difficult, which to some extent limits the efficiency and widespread application of clinical efficacy evaluation.

[0006] With the rapid development of artificial intelligence technology, large language models (such as Deepseek, Tongyi Qianwen, and Wenyan Yixin) have demonstrated powerful natural language processing capabilities in numerous fields. These models, trained on massive amounts of text data, can understand and generate human language, providing intelligent solutions for various complex tasks, such as text generation, question-answering systems, and data analysis. In the healthcare field, utilizing large language models for data processing, analysis, and knowledge extraction helps improve the efficiency and quality of medical services and optimize clinical research processes.

[0007] Currently, existing technical solutions mainly rely on manual or semi-automated observational clinical efficacy evaluation methods based on traditional statistical software and partially automated tools, which are described in detail below: (a) Methods based on traditional statistical software Manual data processing and analysis: Researchers typically use tools such as Excel to perform preliminary data processing, including data cleaning and variable coding. Then, the processed data is imported into professional statistical analysis software (such as SPSS, SAS, R, etc.). Within the statistical software, appropriate statistical analysis methods are selected based on the research design. Statistical models are constructed manually or through a graphical interface, such as performing t-tests, chi-square tests, and regression analyses to evaluate clinical efficacy.

[0008] Results Interpretation and Report Writing: After the analysis, researchers need to interpret the statistical results and, in conjunction with their clinical expertise, determine whether the treatment effect is statistically and clinically significant. Finally, they manually write the clinical evaluation report, organizing the research objectives, methods, results, and conclusions into a document.

[0009] (ii) Methods assisted by partially automated tools Data management tools: Specialized data management software can help researchers collect, store, and manage clinical research data more efficiently. These tools provide data entry templates, data validation functions, and simple data query and export capabilities, reducing human error in data management.

[0010] Automated analysis plugins or scripts: Some statistical analysis software provides pre-built analysis plugins or scripts, allowing researchers to select appropriate plugins for a degree of automation based on the type of research. For example, in the R language, there are ready-made packages (such as the "survival" package for survival analysis) that can simplify the analysis process, but researchers still need some programming knowledge to adjust parameters and interpret results.

[0011] (III) Limitations of existing technical solutions Inefficiency: Whether manual or semi-automatic, the process involves a large number of manual operations, from data processing and analysis to report writing. The entire process is time-consuming and labor-intensive, making it difficult to respond quickly to the timeliness requirements of clinical research.

[0012] High requirements for professional knowledge: In real-world studies, researchers not only need solid clinical medical knowledge and clinical experience, but also need to be proficient in statistical analysis methods and the use of related software. This limits the participation of non-professionals (such as some clinicians or research assistants) in the evaluation of clinical efficacy.

[0013] Lack of intelligence: Existing methods cannot automatically understand and interpret users' research needs, nor can they intelligently select appropriate data processing and analysis strategies according to different research scenarios, lacking flexibility and adaptability.

[0014] Data security risks: During data flow and processing, especially when sensitive patient information is involved, existing methods lack a sound data security protection mechanism, which can easily lead to security problems such as data leakage.

[0015] Existing solutions (such as single AI-assisted statistical tools) only solve local problems (such as data cleaning or report generation) and do not achieve full-chain causal inference. Summary of the Invention

[0016] To overcome the defects and shortcomings of the existing technology, the present invention provides a multi-agent real-world clinical efficacy evaluation and precision decision-making system. The purpose of the invention is to lower the technical threshold for clinicians to use real-world data to conduct observational clinical efficacy evaluation and to rapidly carry out such research.

[0017] This invention includes a research needs analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a results report generation agent. The research needs analysis agent, after understanding the user's input, generates a research plan and transmits it to the data management agent, the statistical analysis and modeling agent, and the results report generation agent. The data security agent and the data management agent perform privacy and encryption processing on the data uploaded by the user, and simultaneously construct additional features according to the user's research needs to form new analysis data for subsequent use by the statistical analysis and modeling agent. After obtaining the data, the statistical analysis and modeling agent obtains the analysis results by calling external software and transmits them to the results report generation agent. Together with the research plan transmitted by the research needs analysis agent, they generate a clinical evaluation report.

[0018] This invention integrates causal inference technology with the powerful language understanding capabilities of large language models to design an intelligent agent for conducting clinical evaluations based on real-world data using complex causal inference techniques. This reduces the research threshold and time, helps clinicians design research for complex causal inference, and enables the medical community to quickly assess the effectiveness and safety of related drugs.

[0019] To address the problems existing in the prior art, the present invention is implemented through the following technical solution.

[0020] This invention provides a multi-agent real-world clinical efficacy evaluation and precision decision-making system, which includes a research needs analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a results report generation agent; each agent can be invoked independently or through collaborative operation to complete the entire process of clinical efficacy evaluation from research needs input to clinical evaluation report output; The research requirement parsing agent receives the research objective or requirement text and related data input by the user. Based on an enhanced semantic parsing large language model in the field of clinical epidemiology, it extracts and transforms the key elements of observational studies in the text to generate a research protocol that conforms to clinical epidemiology standards. The key elements of observational studies include patient inclusion and exclusion criteria, exposure and control group definitions, outcome definitions, confounding variables, bias control, and statistical analysis methods. The enhanced semantic parsing large language model in the field of clinical epidemiology is formed by constructing a corpus based on an observational study knowledge literature base, and then fine-tuning the base large language model multiple times based on this corpus. The data security intelligent agent is used to review sensitive personal information in the dataset uploaded by users. It reviews the data through a sensitive information detection mechanism, and if sensitive information is found, it performs desensitization processing on the data. The data management agent receives user-uploaded data to be processed, performs preliminary understanding and evaluation of the data based on the research plan and user needs, determines rules for data governance and / or feature construction, generates corresponding data processing code from the data processing vertical big model in the backend of the data management agent, calls external data processing tools to run the generated data processing code, performs data governance processing on the data to be processed, and outputs a standard dataset for statistical analysis; the data to be processed includes data processed by the data security agent or compliant data directly uploaded by users; the data processing vertical big model is obtained by optimizing the base big language model through prompting engineering; The statistical analysis modeling agent receives the standard dataset output by the data management agent and the research plan generated by the research requirements parsing agent. It then calls a large statistical analysis language model specifically designed for generating statistical software code to generate code for handling confounding factors, bias control, and analytical modeling. Finally, it calls external statistical analysis software to execute the code, perform statistical analysis, and output the statistical analysis results. The large statistical analysis language model is obtained by fine-tuning the base language model based on a pre-built corpus of statistical analysis code in the field of observational research. The results report generation agent receives the statistical analysis results output by the statistical analysis modeling agent and the research plan generated by the research needs parsing agent. It then uses an enhanced semantic parsing large language model in the field of clinical epidemiology to jointly interpret the results. In accordance with the standard format and logical requirements of clinical epidemiology reports, it generates a clinical evaluation report that includes the research background, purpose, methods, results, and risk of bias assessment. It also provides efficacy difference evaluation and individualized treatment decision suggestions for specific clinical characteristic groups based on age, disease stage, and / or comorbidities.

[0021] Further preferably, the system also includes a task coordination agent, which is connected to a research requirement analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a result report generation agent. The task coordination agent is used to perform semantic understanding of the task requirements input by the user, parse out the specific steps of the clinical efficacy evaluation task, match the task requirements with the corresponding agents based on the functional mapping rules of each agent and determine the execution order, call each agent according to the execution order, monitor the execution status of each agent in real time, handle abnormalities during the execution process and provide feedback for adjustment, and integrate the output results of each agent to generate a clinical evaluation report.

[0022] More preferably, the work coordination agent dynamically allocates tasks to the most suitable agent through an auction algorithm, satisfying the following objectives and constraints: Objective function: ; Constraints: ,and ; in, The utility value of agent j for performing task i is calculated based on the agent's historical task performance and resource utilization. This indicates that task i is assigned to agent j; m represents the total number of agents participating in task assignment, and n represents the total number of independent tasks to be assigned; when the objective function reaches its maximum value at a certain agent, the corresponding task is assigned to that agent.

[0023] Furthermore, the work coordination agent incorporates priority weights. To adapt to the differences in task priorities at different stages of the research, the objective function is dynamically adjusted. ;in The priority weight for task i is determined by the research phase requirements.

[0024] Further preferably, the multiple fine-tuning of the enhanced semantic parsing large language model in the clinical epidemiology field includes the following steps: S101. Corpus Preprocessing: Collect publicly available registration information and research papers from observational studies; clean the text data, correct spelling errors, and / or complete missing fields; establish an annotation framework including population inclusion and exclusion criteria, exposure and control group definitions, outcome definitions, confounding variables, bias control, and statistical analysis methods; use BIO annotation to annotate sentences with entity annotation, refining the annotation granularity to a three-level structure of population characteristics, intervention type, and statistical model; use Clinical BERT technology to tokenize the text, generating input sequences in [CLS]+sentence+[SEP] format; map words to vector space using the WordPiece algorithm; and construct a corpus with a JSON format triple storage structure, which includes research area, research type, and element content. S102, First Fine-tuning: Based on the constructed corpus, the corpus is used as the model output, and the research abstract is used as the model input. The base model parameters are updated using full fine-tuning technology to form a preliminary large language model adapted to the field of clinical epidemiology. S103, Simulation Corpus Generation: Using the large language model initially adapted to the clinical epidemiology field in step S102, an observational study simulation scheme is generated according to the target trial simulation framework. Answers are generated based on zero-shot and few-shot cue word engineering and scored using a scoring formula to screen high-quality simulation corpus. The scoring formula is as follows: ; S104. Second Fine-tuning: Combine the high-quality simulated corpus selected in step S103 with the real corpus in step S101, supplement the zero-time determination rules, simulated randomization grouping, intervention allocation element annotation, and corpus related to differences with observational studies, and perform a second fine-tuning of the large language model initially adapted to the clinical epidemiology field to form an enhanced semantic parsing large language model for the clinical epidemiology field that supports complex causal inferences.

[0025] More preferably, in the data security intelligent agent, the sensitive information detection mechanism refers to reviewing the data based on a vertical large-scale sensitive information detection model fine-tuned according to medical sensitive information-related rules; the vertical large-scale sensitive information detection model is constructed by fine-tuning the base large-scale language model deepseek-r1 based on sensitive information corpus.

[0026] Further optimized The data security intelligent agent identifies sensitive information in user-uploaded datasets using a vertical large-scale sensitive information detection model, and applies pre-approved desensitization rules to desensitize the data containing sensitive information. After desensitization, it generates de-privacy columns to replace the data columns in the original text and returns the data to the user. The construction process of the vertical large-scale sensitive information detection model is as follows: S201. Using existing data sources, construct a prompt word library that includes privacy fields such as patient name, ID number, date of birth, registered address, residential address and / or chief complaint information in hospital admission records; S202. Use the base-based large language model to parse, classify, and / or segment long medical texts; S203. Input the prompt words and the categorized and / or segmented medical text into the base big language model to identify suspected privacy information columns in the medical data; S204. Based on regular expressions, determine whether the privacy information column identified by the base large language model is accurate. If it is not accurate, return an error message and optimize the prompt words until the identification is accurate, thus obtaining the sensitive information detection vertical large model.

[0027] In a further preferred embodiment, the data management agent utilizes prompting engineering to improve the understanding of data processing requirements and code generation capabilities of the base large language model, forming the data processing vertical large model; the data management agent integrates a multi-language parser through the MCP server, isolates the execution environment through Docker, calls external data processing tools, runs the generated data processing code, and generates a standard dataset for statistical analysis.

[0028] More preferably, the specific process by which the data management agent outputs a standard dataset for statistical analysis is as follows: S301. Use a large vertical data processing model to perform a preliminary understanding and evaluation of the data to be processed. The preliminary understanding and evaluation includes determining the data type and the situation of missing data. S302. Utilize the vertical large-scale data processing model to analyze user needs and determine the rules for data governance and / or feature construction, including data transformation, data standardization, missing value imputation, and / or derived feature generation. S303, Generate data processing code that satisfies data governance and / or feature construction rules from a large vertical data processing model; S304 integrates a multilingual parser through the MCP server, isolates the execution environment through Docker, calls external data processing tools, runs the generated data processing code, and generates a standard dataset for statistical analysis.

[0029] More preferably, in the statistical analysis modeling agent, the statistical analysis large language model is obtained by fine-tuning the base model based on a pre-built corpus of statistical analysis codes in the professional field of observational research.

[0030] Further optimized The construction process of the statistical analysis large language model is as follows: S401. Build a code knowledge base by collecting functions and documentation in the target programming language; S402. Construct a statistical analysis code corpus for observational research in a professional field using sample data and sample analysis schemes; S403. Fine-tune the code generation process in complex causal inference methods within the domain to obtain a large language model for statistical analysis.

[0031] Furthermore, the statistical analysis modeling agent, based on user needs, constructs a dedicated task suitable for complex causal inference, namely a dynamic causal inference module based on meta-learning, which adjusts the matching strategy in real time by learning the weight distribution of confounding variables. ; IPM is a measure of the maximum mean difference. For dynamic weighting coefficients, X represents the model parameters; X represents the covariates from different dimensions in the research data; Y represents the outcome variable, i.e., the true value. This represents the average value in the statistical analysis dataset D; This represents the loss function, used to measure the difference between the predicted value of the prediction model and the true value Y; This represents the model's predicted value. This indicates the distribution of covariate X in a population receiving a certain treatment; This represents the distribution of covariate X in the population that did not receive a certain treatment, where T=0 represents the population that did not receive a certain treatment and T=1 represents the population that received a certain treatment. And effect estimation based on the target experiment simulation framework: ; in, This represents the average treatment response and is used to measure the average difference in efficacy between the intervention and the control. This indicates that under covariate X, the intervention... The corresponding latent outcome mean function, ; It represents the distribution measure of the covariate X, reflecting the distribution pattern of patient characteristics in real-world data; Indicates the outcome of receiving a certain treatment; It indicates an outcome of not accepting a certain treatment; For the target experiment simulation framework, reinforcement learning is used to achieve dynamic parameter adjustment and real-time optimization of key parameters in the TTE framework by introducing state. , Where KL represents the Kullback-Leibler divergence of the covariate distributions between the exposed group and the control group; This represents the average treatment effect of the current simulation step. Indicates a measure of bias; The reward function is set to ;in, This represents a known average treatment effect, used to measure... Is it accurate? , and All are weighting coefficients. This represents the simulation convergence speed; finally, the optimal policy is learned through a deep Q-network.

[0032] Furthermore, a preferred approach is to introduce knowledge distillation with causal constraints to reduce the overall deployment scale of the model, while also introducing distillation loss to ensure that complex causal inference results can be efficiently parsed. The objective function is to minimize the KL divergence so that the output distribution of the sub-model approximates the output distribution of the original model as closely as possible. The objective function is ; ; The loss function incorporates causal inference constraints, where IPM is used to measure the difference in covariates between the exposed and control groups. For IPM, the maximum mean difference is used for measurement. ; in, For a family of functions in the reproducing kernel Hilbert space, based on kernel functions The specific calculation formula is as follows:

[0033] Where n represents the sample size of the population receiving intervention; m represents the sample size of the population not receiving intervention; i represents the index variable of the population receiving intervention; and j represents the index variable of the population not receiving intervention. Represents the covariate vector of the i-th treatment group; Represents the covariate vector of patients in the j-th treatment group; Represents the covariate vector of the i-th control group patient; Let represent the covariate vector of the j-th control group patient.

[0034] Compared with the prior art, the beneficial technical effects of the present invention are as follows: This invention constructs a multi-agent system integrating causal inference technology and large language model capabilities. Addressing the technical bottlenecks and shortcomings of traditional observational clinical efficacy evaluation methods, it achieves fully automated, professional, and safe research support. Specific technical effects are as follows: 1. Significantly lowers the technical threshold, enabling non-professionals to participate in research. The research needs analysis agent, based on a domain-enhanced semantic parsing large language model in clinical epidemiology, can automatically transform unstructured needs texts from clinicians (such as "evaluating the preventive effect of Tanreqing on multidrug-resistant bacterial infections in critically ill patients") into research protocols conforming to the STROBE statement. It accurately extracts key elements such as patient inclusion / exclusion criteria, exposure / control group definitions, outcome indicators, and statistical analysis methods, eliminating the need for manual literature review or consultation with statisticians. The data management agent automatically generates data processing code (such as one-hot encoding of text and mean imputation of missing values) through prompting engineering guidance from a vertical large model. The statistical analysis modeling agent outputs code for heterogeneous control and causal inference (such as target trial simulation) based on a finely tuned statistical analysis large language model. Clinicians do not need to master programming skills such as Python or R to complete complex data processing and modeling. The results reporting agent, through a domain-enhanced large language model, jointly interprets the research protocol and statistical results, automatically generating clinical reports including research background, methods, results, and risk of bias assessment. It also provides efficacy difference analyses for subgroups such as age and disease stage, allowing non-professionals to directly obtain clear and professional research conclusions.

[0035] 2. Significantly improves research efficiency and shortens the observational study cycle. Traditional methods require manual integration of multimodal data (electronic medical record text, laboratory indicators, etc.), with semantic parsing accuracy less than 60%. This system automates feature selection, data transformation, and derivative feature construction through a data management agent. Combined with ClinicalBERT technology, the parsing accuracy of unstructured text is improved to over 90%, and the processing time for a single dataset is reduced from weeks to hours. Traditional research relies on manually designed covariate balancing strategies, with an average study time of 6-8 months. This system integrates a dynamic causal inference module (based on meta-learning to adjust matching strategies in real time) with a target trial simulation framework through a statistical analysis modeling agent. It automatically handles high-dimensional confounding variables (the matching failure rate is reduced from 40% to below 5% when the number of variables exceeds 50), shortening the study cycle to days and enabling rapid response to the clinical timeliness requirements of drug safety and efficacy assessments. The work coordination agent dynamically allocates tasks through an auction algorithm, adjusts the execution order based on the priority weight of the research stage, monitors the agent's status in real time and handles anomalies (such as automatically correcting common data format errors), avoiding manual coordination and waiting, and achieving an automation rate of over 95% for the entire process.

[0036] 3. Overcoming the technical bottlenecks of causal inference and improving the reliability of research results. Addressing the high matching failure rate of traditional PSM methods when the number of variables exceeds 50, the statistical analysis modeling agent introduces the maximum mean difference (IPM) to measure the difference in covariate distributions between the exposed and control groups, through an objective function... Real-time optimization of matching strategies ensures covariate balance under high-dimensional data, improving the accuracy of causal inference results by over 30%. Traditional methods require manually defining dozens of indicators (such as randomization windows and grace periods), necessitating cross-disciplinary expert collaboration. This system uses reinforcement learning to dynamically optimize key parameters and automatically complete the simulation process, enabling clinicians to conduct high-quality causal inference research without needing to master complex theories. The research requirement analysis agent undergoes two fine-tuning processes (first optimizing the professional style based on real corpora, then enhancing causal inference capabilities by combining simulated corpora), and employs... By screening high-quality training data, the "illusion" rate of model-generated research plans has been reduced from 25% in traditional large models to below 5%, ensuring that the research design complies with clinical epidemiological standards.

[0037] 4. Construct a full-chain data security mechanism to ensure compliance with medical privacy regulations. The data security intelligent agent is based on a vertical large-scale model (using Deepseek-R1 as a base) finely tuned from a corpus of medical sensitive information. Combined with a prompt word library containing privacy fields such as patient names and ID numbers, it achieves a 98% accuracy rate in detecting sensitive information through a process of "suspected column identification - text segmentation - regular expression verification," avoiding the omissions of traditional rule engines in detecting hidden privacy information (such as indirect identity descriptions in the admission complaint). Differentiated strategies are adopted for different privacy information: names and ID numbers are encrypted using MD5 and processed locally to generate unique identification codes (eliminating encryption traces), addresses are completely replaced with "*", and the last four digits of contact information are retained. While complying with the "Personal Information Protection Law" and the "Medical Data Security Guidelines," it ensures that the de-identified data can still be used for subsequent statistical analysis (such as subgroup population feature matching). The data management agent builds an isolated execution environment using Docker containers, limiting CPU and memory resource usage and granting only data read (specifying the input set) and result write (specifying the output path) permissions. Combined with static code analysis tools, it detects dangerous operations (such as deleting files or remote connections) to avoid privacy leaks during data processing.

[0038] 5. Support precise clinical decision-making and promote the clinical translation of research findings. The results report generation agent automatically summarizes the efficacy differences among different population groups (age <65 years / ≥65 years, presence or absence of comorbidities, etc.) based on statistical analysis results (e.g., "The infection rate of multidrug-resistant bacteria decreased by 11% in critically ill patients under 65 years old after using Tanreqing, and by 6% in patients over 65 years old"), providing data support for individualized treatment. The system conducts analysis based on real-world data (electronic medical records, laboratory indicators, etc.), and the generated clinical evaluation report includes "risk of bias assessment" and "real-world applicability analysis," avoiding the disconnect between traditional randomized controlled trial results and actual clinical scenarios, and helping to quickly translate research conclusions into clinical treatment plans (e.g., precise targeting of drug users). Through causal constraint knowledge distillation technology, the model deployment scale is reduced by 70% while ensuring the accuracy of causal inference, supporting local hospital deployment (data is not transmitted outside the institution), adapting to the computing power limitations and data privacy requirements of clinical scenarios.

[0039] In summary, this invention, through the deep integration of multi-agent collaboration and AI technology, not only solves the problems of efficiency, professional threshold and technical bottlenecks in traditional observational clinical efficacy evaluation, but also provides an efficient and reliable technical solution for the large-scale implementation and clinical translation of real-world research through precise decision support and safe and compliant design. Attached Figure Description

[0040] Figure 1 This is a framework diagram of the multi-agent real-world clinical efficacy evaluation and precision decision-making system of the present invention; Figure 2This is a schematic diagram illustrating the principle of fine-tuning the enhanced semantic parsing large language model in the field of clinical epidemiology. Figure 3 This invention provides a schematic diagram of the intelligent agent principle to illustrate the research requirements of the present invention. Figure 4 This is a schematic diagram of the data management intelligent agent of the present invention; Figure 5 This is a schematic diagram of the intelligent agent principle for statistical analysis and modeling in this invention; Figure 6 Generate an intelligent agent principle diagram for the results report of this invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0042] Example 1 As a preferred embodiment of the present invention, please refer to the appendix to the specification. Figure 1 As shown in the figure, this embodiment discloses a multi-agent real-world clinical efficacy evaluation and precision decision-making system. The system includes a research needs analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a result report generation agent. Each agent can be called independently or through collaborative operation to complete the entire process of clinical efficacy evaluation from research needs input to clinical evaluation report output. The research requirement parsing agent receives the research objective or requirement text and related data input by the user. Based on an enhanced semantic parsing large language model in the field of clinical epidemiology, it extracts and transforms the key elements of observational studies in the text to generate a research protocol that conforms to clinical epidemiology standards. The key elements of observational studies include patient inclusion and exclusion criteria, exposure and control group definitions, outcome definitions, confounding variables, bias control, and statistical analysis methods. The enhanced semantic parsing large language model in the field of clinical epidemiology is constructed based on a corpus of observational study knowledge literature, and the base large language model is formed after multiple fine-tunings based on this corpus. As an example, the base large language model uses deepseek-r1. The data security intelligent agent is used to review sensitive personal information in the dataset uploaded by users. It reviews the data through a sensitive information detection mechanism, and if sensitive information is found, it performs desensitization processing on the data. The data management agent receives user-uploaded data to be processed, performs preliminary understanding and evaluation of the data based on the research plan and user needs, determines the rules for data governance and / or feature construction, generates corresponding data processing code from the data processing vertical big model in the backend of the data management agent, calls external data processing tools to run the generated data processing code, performs data governance processing on the data to be processed, and outputs a standard dataset for statistical analysis; the data to be processed includes data processed by the data security agent or compliant data directly uploaded by the user; the data processing vertical big model is obtained by optimizing the base big language model through prompting engineering. As an example, the base big language model uses deepseek-r1; The statistical analysis modeling agent receives the standard dataset output by the data management agent and the research plan generated by the research requirements parsing agent. It then calls a large statistical analysis language model specifically designed for generating statistical software code to generate code for handling confounding factors, bias control, and analytical modeling. Finally, it calls external statistical analysis software to execute the code, perform statistical analysis, and output the statistical analysis results. The large statistical analysis language model is based on a pre-built corpus of statistical analysis code in the field of observational research. It is obtained by fine-tuning the base language model. As an example, the base language model uses Qwen3-coder. The results report generation agent receives the statistical analysis results output by the statistical analysis modeling agent and the research plan generated by the research needs parsing agent. It then uses an enhanced semantic parsing large language model in the field of clinical epidemiology to jointly interpret the results. In accordance with the standard format and logical requirements of clinical epidemiology reports, it generates a clinical evaluation report that includes the research background, purpose, methods, results, and risk of bias assessment. It also provides efficacy difference evaluation and individualized treatment decision suggestions for specific clinical characteristic groups based on age, disease stage, and / or comorbidities.

[0043] As an example in this embodiment, a clinician wants to conduct a study with the initial goal of evaluating the effectiveness of using Tanreqing (a traditional Chinese medicine) versus not using it in preventing multidrug-resistant bacterial infections after hospital admission in a critically ill patient population. However, he is unclear about how to screen patients, define medication, set up follow-up procedures, and use statistical analysis methods in a real-world data environment. He can first input his requirements through the dialog box in this invention: How can I evaluate the effectiveness of using Tanreqing versus not using it in preventing multidrug-resistant bacterial infections after hospital admission in a critically ill patient population using hospital electronic medical record data? What statistical analysis methods should I use? At this point, the work coordination agent determines, based on the user's question, that the user is currently in the early design phase of clinical research and transmits the question to the "research needs analysis agent." This agent can provide a comprehensive research plan for the user's reference based on prior knowledge and user feedback. The user can prepare data according to the research plan and transmit the data through the data upload window of this invention. The "research needs analysis agent" determines, based on the user's actions, that the data should be allocated to the "data security agent" and the "data management agent." After data security monitoring (or processing by the data security agent), the "data management agent" performs data standardization, missing value imputation, etc., forming an analysis dataset that is transmitted to the "statistical analysis and modeling agent." The "statistical analysis agent" will provide a statistical analysis plan based on the data format and research plan, conduct statistical analysis and modeling, call the analysis software, provide analysis results, and transmit them to the "results report generation agent" to summarize and form a clinical evaluation report, providing the clinical efficacy of the target drug, as well as efficacy evaluation content for a specific characteristic population, providing precise decision-making content.

[0044] Example 2 As another preferred embodiment of the present invention, this embodiment is a further detailed supplement and explanation of the technical solution of the present invention based on the above embodiment 1. In this embodiment, as shown in the appendix... Figure 1 As shown, the system also includes a work coordination agent, which is connected to a research requirement analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a result report generation agent. The work coordination agent is used to perform semantic understanding of the task requirements input by the user, parse out the specific steps of the clinical efficacy evaluation task, match the requirements with the corresponding agents based on the functional mapping rules of each agent and determine the execution order, call each agent according to the execution order, monitor the execution status of each agent in real time, handle abnormalities during the execution process and provide feedback for adjustment, and integrate the output results of each agent to generate a clinical evaluation report.

[0045] When a user inputs task requirements into the task coordination agent, the agent first uses a large language model to semantically understand the user's input, parsing out the specific steps of the clinical efficacy evaluation task the user expects to complete, such as analyzing the effectiveness of a drug in a specific population, and determining which data management agent and statistical analysis modeling agent need to be invoked. Based on preset agent function mapping rules, the agent matches the user's requirements with the corresponding agents, clarifying the execution order. Following the predetermined execution order of the agents, the work coordination agent sequentially calls the corresponding agents. When calling each agent, necessary parameters and data are passed through standardized interfaces; for example, when calling the data management agent, de-anonymized data and variable identification rules are transmitted. Simultaneously, the execution status of each agent is monitored in real time to ensure the process proceeds as planned.

[0046] The task coordination agent dynamically allocates tasks to the most suitable agent through an auction algorithm, satisfying the following objectives and constraints: Objective function: ; Constraints: ,and ; in, The utility value of agent j for performing task i is calculated based on the agent's historical task performance and resource utilization. This indicates that task i is assigned to agent j; m represents the total number of agents participating in task assignment, and n represents the total number of independent tasks to be assigned; when the objective function reaches its maximum value at a certain agent, the corresponding task is assigned to that agent; constraints ensure that each task is assigned to only one agent.

[0047] Furthermore, in a typical research process, different research stages have different priorities; therefore, this embodiment introduces priority weights. To adapt to the differences in task priorities at different stages of the research, the objective function is dynamically adjusted. ;in The priority weight for task i is determined by the research phase requirements.

[0048] If an agent encounters an anomaly during execution, such as the data management agent being unable to process data due to an error in data format, the work coordination agent will immediately capture the anomaly information. On one hand, it will promptly provide the user with details of the problem, such as displaying a pop-up dialog box indicating the specific location of the data format error; on the other hand, according to the preset anomaly handling strategy, it will automatically attempt to adjust parameters and re-execute or pause the process to await user intervention. For example, it may automatically correct common format errors and then re-call the data management agent. If the process still fails, it will prompt the user to manually correct the errors and re-upload the data.

[0049] Once all agents have successfully completed their tasks according to the workflow, the work coordination agent integrates the outputs of each agent and presents the final clinical evaluation report to the user according to the report generation logic. Simultaneously, it cleans up temporary data and resources related to this task, ending the entire workflow. If the task is interrupted due to an anomaly and cannot be recovered, the work coordination agent will also record the execution status and anomaly details of this task to facilitate subsequent troubleshooting and process optimization.

[0050] Example 3 As another preferred embodiment of the present invention, this embodiment is a further detailed supplement and explanation of the technical solution of the present invention based on the above-described Embodiment 1 or Embodiment 2. In this embodiment, reference is made to the appendix to the specification. Figure 2 As shown, the research needs analysis agent parses the text description, research data, and possible research plans input by the user, and extracts key elements such as inclusion and exclusion criteria, group definitions, outcome definitions, key variables, bias control and assessment, and statistical analysis methods.

[0051] The research needs analysis agent first constructs an enhanced semantic parsing large language model for the clinical epidemiology field based on a vast real-world research knowledge literature base. Through the large knowledge base and fine-tuning that conforms to the style of the professional field, the illusion of a large model is reduced. The model is then used to generate simulated corpora, and combined with real-world corpora for secondary fine-tuning, so as to realize the generation of observational research designs based on complex causal inference methods.

[0052] The research needs analysis agent constructs an "enhanced semantic parsing large language model for complex causal inference in the clinical epidemiology field" to meet users' research design needs for observational studies based on real-world data and provides key elements based on user queries. One way this enhanced semantic parsing large language model for the clinical epidemiology field is implemented is by constructing a professional knowledge base: collecting a large amount of publicly available registration information and research papers from observational studies, performing preprocessing operations such as cleaning and annotation on this text data to improve data quality and consistency, and constructing a JSON-formatted optimized corpus for the clinical epidemiology large language model. This corpus, based on the characteristics of observational studies, is divided into the following key research elements: inclusion and exclusion criteria, definitions of exposure and control groups, definition of outcome, confounding variables, bias control, and statistical analysis methods. Regular expressions were used to process unstructured text, removing special characters and redundant spaces; TF-IDF algorithm was used to detect and delete duplicate paragraphs; medical dictionaries (such as UMLS) were used to correct spelling errors; KNN algorithm was used to complete missing fields; a labeling framework containing 6 key elements was established (inclusion criteria, group definition, outcome definition, confounding variables, bias control, and statistical methods), and BIO labeling was used to annotate each sentence with entity annotation, with the labeling granularity refined to a three-level structure of "population characteristics - intervention type - statistical model"; ClinicalBERT technology was used to tokenize the text, generating an input sequence in the format of [CLS] + sentence + [SEP], where [CLS] is used for classification tasks and [SEP] separates different research elements; the WordPiece algorithm was used to map words into a vector space, and finally a JSON format triplet storage structure (research area, research type, element content) was constructed.

[0053] In addition to building a knowledge base, the parameters of the base model are updated through full-scale fine-tuning techniques to obtain an enhanced semantic parsing large language model for the clinical epidemiology field. This fine-tuning enables the model to better understand the professional terminology and concepts in the clinical epidemiology field. Figure 2 Based on the constructed JSON corpus, the original corpus is used as the model output, and the research abstract is used as the user input. The input format is: input=f”[INST]{instruction}[SEP]{study_abstract}[ / INST]” output=f”{element_dict}”. The user instruction is: “As a clinical epidemiologist, please design a research protocol for {research area / disease area} targeting {patient type} using {drug name} regarding {study outcome}, and comply with the requirements of the strobe statement.”

[0054] After obtaining the large language model, a simulation study plan based on observational research is generated using the large language model according to the target experimental simulation framework. Simultaneously, answers are generated based on different cue word engineering techniques: zero-shot and few-shot, and scores are applied using the formula: High-quality, finely tuned corpora are selected based on scores.

[0055] Based on the selected high-scoring corpus, and building upon the six key elements mentioned above, additional annotations are added for time-instance determination rules (nested trials, grace periods, cloning, etc.), simulated randomization grouping, and intervention allocation. Simultaneously, the newly added corpus, "differences from observational studies," is used as source material for fine-tuning the large language model in the previous step. Through these two steps, an enhanced semantic parsing large language model for clinical epidemiology, applicable to complex causal inferences, is constructed.

[0056] As another implementation method of this embodiment, in addition to utilizing the enhanced semantic parsing large language model in the field of clinical epidemiology, domain-specific models trained based on traditional machine learning algorithms can also be used. For example, classification models built using algorithms such as support vector machines (SVM) and random forests can extract key elements from user input text, although they may be slightly inferior in terms of semantic understanding depth and complex text processing capabilities.

[0057] Example 4 As another preferred embodiment of the present invention, this embodiment is a further detailed supplement and explanation of the technical solution of the present invention based on the above-described embodiments 1, 2, or 3. In this embodiment, reference is made to the appendix to the specification. Figure 3 As shown, the data security intelligent agent uses a pre-built sensitive information corpus to construct a large-scale vertical model for detecting sensitive information in sensitive data. If sensitive information is found, the user is prompted to process and re-upload the data, or an encryption tool is invoked to de-identify the data. Figure 4 ).

[0058] Prompt engineering is a method specifically designed to optimize large language models. It guides the model to generate more specialized and targeted text by constructing domain-specific prompts, thereby achieving efficient utilization of large language models. This module primarily addresses privacy anonymization for medical big data and includes the following steps: a) Using existing data sources, construct a prompt word library that includes privacy fields such as patient name, ID number, date of birth, registered address, residential address and / or chief complaint information in hospital admission records; b) Use a base-based large language model to parse, classify, and / or segment long medical texts; c) Input the prompt words and the categorized and / or segmented medical text into the base big language model to identify suspected privacy information columns in the medical data; d) Based on regular expressions, determine whether the privacy information column identified by the base large language model is accurate. If it is not accurate, return an error message and optimize the prompt words until the identification is accurate, thus obtaining the sensitive information detection vertical large model.

[0059] e) The desensitization rules are set as follows: For patient names and ID numbers, a unique identification code is generated using the MD5 encryption method, and traces of algorithm generation are eliminated. Data security is ensured based on localized processing. For registered addresses and residential addresses, all text is replaced with *. For contact information, the last 4 digits are retained while other numbers are replaced with *.

[0060] f) After the desensitization process is completed, a privacy-free column will be generated to replace the data column in the original text, and then returned to the user.

[0061] As another implementation method of this embodiment, in addition to constructing a dedicated prompt system and a large-scale vertical model for sensitive information detection for data review, a sensitive information detection rule engine based on regular expression matching can also be used. This engine scans the data according to preset sensitive information format rules (such as the numerical arrangement rules of ID card numbers and phone numbers), but its ability to identify complex and hidden sensitive information may be relatively weak.

[0062] In addition to MD5, quasi-identifier anonymization, and existing data desensitization tools, homomorphic encryption technology can be used to encrypt data while allowing for a certain degree of computational operations, thus better protecting data privacy. However, the computational performance overhead may be greater.

[0063] Example 5 As another preferred embodiment of the present invention, this embodiment is a further detailed supplement and explanation of the technical solution of the present invention based on the above-described embodiments 1, 2, 3, or 4. In this embodiment, reference is made to the appendix to the specification. Figure 4 As shown, the data management agent comprises three parts. First, it performs preliminary understanding and evaluation of the data, including determining data type and missing value status. Then, based on user requirements, it determines rules for data governance and feature construction, including variable transformation, data standardization, derived feature generation, and missing value imputation. Finally, it calls data processing tools to perform the above operations, generating a dataset for subsequent statistical analysis. This module constructs an interactive interface between the user and the large model to upload data expressing needs. It leverages the understanding and generation capabilities of the backend large model to generate data evaluation and processing code, and calls data processing software to achieve the purpose of data management. The specific process by which the data management agent outputs a standard dataset for statistical analysis is shown below: S301. Use a large vertical data processing model to perform a preliminary understanding and evaluation of the data to be processed. The preliminary understanding and evaluation includes determining the data type and the situation of missing data. S302. Utilize the vertical large-scale data processing model to analyze user needs and determine the rules for data governance and / or feature construction, including data transformation, data standardization, missing value imputation, and / or derived feature generation. S303, Generate data processing code that satisfies data governance and / or feature construction rules from a large vertical data processing model; S304 integrates a multilingual parser through the MCP server, isolates the execution environment through Docker, calls external data processing tools, runs the generated data processing code, and generates a standard dataset for statistical analysis.

[0064] Enhancing the ability of the Deepseek-R1 large language model to understand and generate data processing requirements through prompting engineering, system prompt keywords are constructed for data transformation and missing data tasks. The system prompt content for data transformation is: "As a data analyst, you are preparing to process a data transformation task. The columns to be processed include {column name}. The transformation mode is to sort the text data and then perform the transformation through one-hot encoding. For numerical data, the transformation is performed from 1 to the category length according to the user's requirements, and the data is returned in {data format}. For example: {Occupation: Teacher, Worker, Doctor, Other; Occupation: 1,2,4,3}". For data imputation, the system prompts: "You are a data analyst. You are preprocessing data containing null values ​​for attributes. The columns to be processed include: {column name}. Please help me decide how to impute null values ​​for each attribute based on content, statistics, and semantics. Mean imputation is represented by 1, median imputation by 2, pattern imputation by 3, introducing new categories to represent the unknown by 4, and interpolation imputation by 5. Only data is returned in JSON format, without any other explanation or content, and then the corresponding function is called to impute the data."

[0065] Users upload analytical data through an interactive interface and express their data processing needs in text form. The user's question needs to provide "column name", "data format" and "expected data processing method" as input to the large language model. After being uploaded, the large language model stores the data in the data storage module and generates data governance code text according to the user's needs.

[0066] The data management agent integrates a multilingual parser through the MCP server, isolates the execution environment through Docker, calls external data processing software, runs the generated data governance code, and generates analysis datasets that meet user needs.

[0067] As another implementation method in this embodiment, besides calling external general-purpose data processing software such as Python and R, customized data processing programs can be developed or commercial data processing software such as SAS and SPSS can be used. However, this may face software cost and compatibility issues. In addition to the knowledge association based on the large model, the feature construction logic can introduce an expert rule base to assist in feature construction. Domain experts pre-define common and clinically significant feature construction rules, which, combined with the feature construction logic generated by the large model, further enhance the analytical value of the features.

[0068] Example 6 As another preferred embodiment of the present invention, this embodiment is a further detailed supplement and explanation of the technical solution of the present invention based on the above-described embodiments 1, 2, 3, 4, or 5. In this embodiment, reference is made to the appendix to the specification. Figure 5 As shown, the statistical analysis modeling agent is based on a pre-built corpus of statistical analysis code in the professional field of observational research. In particular, it constructs a code framework for professional statistical knowledge such as confounding factor adjustment, bias control, and causal inference. It mainly uses R language for targeted model fine-tuning to build an enhanced large language model for observational research statistical analysis, which can accurately understand the statistical analysis requirements and generate corresponding high-quality code. Figure 5 ).

[0069] a) Construct a large-scale model for generating code for statistical analysis in observational studies. This model is built by collecting functions and documentation in the R language to create a code knowledge base. The code knowledge base is formatted as follows: {function category}, {function name}, {parameters}, {detailed description}, {function example}. b) Construct a large-scale model fine-tuning corpus using sample data and sample analysis schemes. The corpus format is: {Data Column Information}, {Modeling Requirements}, {Modeling Code}, {Modeling Instructions}, {Output Results}, {Result Explanation}, {Output Result Formatting}; c) For code generation processes in complex causal inference methods such as target simulation experiments, domain-specific fine-tuning can be performed, and multiple different subsequent process codes can be returned depending on the task; d) Generate code for the corresponding task based on the user's requirements, and save the output code in intermediate storage or return it to the user.

[0070] e) The statistical analysis modeling agent constructs specialized tasks suitable for complex causal inference based on user needs: including a dynamic causal inference module based on meta-learning, which adjusts the matching strategy in real time by learning the weight distribution of confounding variables. , IPM is a measure of the maximum mean difference. For dynamic weighting coefficients, X represents the model parameters; X represents the covariates from different dimensions in the research data; Y represents the outcome variable, i.e., the true value. This represents the average value in the statistical analysis dataset D; This represents the loss function, used to measure the difference between the predicted value of the prediction model and the true value Y; This represents the model's predicted value. This indicates the distribution of covariate X in a population receiving a certain treatment; This represents the distribution of covariate X in the population that did not receive a certain treatment, where T=0 represents the population that did not receive a certain treatment and T=1 represents the population that received a certain treatment. And effect estimation based on the target experiment simulation framework: , in, This represents the average treatment response and is used to measure the average difference in efficacy between the intervention and the control. This indicates that under covariate X, the intervention... The corresponding latent outcome mean function, ; It represents the distribution measure of the covariate X, reflecting the distribution pattern of patient characteristics in real-world data; Indicates the outcome of receiving a certain treatment; It indicates an outcome of not accepting a certain treatment; For the target experiment simulation framework, this method combines reinforcement learning to achieve dynamic parameter adjustment and real-time optimization of key parameters in the TTE framework (such as randomization window, grace period, and cloning rule). This is achieved by introducing state... S t :

[0071] Wherein, KL: Kullback-Leibler divergence of covariate distributions between the exposed and control groups; ATE(t): average treatment effect of the current simulation step; Bias(t): bias measure.

[0072] The reward function is set to r t :

[0073] in, This represents a known average treatment effect, used to measure... Is it accurate? , and All are weighting coefficients. This represents the simulation convergence speed; finally, the optimal policy is learned through a deep Q-network.

[0074] In observational studies, the method for achieving causal inference is to construct a comparable research population. This step requires matching treated and untreated groups, comparing outcomes among groups with similar characteristics. Therefore, this part mainly presents how the agent uses data features to control the model used, where λ>0 is an adjustable coefficient. When λ=0, the formula degenerates into ordinary supervised learning, only concerned with prediction accuracy and not considering causal balance. As λ increases, the model will place greater emphasis on covariate balance, sacrificing some prediction accuracy for stronger causal inference ability. In other words, it combines the prediction task (ℓ(Y,f θ(X))) and the causal balance task (IPM(PX∣T=1,PX∣T=0)) through a unified objective function for joint optimization.

[0075] f) Call external statistical analysis software, such as R, import the generated code into the software, and use the software's calculation engine to execute the code to perform corresponding statistical analysis, including descriptive statistics, correlation analysis, hypothesis testing, etc., and finally output detailed statistical analysis results, such as statistical statistics, etc. p Values, confidence intervals, etc.

[0076] As another implementation method of this embodiment, in addition to fine-tuning the domain-enhanced large language model using R software modeling code, statistical analysis code based on the Python language ecosystem (such as code written using libraries like stats models and scikit-learn) can be used as the fine-tuning corpus to train a code generation model adapted to different programming environments, thus meeting the technical stack needs of different teams. Based on existing statistical analysis methods, more cutting-edge causal inference algorithms, such as Double Machine Learning, can be integrated to address more complex confounding factors and bias control scenarios, but this requires a corresponding increase in model training and computational resource investment.

[0077] Example 7 As another preferred embodiment of the present invention, this embodiment further supplements and elaborates on the technical solution of the present invention based on the above-described embodiments 1, 2, 3, 4, 5, or 6. In this embodiment, referring to Appendix 6 of the specification, the result report generation agent combines the research scheme with the results of statistical analysis, summarizes the clinical epidemiology domain enhanced semantic parsing large language model obtained by fine-tuning the research needs parsing agent, and forms the final clinical evaluation report, as shown in the appendix. Figure 6 As shown.

[0078] a) The statistical analysis results generated by the statistical analysis modeling agent and the research plan generated by the research needs parsing agent are used as input. Based on the standardized format and logical requirements of clinical epidemiological reports, logical rules for report generation are established, clarifying the content and presentation order of each section, such as study overview, baseline description, and main analysis results. For subgroup and specific analysis results, population characteristics, intervention effects, and statistically significant covariates are summarized based on the data to provide recommendations for precise decision-making. b) Through the inductive and summarizing function of the enhanced semantic parsing large language model in the field of clinical epidemiology, the input statistical analysis results and research plan are integrated and transformed according to the preset report generation logic to generate a fluent, accurate clinical evaluation report text that meets the requirements of medical professionals. The report content covers key parts such as the background, purpose, methods, results and risk of bias assessment of the study, and uses visualization charts and other methods to help present the analysis results for easy and intuitive interpretation by users.

[0079] As another implementation method of this embodiment, in addition to following the preset report generation logic rules, template engine technology can be introduced to provide a variety of report templates for users to choose from. Based on different research types and audience needs, diverse and customized clinical evaluation reports can be quickly generated, while the summary content of the large model can be incorporated into the corresponding templates. In addition to presenting results in common formats such as text and charts, interactive visualization of results can be added, such as developing a web-based interactive data visualization interface, allowing users to dynamically filter and explore the analysis results, enhancing the readability and usability of the results.

[0080] This invention aims to create a user-deployable observational research multi-agent system based on a complex causal inference framework. It incorporates knowledge distillation with causal constraints to reduce the overall model deployment size, while introducing distillation loss to ensure efficient parsing of complex causal inference results. As shown in the following formula, the goal is to minimize the KL divergence, making the output distribution of the sub-models approximate the output distribution of the original model as closely as possible: ; ; The loss function incorporates causal inference constraints, where IPM is used to measure the difference in covariates between the exposed and control groups. For IPM, the maximum mean difference is used as the measure, i.e.:

[0081] in F It is a family of functions in the regenerating kernel Hilbert space (RKHS).

[0082] The specific calculation formula is as follows: 。

Claims

1. A multi-agent real-world clinical efficacy evaluation and precision decision-making system, characterized by: The system includes a research needs analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a results report generation agent; each agent can be invoked independently or through collaborative operation to complete the entire process of clinical efficacy evaluation from research needs input to clinical evaluation report output. The research requirement parsing agent receives the research objective or requirement text and related data input by the user. Based on an enhanced semantic parsing large language model in the field of clinical epidemiology, it extracts and transforms the key elements of observational studies in the text to generate a research protocol that conforms to clinical epidemiology standards. The key elements of observational studies include patient inclusion and exclusion criteria, exposure and control group definitions, outcome definitions, confounding variables, bias control, and statistical analysis methods. The enhanced semantic parsing large language model in the field of clinical epidemiology is formed by constructing a corpus based on an observational study knowledge literature base, and then fine-tuning the base large language model multiple times based on this corpus. The data security intelligent agent is used to review sensitive personal information in the dataset uploaded by users. It reviews the data through a sensitive information detection mechanism, and if sensitive information is found, it performs desensitization processing on the data. The data management agent receives user-uploaded data to be processed, performs preliminary understanding and evaluation of the data based on the research plan and user needs, determines rules for data governance and / or feature construction, generates corresponding data processing code from the data processing vertical big model in the backend of the data management agent, calls external data processing tools to run the generated data processing code, performs data governance processing on the data to be processed, and outputs a standard dataset for statistical analysis; the data to be processed includes data processed by the data security agent or compliant data directly uploaded by users; the data processing vertical big model is obtained by optimizing the base big language model through prompting engineering; The statistical analysis modeling agent receives the standard dataset output by the data management agent and the research plan generated by the research requirements parsing agent. It then calls a large statistical analysis language model specifically designed for generating statistical software code to generate code for handling confounding factors, bias control, and analytical modeling. Finally, it calls external statistical analysis software to execute the code, perform statistical analysis, and output the statistical analysis results. The large statistical analysis language model is obtained by fine-tuning the base language model based on a pre-built corpus of statistical analysis code in the field of observational research. The results report generation agent receives the statistical analysis results output by the statistical analysis modeling agent and the research plan generated by the research needs parsing agent. It then uses an enhanced semantic parsing large language model in the field of clinical epidemiology to jointly interpret the results. In accordance with the standard format and logical requirements of clinical epidemiology reports, it generates a clinical evaluation report that includes the research background, purpose, methods, results, and risk of bias assessment. It also provides efficacy difference evaluation and individualized treatment decision suggestions for specific clinical characteristic groups based on age, disease stage, and / or comorbidities.

2. The multi-agent real-world clinical efficacy evaluation and precise decision-making system as described in claim 1, characterized in that: The system also includes a task coordination agent, which is connected to a research requirement analysis agent, a data security agent, a data management agent, a statistical analysis and modeling agent, and a result report generation agent. The task coordination agent is used to perform semantic understanding of the task requirements input by the user, parse out the specific steps of the clinical efficacy evaluation task, match the task requirements with the corresponding agents based on the functional mapping rules of each agent and determine the execution order, call each agent according to the execution order, monitor the execution status of each agent in real time, handle abnormalities during the execution process and provide feedback for adjustment, and integrate the output results of each agent to generate a clinical evaluation report.

3. The multi-agent real-world clinical efficacy evaluation and precise decision-making system as described in claim 2, characterized in that: The task coordination agent dynamically allocates tasks to the most suitable agent through an auction algorithm, satisfying the following objectives and constraints: Objective function: ; Constraints: ,and ; in, The utility value of agent j for performing task i is calculated based on the agent's historical task performance and resource utilization. This indicates that task i is assigned to agent j; m represents the total number of agents participating in task assignment, and n represents the total number of independent tasks to be assigned; when the objective function reaches its maximum value at a certain agent, the corresponding task is assigned to that agent.

4. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in claim 3, characterized in that: The work coordination agent introduces priority weights. To adapt to the differences in task priorities at different stages of the research, the objective function is dynamically adjusted. ;in The priority weight for task i is determined by the research phase requirements.

5. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in any one of claims 1-4, characterized in that: The multiple fine-tuning of the enhanced semantic parsing large language model in the clinical epidemiology field includes the following steps: S101. Corpus Preprocessing: Collect publicly available registration information and research papers from observational studies; clean the text data, correct spelling errors, and / or complete missing fields; establish an annotation framework including population inclusion and exclusion criteria, exposure and control group definitions, outcome definitions, confounding variables, bias control, and statistical analysis methods; use BIO annotation to annotate sentences with entity annotation, refining the annotation granularity to a three-level structure of population characteristics, intervention type, and statistical model; use Clinical BERT technology to tokenize the text, generating input sequences in [CLS]+sentence+[SEP] format; map words to vector space using the WordPiece algorithm; and construct a corpus with a JSON format triple storage structure, which includes research area, research type, and element content. S102, First Fine-tuning: Based on the constructed corpus, the corpus is used as the model output, and the research abstract is used as the model input. The base model parameters are updated using full fine-tuning technology to form a preliminary large language model adapted to the field of clinical epidemiology. S103, Simulation Corpus Generation: Using the large language model initially adapted to the clinical epidemiology field in step S102, an observational study simulation scheme is generated according to the target trial simulation framework. Answers are generated based on zero-shot and few-shot cue word engineering and scored using a scoring formula to screen high-quality simulation corpus. The scoring formula is as follows: ; S104. Second Fine-tuning: Combine the high-quality simulated corpus selected in step S103 with the real corpus in step S101, supplement the zero-time determination rules, simulated randomization grouping, intervention allocation element annotation, and corpus related to differences with observational studies, and perform a second fine-tuning of the large language model initially adapted to the clinical epidemiology field to form an enhanced semantic parsing large language model for the clinical epidemiology field that supports complex causal inferences.

6. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in any one of claims 1-4, characterized in that: The data security intelligent agent identifies sensitive information in user-uploaded datasets using a vertical large-scale sensitive information detection model, and applies pre-approved desensitization rules to desensitize the data containing sensitive information. After desensitization, it generates de-privacy columns to replace the data columns in the original text and returns the data to the user. The construction process of the vertical large-scale sensitive information detection model is as follows: S201. Using existing data sources, construct a prompt word library that includes privacy fields such as patient name, ID number, date of birth, registered address, residential address and / or chief complaint information in hospital admission records; S202. Use the base-based large language model to parse, classify, and / or segment long medical texts; S203. Input the prompt words and the categorized and / or segmented medical text into the base big language model to identify suspected privacy information columns in the medical data; S204. Based on regular expressions, determine whether the privacy information column identified by the base large language model is accurate. If it is not accurate, return an error message and optimize the prompt words until the identification is accurate, thus obtaining the sensitive information detection vertical large model.

7. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in any one of claims 1-4, characterized in that: The data management agent utilizes prompting engineering to improve the understanding of data processing requirements and code generation capabilities of the base large language model, forming the data processing vertical large model; the data management agent integrates a multi-language parser through the MCP server, isolates the execution environment through Docker, calls external data processing tools, runs the generated data processing code, and generates a standard dataset for statistical analysis.

8. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in claim 7, characterized in that: The specific process by which the data management agent outputs a standard dataset for statistical analysis is as follows: S301. Use a large vertical data processing model to perform a preliminary understanding and evaluation of the data to be processed. The preliminary understanding and evaluation includes determining the data type and the situation of missing data. S302. Utilize the vertical large-scale data processing model to analyze user needs and determine the rules for data governance and / or feature construction, including data transformation, data standardization, missing value imputation, and / or derived feature generation. S303, Generate data processing code that satisfies data governance and / or feature construction rules from a large vertical data processing model; S304 integrates a multilingual parser through the MCP server, isolates the execution environment through Docker, calls external data processing tools, runs the generated data processing code, and generates a standard dataset for statistical analysis.

9. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in any one of claims 1-4, characterized in that: The construction process of the statistical analysis large language model is as follows: S401. Build a code knowledge base by collecting functions and documentation in the target programming language; S402. Construct a statistical analysis code corpus for observational research in a professional field using sample data and sample analysis schemes; S403. Fine-tune the code generation process in complex causal inference methods within the domain to obtain a large language model for statistical analysis.

10. The multi-agent real-world clinical efficacy evaluation and precision decision-making system as described in any one of claims 1-4, characterized in that: The statistical analysis modeling agent constructs a dedicated task suitable for complex causal inference based on user needs. This is a dynamic causal inference module based on meta-learning, which adjusts the matching strategy in real time by learning the weight distribution of confounding variables. ; IPM is a measure of the maximum mean difference. For dynamic weighting coefficients, X represents the model parameters; X represents the covariates from different dimensions in the research data; Y represents the outcome variable, i.e., the true value. This represents the average value in the statistical analysis dataset D; This represents the loss function, used to measure the difference between the predicted value of the prediction model and the true value Y; This represents the model's predicted value. This indicates the distribution of covariate X in a population receiving a certain treatment; This represents the distribution of covariate X in the population that did not receive a certain treatment, where T=0 represents the population that did not receive a certain treatment and T=1 represents the population that received a certain treatment. And effect estimation based on the target experiment simulation framework: ; in, This represents the average treatment response and is used to measure the average difference in efficacy between the intervention and the control. This indicates that under covariate X, the intervention... The corresponding latent outcome mean function, ; It represents the distribution measure of the covariate X, reflecting the distribution pattern of patient characteristics in real-world data; Indicates the outcome of receiving a certain treatment; It indicates an outcome of not accepting a certain treatment; For the target experiment simulation framework, reinforcement learning is used to achieve dynamic parameter adjustment and real-time optimization of key parameters in the TTE framework by introducing state. , Where KL represents the Kullback-Leibler divergence of the covariate distributions between the exposed group and the control group; This represents the average treatment effect of the current simulation step. Indicates a measure of bias; The reward function is set to ;in, This represents a known average treatment effect, used to measure... Is it accurate? , and All are weighting coefficients. This represents the simulation convergence speed; finally, the optimal policy is learned through a deep Q-network.

Citation Information

Patent Citations

  • Clinical rapid evidence-based decision-making system and method

    CN117393151A

  • Reinforcement learning algorithm model construction method and equipment for medical data

    CN117521852A

  • Multi-agent interactive efficient data analysis system

    CN119988421A

  • Multi-agent medical diagnosis and treatment system

    CN120126807A

  • Life omics research method and device based on artificial intelligence, equipment and medium

    CN120748512A

Cited By

  • New traditional Chinese medicine research and development conversion decision-making method and system and electronic equipment

    CN121687560A

  • Multi-agent medical care teaching virtual tutor system, teaching method and medium

    CN122347892A

  • Multi-agent medical care teaching virtual tutor system, teaching method and medium

    CN122347892B