Acquisition and arrangement condition optimization method based on key factor mining, medium and equipment
By generating an initial set of inclusion and exclusion conditions, constructing a multi-dimensional database queue, and analyzing feature contribution, this approach addresses the shortcomings of existing technologies in multi-omics data fusion and dynamic optimization for inclusion and exclusion condition optimization, thereby achieving efficient and precise design of clinical trials.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YUYANG ZHISHU (BEIJING) TECHNOLOGY CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods for optimizing inclusion and exclusion conditions fail to effectively integrate multi-omics data, lack quantitative support, and have weak dynamic optimization capabilities, resulting in prolonged clinical trial recruitment cycles, high costs, low success rates, and difficulty in achieving precise design.
An initial set of inclusion and ranking conditions is generated by parsing the target text. An initial simulated queue is constructed by combining multi-dimensional database retrieval. The feature matrix of the panoramic object is analyzed. The contribution of feature variables is evaluated by using an interpretable machine learning model. The inclusion and ranking conditions are dynamically adjusted to meet the preset optimization conditions.
Significantly reduce the costs and time required for early-stage clinical trials, enhance the scientific rigor and relevance of inclusion and exclusion condition design, improve the success rate of clinical trials, and achieve efficient and feasible designs in the context of precision medicine.
Smart Images

Figure CN121964020A_ABST
Abstract
Description
A method, medium, and device for optimizing absorbance and scattering conditions based on key factor mining. Technical Field
[0001] This invention relates to the field of experimental design technology, and in particular to a method, medium, and equipment for optimizing absorbance and scattering conditions based on key factor mining. Background Technology
[0002] With the deep penetration of artificial intelligence technology into the field of pharmaceutical research and development, clinical trials, as a core part of new drug marketing, are gradually transforming from traditional experience-driven to data-driven approaches. As a key element of clinical trial design, the scientific validity of inclusion and exclusion criteria directly determines the efficiency of subject recruitment, the reliability of trial results, and the success or failure of research and development.
[0003] Currently, existing methods for optimizing inclusion and exclusion conditions have made some progress. Some technologies have provided preliminary support for setting inclusion and exclusion conditions by integrating clinical data to construct patient cohorts or by using historical trial databases as a basic reference. However, in practical applications, existing methods still have many insurmountable shortcomings: First, the data integration dimensions are limited, mostly relying only on clinical phenotypic data, failing to effectively integrate multi-omics data such as genomics, proteomics, and single-cell omics, making it difficult to comprehensively capture the complex molecular mechanisms of disease progression and treatment response, resulting in a lack of in-depth biological basis for inclusion and exclusion conditions; Second, the identification of key factors lacks quantitative support. Existing technologies are unable to systematically mine and quantify key features affecting clinical outcomes, relying heavily on manual experience to select indicators, resulting in insufficient correlation between inclusion and exclusion conditions and trial objectives; Third, the dynamic optimization capability is weak, unable to adjust inclusion and exclusion conditions in real time based on patient characteristic distribution and key factor analysis results, making it difficult to achieve a balance between patient recruitment scale and trial scientific rigor; Fourth, feature association analysis is insufficient, failing to establish a clear mapping relationship between feature variables and key clinical outcomes, resulting in a lack of precise data-driven guidance for the optimization of inclusion and exclusion conditions. These problems directly lead to longer clinical trial recruitment cycles, higher R&D costs, and difficulty in guaranteeing trial success rates, severely restricting the translation efficiency of innovative therapies.
[0004] Therefore, how to achieve precise and scientific design of emission standards has become an urgent problem to be solved. Summary of the Invention
[0005] To address the aforementioned technical problems, the present invention adopts a method for optimizing inclusion and exclusion conditions based on key factor mining. This method includes the following steps: S10, parsing the target text to generate an initial set of inclusion and exclusion conditions, wherein the target text includes experimental grouping criteria, inclusion conditions, and exclusion conditions.
[0006] S20, based on the initial set of inclusion and exclusion conditions, retrieve matching object data from at least one database to obtain an initial simulation queue and baseline feature distribution, wherein the baseline feature distribution includes at least two of the following: age distribution, gender distribution, disease stage distribution, and biomarker expression distribution.
[0007] S30: Perform pre-processing on the object data related to the target domain corresponding to the initial simulation queue to obtain the panoramic object feature matrix.
[0008] S40, Analyze the feature matrix of the panoramic object to obtain the contribution of each feature variable to the prediction of the preset key outcome, wherein the preset key outcome includes at least one of effect response, survival time, and adverse reaction.
[0009] S50: If the initial simulation queue, baseline feature distribution, and key feature variable list generated according to contribution satisfy the preset optimization conditions, then the current initial inclusion and exclusion condition set is determined as the target inclusion and exclusion condition set; otherwise, the initial inclusion and exclusion condition set is adjusted according to the initial simulation queue, baseline feature distribution, and key feature variable list, and the adjusted inclusion and exclusion condition set is used to replace the initial inclusion and exclusion condition set, and the process returns to step S20.
[0010] The present invention also provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described method for optimizing the inclusion and exclusion conditions based on key factor mining.
[0011] The present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0012] This invention offers at least the following advantages: By parsing target text to generate a structured initial inclusion / exclusion condition set, and combining this with database retrieval to construct an initial simulated cohort and analyze baseline feature distribution, a representative research population can be obtained without real-world subject recruitment, significantly reducing the time and manpower costs in the early stages of clinical trials; by integrating multi-dimensional data such as clinical phenotypes, genomics, and proteomics to construct a panoramic object feature matrix, it overcomes the limitations of traditional single baseline features, achieving comprehensive coverage of object features and providing a comprehensive and systematic data foundation for key factor mining; by analyzing the contribution of each feature variable in the panoramic object feature matrix to the preset key outcomes, it accurately identifies key feature variables strongly correlated with efficacy, survival, and adverse reactions, shifting inclusion / exclusion condition optimization from experience-driven to data-driven, thus improving the scientific rigor and relevance of inclusion / exclusion condition design; through comprehensive judgment and dynamic adjustment of the initial simulated cohort, baseline feature distribution, and key feature variable list, a closed-loop optimization process is formed, ensuring that the target inclusion / exclusion condition set balances cohort representativeness, feature relevance, and distribution rationality, providing an efficient and feasible technical solution for clinical trial design in the context of precision medicine, thereby improving the success rate and research value of clinical trials. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 is a flowchart of a method for optimizing inclusion and exclusion conditions based on key factor mining provided in Embodiment 1 of the present invention; Figure 2 is a flowchart of a method for optimizing inclusion and exclusion conditions based on experimental simulation provided in Embodiment 2 of the present invention; Figure 3 is a flowchart of a method for joint optimization of inclusion and exclusion conditions provided in Embodiment 3 of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "including," "having," and any variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0017] Example 1 provides a method for optimizing inclusion and exclusion conditions based on key factor mining. As shown in Figure 1, the method for optimizing inclusion and exclusion conditions based on key factor mining includes the following steps: S10, parsing the target text to generate an initial set of inclusion and exclusion conditions, wherein the target text includes experimental grouping criteria, inclusion conditions and exclusion conditions.
[0018] The target text is a natural language text containing the core rules of clinical trial design. It serves as the original input data for generating the initial set of inclusion and exclusion criteria. Its content must clearly define key information such as the clinical trial's screening requirements for participants and grouping rules. For example, "This trial includes patients aged 18-70 years with pathologically confirmed locally advanced non-small cell lung cancer and an ECOG score of 0-1; patients who have previously received targeted or immunotherapy, as well as those with severe liver or kidney dysfunction, are excluded; the trial is divided into an experimental group (receiving XX drug treatment) and a control group (receiving placebo treatment)." Specifically, the target text includes trial grouping criteria, inclusion criteria, and exclusion criteria. The trial grouping criteria, as specified in the target text, are the basis for dividing eligible participants into different experimental groups (e.g., experimental group, control group). This clarifies the grouping logic of participants, provides a basis for subsequent initial simulated cohort grouping and propensity score matching, and ensures that the division between the experimental and control groups meets the requirements of the clinical trial design. For example, treatment regimen grouping: subjects are divided according to the treatment they receive, such as "those receiving targeted therapy are in the experimental group, and those receiving conventional chemotherapy are in the control group"; population characteristic grouping: subjects are divided according to their pathological and molecular characteristics, such as "patients with EGFR-sensitive mutations are in the experimental group, and patients with EGFR wild-type are in the control group". Inclusion criteria are all the positive screening requirements specified in the target text that subjects must meet to participate in the clinical trial. Only subjects who meet all inclusion criteria are eligible to enter the clinical trial. These criteria are used to screen potential subjects who meet the clinical trial indications and are physically fit, ensuring the homogeneity of the enrolled population and laying the foundation for accurate assessment of the trial's efficacy and safety. Examples include "includes individuals aged 18-70 years", "pathologically diagnosed with non-small cell lung cancer", "ECOG score 0-1", and "expected survival ≥ 3 months". Exclusion criteria are reverse screening requirements specified in the target text to exclude subjects from participating in clinical trials. If a subject meets any of the exclusion criteria, even if they meet the inclusion criteria, they must be removed from the clinical trial. This is to exclude subjects whose comorbidities, previous treatment history, physical condition, or other factors may interfere with the trial results or pose a risk to treatment safety, thus ensuring the scientific rigor of the clinical trial and the safety of the subjects.
[0019] By analyzing the experimental grouping criteria, inclusion conditions, and exclusion conditions contained in the target text, several initial inclusion and exclusion conditions are obtained, forming an initial inclusion and exclusion condition set.
[0020] In one specific implementation, S10 includes the following steps: S101, receiving target text input by the user through file upload or manual input, wherein the target text is in natural language form.
[0021] S102, use a preset large language model to parse the target text and generate an initial arrangement condition set including several initial arrangement conditions. The preset large language model is a large language model that has been fine-tuned with medical domain knowledge and optimized with professional prompts. Each initial arrangement condition contains logical nodes and entity nodes. Logical nodes are used to define the AND / OR logical relationship between conditions, and entity nodes contain entity type, attribute name, operator and corresponding value.
[0022] Users can input target text via file upload or manual entry, and real-time editing and modification are supported.
[0023] By inputting medical data such as clinical trial guidelines, disease diagnostic criteria, drug development professional data, and historical clinical trial texts into a pre-selected basic large language model, and combining this with prompts tailored to the parsing needs of clinical trial conditions, the model is explicitly required to identify logical relationships, extract entity types / attributes / operators / values, and adjust model parameters. This allows the large language model to become familiar with medical terminology, the expression habits and logical rules of clinical trial scenarios, and strengthens its adherence to structured output formats, resulting in a preset large language model. This improves the accuracy of parsing medical professional texts and avoids misinterpretation of professional terms by general models. Those skilled in the art will understand that the structure and fine-tuning methods of any existing large language model fall within the protection scope of this invention, and will not be elaborated further here.
[0024] Entity types include demographic information, disease diagnosis, past treatment, biomarkers, etc.; attribute names are specific subcategories of entity types, such as "age" under demographic information, "lung cancer subtype" under disease diagnosis, etc.; operators are used to limit the comparison rules of attributes, such as "between", "≤", "=", "except", etc.; corresponding values are the specific parameters corresponding to the operators, such as "18-65 years old", "0-1 points", "non-small cell lung cancer", etc.
[0025] As described above, by supporting both file upload and manual input, users can flexibly provide target text according to their actual needs, lowering the operational threshold. Fine-tuning the large language model with medical domain knowledge enables it to accurately identify medical terminology and special expressions from clinical trial scenarios, avoiding parsing bias and improving the accuracy of initial inclusion and exclusion conditions. Optimizing and guiding the model to output structured results through professional prompts transforms natural language text into a standardized set of conditions containing logical and entity nodes, providing a computational foundation for subsequent database retrieval and dynamic condition adjustment. Full-process automated parsing replaces the traditional method of manually compiling structured conditions, improving the efficiency of inclusion and exclusion condition generation while reducing logical omissions and parameter errors caused by manual operation.
[0026] S20, based on the initial set of inclusion and exclusion conditions, retrieve matching object data from at least one database to obtain an initial simulation queue and baseline feature distribution, wherein the baseline feature distribution includes at least two of the following: age distribution, gender distribution, disease stage distribution, and biomarker expression distribution.
[0027] Among them, the database is a clinical database that stores data on the entire clinical diagnosis and treatment process of patients (including demographic information, disease diagnosis, treatment history, laboratory tests, biomarker detection, etc.). It can be a single-center hospital database, a multi-center shared database, or a public clinical database (such as TCGA, GEO), etc., and provides the original data source for patients who meet the inclusion and exclusion criteria.
[0028] The initial simulated cohort is a dataset of objects retrieved from the database that meet the initial inclusion and exclusion criteria and have undergone de-privacying and data cleaning. It is a virtual research cohort that simulates the enrollment population in real clinical trials. It is used to replace the population recruitment stage in the early stage of real clinical trials, providing research subjects for subsequent steps such as propensity score matching and effect simulation, thereby reducing the time and cost risks of trial design.
[0029] Baseline feature distribution refers to the numerical distribution or categorical composition of the core baseline features of the subjects in the initial simulated cohort. It is used to visually reflect the population characteristics of the initial cohort and provides a direct basis for determining whether the cohort meets the preset optimization conditions and for subsequent adjustments to the inclusion and exclusion conditions. Among them, baseline features are the quantifiable or categorizable clinical and biological characteristics that already exist when the subjects are included in the simulated cohort. They are key variables for assessing treatment efficacy and controlling confounding factors, including demographic characteristics (age, sex, race, etc.), disease-related characteristics (disease stage, pathological type, tumor size, etc.), biological characteristics (biomarker expression levels, genotyping, etc.), and functional status characteristics (ECOG score, KPS score, etc.).
[0030] Specifically, using a structured initial set of inclusion and exclusion criteria as the screening basis, a retrieval association with the database is established. Through precise matching, object data that meets the criteria is screened to form an initial cohort simulating the clinical trial population. At the same time, the distribution patterns of the core baseline characteristics of the cohort are explored to provide data support for subsequent baseline balance verification and inclusion and exclusion condition optimization, thus realizing a closed loop of transformation of conditions, data, and population characteristics.
[0031] In one specific implementation, S20 includes the following steps: S201, converting the initial set of inclusion and exclusion conditions into a query statement compatible with at least one selected database.
[0032] S202, execute the query and de-privacy operation according to the query statement, and integrate to obtain the initial simulation queue.
[0033] S203, extract the baseline feature distribution corresponding to the initial simulation queue.
[0034] The structured initial set of inclusion and exclusion criteria cannot be directly recognized by the database. Therefore, it is necessary to establish a connection between the criteria and database fields through rule mapping, transforming the abstract filtering conditions into a syntax format supported by the database to ensure that the retrieval logic is accurate and that no conditions are omitted. Specifically, first, the type of the selected database (e.g., relational database MySQL, clinical data platform) and the supported query languages (e.g., SQL, SPARQL) are determined, the retrieval syntax is established, and the attribute names of entity nodes in the initial inclusion and exclusion criteria are mapped to the corresponding field names in the database. The AND / OR logic of logical nodes and the operators of entity nodes are converted into logical operators and comparison operators in the query statement. The mapped fields, logical relationships, and parameter values are integrated to generate a complete query statement, and the statement's syntax correctness and condition completeness (e.g., whether any key filtering requirements are omitted) are verified to ensure that the statement can be correctly parsed and executed by the database.
[0035] The database is retrieved by querying a statement, filtering out all records that meet the criteria. Personal privacy information is removed while retaining core clinical information. The data is then cleaned and integrated into a unified format dataset, ultimately forming the initial simulation queue. This process ensures both the accuracy of the search results and the compliance of data use. Those skilled in the art will recognize that any de-identification operation in the prior art falls within the scope of this invention. Examples include removing direct identifying information (name, ID number, contact information, etc.), and desensitization / pseudo-name processing (encrypting or replacing key identifiers such as medical record numbers with virtual identifiers) to ensure that no specific individual can be associated with the data. These details will not be elaborated upon here.
[0036] From the massive data of the initial simulated cohort, we focus on core baseline features related to clinical trial efficacy evaluation and baseline balancing. Statistical methods are used to analyze their distribution patterns, transforming scattered object data into intuitive descriptions of population characteristics, providing interpretable evidence for subsequent decision-making. Specifically, we select pre-defined core baseline features from the initial simulated cohort, including at least age, sex, disease stage, and biomarker expression. We can also supplement these with ECOG scores, laboratory indicators (such as CEA and creatinine), and metastatic lesion status, depending on the disease domain. For different types of baseline features, we employ corresponding statistical methods. For continuous features (such as age and quantitative biomarker expression), we calculate the mean, standard deviation, median, or quartile range to analyze the central tendency and dispersion of the numerical distribution. For categorical features (such as sex, disease stage, and qualitative biomarker expression), we statistically analyze the frequency and proportion distribution of each category to clarify the population size of different subgroups.
[0037] It should be noted that quantitative biomarker expression refers to biomarkers whose expression levels are expressed as "specific numerical values," allowing for comparisons of magnitude and intensity. Qualitative biomarker expression, on the other hand, refers to biomarkers whose expression levels are expressed as "categories," making numerical comparisons impossible; they can only distinguish between "present / absent," "positive / negative," and "wild-type / mutant."
[0038] Furthermore, the statistical analysis results can be transformed into interpretable forms, such as structured tables (feature name-statistic-value) or visualization charts (histograms, pie charts, box plots), ultimately forming baseline feature distribution data, which is then displayed on the interactive interface.
[0039] As described above, by converting the initial set of inclusion and exclusion conditions into database-compatible query statements, retrieval logic deviations were avoided; by performing de-anonymization operations after retrieval, the personal identity information of the subjects was effectively protected, and the entire data retrieval process complied with medical data privacy regulations, reducing data compliance risks; by extracting core baseline features and analyzing their distribution, the population structure of the initial simulated queue was presented intuitively, making it easy to quickly determine whether the queue meets the preset optimization conditions, and providing direct data support for adjusting inclusion and exclusion conditions.
[0040] S30, perform pre-processing on the object data related to the target domain corresponding to the initial simulation queue to obtain a panoramic object feature matrix, wherein the object data includes at least two of the following: clinical phenotype data, genomics data, proteomics data, and single-cell omics data corresponding to each object.
[0041] Clinical phenotypic data is a set of observable and quantifiable indicators obtained through routine clinical methods such as clinical history taking, physical examination, laboratory testing, and imaging examinations. These indicators are directly related to the subject's disease state, physiological characteristics, and treatment process and are used to reflect the subject's clinical status.
[0042] High-throughput detection technologies are used to obtain data at the biomolecular level. Specifically, genomic data reflects genetic information such as gene sequence variations and copy number changes; proteomics data reflects functional information such as protein expression levels and modification states; and single-cell omics data is data on gene expression and protein levels detected at the single-cell level, which can reflect cellular heterogeneity.
[0043] The panoramic object feature matrix is a two-dimensional data matrix with individual objects as row indices and multi-dimensional feature variables as column indices. Each element in the matrix represents the specific value of a certain feature variable of a certain object. It integrates scattered clinical phenotype, genomics and other data into a unified structured data carrier, which facilitates subsequent batch feature contribution analysis.
[0044] Therefore, this embodiment targets multi-dimensional object data in the target domain, eliminates data noise and technical bias through standardized preprocessing, then selects core variables that are strongly correlated with preset key outcomes through targeted feature extraction, and finally integrates multi-source data in matrix form to achieve a systematic and structured presentation of object features, providing a high-quality data foundation for subsequent key feature contribution analysis.
[0045] In one specific embodiment, S30 includes the following steps: S301, for any object corresponding to the object data, preprocessing is performed on the object data of the current object for each preset type to obtain intermediate data of the current object for each preset type, wherein the preprocessing includes batch effect correction, normalization and quality filtering, and the preset types include at least two of clinical phenotype, genomics, proteomics and single-cell omics.
[0046] S302, extract features from the intermediate data of the current object for each preset type to obtain several feature variable data corresponding to the current object. Among them, the feature variables corresponding to the clinical phenotype include at least three of age, gender, disease stage, past treatment history, ECOG score, and laboratory test results; the feature variables corresponding to genomics include at least two of gene mutation status, tumor mutation burden, gene fusion, copy number variation, and pathway enrichment score; the feature variables corresponding to proteomics include at least one of differentially expressed protein and phosphorylation modification level; and the feature variables corresponding to single-cell omics include at least one of cell subpopulation ratio and differentially expressed gene module score.
[0047] S303 constructs a panoramic object feature matrix with objects as rows and feature variables as columns, based on several feature variable data corresponding to each object.
[0048] Batch effect correction is used to eliminate non-biological differences caused by variations in experimental batches, testing platforms, and operators, thereby avoiding interference from batch differences on the authenticity of feature variables and ensuring that data differences reflect the biological differences of individual subjects rather than experimental technique biases. Normalization is used to transform feature data with different dimensions and value ranges into a unified range (such as [0, 1] or a standard normal distribution) to eliminate analytical bias caused by dimensional differences and ensure the fairness and accuracy of subsequent analyses. Quality filtering is used to remove low-quality and meaningless data, such as deleting samples that failed detection, filtering genes with excessively low expression levels, and removing outliers, to improve the overall quality of the data, reduce the interference of noisy data on subsequent feature analysis, and reduce analytical errors. Those skilled in the art will understand that any batch effect correction, normalization, and quality filtering methods in the prior art fall within the protection scope of this invention, and will not be elaborated further here.
[0049] For different types of intermediate data, the core feature variables are extracted. The data are arranged with objects as rows and feature variables as columns. Each feature variable of each object is filled into the corresponding position in the matrix. If a feature variable of an object has missing values, the mean, median or deletion methods can be used to fill the missing values. The panoramic object feature matrix is then constructed.
[0050] As described above, by integrating clinical phenotypes and multi-omics data, subsequent key feature mining can simultaneously consider the clinical status and molecular mechanisms of the subjects, providing a more comprehensive decision-making basis for the precise optimization of inclusion and exclusion conditions. By performing batch effect correction, normalization, and quality filtering preprocessing on multi-type subject data, technical biases and noise in the original data are effectively eliminated, ensuring the accuracy and reliability of subsequent feature extraction. By constructing a panoramic subject feature matrix, scattered multi-source data are integrated into a structured carrier, and the panoramic subject feature matrix takes into account both clinical practicality and molecular biological depth, breaking through the limitations of traditional single feature analysis, facilitating subsequent batch feature contribution calculations, and improving analysis efficiency.
[0051] S40, analyze the feature matrix of the panoramic object to obtain the contribution of each feature variable to the prediction of the preset key outcome.
[0052] Among them, the pre-specified critical outcome is a core observation indicator that can directly reflect the clinical value, benefit and safety of the investigational drug / therapy in clinical trials or clinical research. Therefore, the panoramic subject feature matrix is analyzed to obtain the contribution of each feature variable to the prediction of the pre-specified critical outcome, which serves as a key basis for evaluating the success or failure of the trial and the effectiveness of the intervention.
[0053] In this embodiment, the pre-defined key outcomes specifically include at least one of the following: response, survival, and adverse reactions. Response refers to the degree of improvement or change in the disease state of the subject after receiving the intervention, and is a core indicator for evaluating the effectiveness of the intervention. Survival is the time span from enrollment in the trial or the start of the intervention until the occurrence of a pre-defined endpoint event, and is an indicator reflecting the long-term benefit of the intervention. Adverse reactions are unexpected medical events related to the intervention that occur during or after the intervention, and are a core indicator for evaluating the safety of the intervention.
[0054] In one specific implementation, S40 includes the following steps: S401, using a preset key outcome as the prediction target, training a preset interpretable machine learning model with a panoramic object feature matrix to obtain a trained interpretable machine learning model.
[0055] S402, based on a trained interpretable machine learning model, uses a model interpretability framework to perform attribution analysis on each feature variable in the panoramic object feature matrix, and obtains the contribution of each feature variable to the prediction of the preset key outcome.
[0056] Interpretable machine learning models are those that possess both high-precision predictive capabilities and the ability to clearly explain "how feature variables affect the prediction results." Unlike traditional "black box" models (such as deep neural networks), they not only predict preset key outcomes but also provide a foundation for subsequent feature contribution attribution analysis, ensuring the reliability and interpretability of the contribution calculation results. Those skilled in the art will recognize that any interpretable machine learning model and its training method in the prior art falls within the scope of this invention, such as logistic regression models, decision tree models, random forest models, gradient boosting tree models, etc., which will not be elaborated upon here. Commonly used models in this embodiment include...
[0057] This embodiment uses a preset key outcome as the prediction target (dependent variable) and feature variables from the panoramic object feature matrix as inputs (independent variables). A machine learning algorithm is used to enable a preset interpretable machine learning model to learn the correlation between the feature variables and the preset key outcome. Specifically, feature variables are extracted from the panoramic object feature matrix as the model input set X, and the corresponding preset key outcome is used as the model output set y (e.g., effect response: effective = 1, ineffective = 0; survival time: censored data + survival time; adverse reaction: occurred = 1, did not occur = 0). The input set X and output set y are divided into a training set (70%-80%) and a validation set (20%-30%) for model training and performance evaluation.
[0058] Then, based on the type of the preset key outcome, an appropriate interpretable model is selected. If the preset key outcome is a binary classification outcome (e.g., effective / ineffective, adverse reaction occurring / not occurring), logistic regression, random forest, or Boost classifier models are selected. If the preset key outcome is a survival outcome (e.g., survival time), tree-based survival analysis algorithms (e.g., XGBoost Survival) are selected. Furthermore, grid search and cross-validation methods are used to optimize the hyperparameters of the interpretable machine learning model (e.g., decision tree depth, number of random forest trees) to ensure that the prediction accuracy of the interpretable machine learning model on the validation set meets the preset standard. Corresponding evaluation metrics are used to verify the model performance. Specifically, AUC, accuracy, and recall are used for binary classification outcomes; C-index (consistency index) is used for survival outcomes. If the model performance does not meet the standard, the process is repeated to adjust feature variables or replace the model until an interpretable machine learning model that meets the accuracy requirements is trained, thus obtaining a trained interpretable machine learning model.
[0059] The model interpretability framework is an analytical toolset used to quantify the impact of feature variables on model prediction results. It is divided into two categories: global interpretation (measuring the contribution of a feature to the overall model) and local interpretation (measuring the contribution of a feature to the prediction of a single sample). It transforms the model's prediction logic into quantifiable contribution values, clarifying the weight of each feature variable in the prediction of preset key outcomes. Attribution analysis, for a trained interpretable machine learning model, analyzes the direction (positive / negative) and degree (contribution) of the impact of each input feature variable on the output prediction results. It is used to accurately identify "which feature variables are the core factors affecting clinical outcomes such as efficacy and survival," providing quantitative standards for key feature selection.
[0060] This embodiment preferentially selects a global interpretability framework for computation, and the calculation method is the sum of the improvement of node purity by the feature variable in all decision trees. Those skilled in the art will know that any existing global interpretability framework and its pre-training method fall within the protection scope of this invention, such as the SHAP global interpretation module (Shapley AdditiveexPlanations, SHAP), the tree model built-in feature importance scoring (Random Forest / XGBoost / LightGBM), etc.
[0061] Specifically, the panoramic object feature matrix is input into a pre-defined model interpretability framework, and the contribution of each feature variable (such as the absolute mean of SHAP values and feature importance scores) is output. The contribution values are standardized (e.g., normalized to the [0, 1] interval) to ensure the comparability of contributions between different feature variables. A correspondence between feature variable names and contribution values is established, and the direction of influence of feature variables on pre-defined key outcomes is marked, including positive contributions (the higher the feature value, the better the effect) and negative contributions (the higher the feature value, the higher the risk of adverse reactions).
[0062] Contribution is a quantitative indicator that measures the degree of influence of a single feature variable on the prediction of a predefined critical outcome. It is usually a non-negative number. The larger the value, the more important the feature variable is to the prediction of the critical outcome. It can be used to distinguish between critical feature variables and redundant feature variables.
[0063] As described above, by incorporating clinical phenotypes and multi-omics features into the same model analysis, key feature mining takes into account both clinical practice and molecular mechanisms, avoiding the limitations of single-dimensional analysis and improving the comprehensiveness of inclusion and exclusion condition optimization. By adopting an interpretable machine learning model instead of a black-box model, the calculation process of feature variable contribution is traceable and the results are verifiable, avoiding the bias of subjective experience judgment and improving the scientific nature of key feature mining. By introducing a model interpretability framework to conduct attribution analysis, the influence of each feature variable on the preset key outcome is accurately quantified, realizing the leap from qualitative association to quantitative contribution and providing a decision-making basis for inclusion and exclusion condition optimization.
[0064] In one specific implementation, S40 further includes the following step: S403, sorting all feature variables in descending order of contribution to obtain a sorting result.
[0065] S404 Select the feature variables whose ranking is lower than the preset ranking threshold in the ranking results to form a list of key feature variables.
[0066] The ranking results include the name of the feature variable, the contribution value, and the ranking number. For example: "Tumor mutation burden (TMB) - contribution 0.85 - ranking 1; ECOG score - contribution 0.72 - ranking 2".
[0067] By setting a preset ranking threshold, key feature variables with the highest contribution are selected, while redundant feature variables with low contribution are eliminated, forming a concise and high-value list of key feature variables. This provides a direct basis for adjusting the inclusion and ranking criteria. The specific value of the preset ranking threshold can be set by the implementer according to the actual situation; for example, the preset ranking threshold can be 3 or 5.
[0068] As described above, by sorting and filtering feature variables according to their contribution, high-value key feature variables and low-value redundant feature variables are effectively distinguished. The resulting list of key feature variables is concise and highly targeted, providing accurate decision-making basis for optimizing inclusion and exclusion conditions.
[0069] S50: If the initial simulation queue, baseline feature distribution, and key feature variable list generated according to contribution satisfy the preset optimization conditions, then the current initial inclusion and exclusion condition set is determined as the target inclusion and exclusion condition set; otherwise, the initial inclusion and exclusion condition set is adjusted according to the initial simulation queue, baseline feature distribution, and key feature variable list, and the adjusted inclusion and exclusion condition set is used to replace the initial inclusion and exclusion condition set, and the process returns to step S20.
[0070] Among them, based on multi-dimensional quantitative judgment criteria, the initial simulation queue, baseline feature distribution and key feature variable list generated under the current set of inclusion and exclusion conditions are comprehensively evaluated. If the preset optimization conditions are met, the set of inclusion and exclusion conditions is determined as the final solution; if not, the inclusion and exclusion conditions are adjusted in reverse based on the evaluation results. Through iterative retrieval and simulation, the inclusion and exclusion conditions are accurately optimized.
[0071] In one specific implementation, the preset optimization conditions include queue size conditions, feature association conditions, and reasonable distribution conditions.
[0072] The queue size condition is: the total number of objects in the initial simulated queue is greater than or equal to the second preset threshold.
[0073] The feature association condition is: the proportion of the number of feature variables in the list of key feature variables that rank in the top N in terms of contribution and are included in the current set of inclusion and exclusion conditions is greater than or equal to the third preset threshold, where N is a positive integer.
[0074] The reasonable distribution condition is that, in at least one preset key dimension, the statistical distribution difference between the baseline feature distribution and the reference population distribution in the target field is less than the fourth preset threshold.
[0075] The preset optimization conditions include three core dimensions: cohort size, feature association, and reasonable distribution. When all three conditions are met, it means that the population selected by the current initial inclusion and exclusion condition set is not only statistically feasible, but also fully covers the key features and is representative of the population. The current initial inclusion and exclusion condition set is directly determined as the target inclusion and exclusion condition set and is used as the final clinical trial inclusion and exclusion plan output. At the same time, the corresponding initial simulated cohort, baseline feature distribution, and list of key feature variables are archived as evidence for the plan.
[0076] Specifically, the cohort size condition is used to ensure that the cohort has sufficient statistical power and to avoid the randomness of subsequent feature analysis and effect prediction results due to an insufficient sample size. The second preset threshold can be set according to the incidence rate of the target disease and the statistical requirements of the clinical trial design. For example, it is usually set at 200 for oncology clinical trials and can be appropriately reduced to 50 for rare disease trials. If the cohort size does not meet the standard, it indicates that the current inclusion and exclusion conditions are too strict, and non-core restrictions need to be relaxed to expand the search scope, such as expanding the age range from 18-65 years old to 18-70 years old.
[0077] Feature association conditions are used to verify whether the current inclusion and exclusion conditions sufficiently include feature variables that have a core impact on the preset key outcomes, ensuring the scientific validity and relevance of the inclusion and exclusion conditions. The N value can be set according to the total length of the list of key feature variables; for example, selecting the top 8 or top 10 feature variables by contribution. The third preset threshold is usually set to 70%, requiring that at least 70% of the top N feature variables be included in the current inclusion and exclusion conditions. Correspondingly, if only 4 of the top 10 feature variables are included in the inclusion and exclusion condition set (40% < 70%), it indicates that the current initial inclusion and exclusion condition set does not sufficiently cover the key features. In this case, the top N feature variables in the key feature variable list that were not included in the initial inclusion and exclusion condition set need to be added as inclusion and exclusion conditions.
[0078] Reasonable distribution conditions are used to ensure that the baseline characteristic distribution of the initial simulated cohort is representative of the population, avoiding the inability to extrapolate clinical trial results to the real-world population due to cohort distribution bias. Specifically, the preset key dimensions are baseline characteristic dimensions strongly correlated with clinical outcomes, such as age, disease stage, and biomarker expression distribution; the reference population distribution is the baseline characteristic distribution data of the real-world population in the target domain, which can be obtained from authoritative databases or published epidemiological studies; and the statistical distribution difference value is a statistical method used to quantify the difference between two distributions, such as using the Kolmogorov-Smirnov test (KS test) for continuous characteristics (age) and the chi-square test for categorical characteristics (disease stage). 2 The fourth preset threshold can be set according to the significance level of the statistical test, for example, a KS test D value < 0.1, or a chi-square test P value > 0.05. If the distribution is not reasonable, the screening range of the core baseline characteristics is adjusted (e.g., the disease stage is expanded from "stage III only" to "stages II-III") to reduce the distribution difference with the reference population.
[0079] Using the adjusted set of inclusion and exclusion conditions, return to step S20 and re-execute the database retrieval, queue construction, feature matrix construction, and contribution analysis processes until the generated solution meets all preset optimization conditions.
[0080] As described above, by setting the cohort size condition, the final inclusion and exclusion scheme can ensure the statistical power of the clinical trial and avoid trial failure due to insufficient sample size; by validating the feature association condition, the inclusion and exclusion conditions fully incorporate feature variables that have a core impact on the preset key outcomes, ensuring the scientific nature and relevance of the scheme and enhancing the benefit potential of the clinical trial; by constraining the distribution condition, the baseline characteristics of the initial simulated cohort are consistent with the real-world reference population, ensuring the extrapolation of the clinical trial results and their clinical application value; and by using a closed-loop optimization mechanism, the optimization process of the inclusion and exclusion conditions is based on objective data rather than human experience judgment, significantly improving the reliability and stability of the scheme, thereby improving the design quality and success rate of the clinical trial.
[0081] In one specific implementation, before returning to the execution step S20, S50 further includes the following step: saving the currently adjusted set of inbound and outbound conditions, the corresponding baseline feature distribution, the corresponding initial simulation queue, and the corresponding list of key feature variables as a record of one optimization iteration.
[0082] The interactive interface displays records of at least two optimization iterations, where the baseline feature distribution, initial simulation queue, and list of key feature variables are displayed in the form of distribution histograms, pie charts, or box plots.
[0083] Each round of adjustment of the inclusion and exclusion conditions corresponds to a unique data chain. Saving this chain enables full traceability of the iteration process, making it convenient for researchers to trace back the adjustment logic and verify the adjustment effect. At the same time, it provides a basis for multi-round data comparison and avoids iteration interruption or repeated experiments due to data loss.
[0084] Interactive interfaces are visual interfaces that allow researchers to interact with the optimization system, such as desktop applications and web pages. They support functions such as data viewing, parameter setting, retrieval of iteration records, and chart interaction. Compared to plain text data, visual charts can more intuitively present the differences between multiple iterations, such as changes in queue size, differences in baseline distribution balance, and differences in key feature variables. This reduces the data interpretation costs for researchers, helps them quickly identify the optimal adjustment direction, and improves decision-making efficiency.
[0085] The above-mentioned approach generates a structured initial inclusion / exclusion condition set by parsing the target text, constructs an initial simulated cohort by combining database retrieval, and analyzes the baseline feature distribution. This allows for the acquisition of a representative research population without the need for real-world subject recruitment, significantly reducing the time and manpower costs in the early stages of clinical trials. By integrating multi-dimensional data such as clinical phenotypes, genomics, and proteomics to construct a panoramic object feature matrix, the approach overcomes the limitations of traditional single baseline features, achieving comprehensive coverage of object features and providing a comprehensive and systematic data foundation for key factor mining. By analyzing the contribution of each feature variable in the panoramic object feature matrix to the preset key outcomes, the approach accurately identifies key feature variables strongly correlated with efficacy, survival, and adverse reactions, shifting the inclusion / exclusion condition optimization from experience-driven to data-driven, thus improving the scientific rigor and relevance of inclusion / exclusion condition design. Through comprehensive judgment and dynamic adjustment of the initial simulated cohort, baseline feature distribution, and key feature variable list, a closed-loop optimization process is formed, ensuring that the target inclusion / exclusion condition set balances cohort representativeness, feature relevance, and distribution rationality. This provides an efficient and feasible technical solution for clinical trial design in the context of precision medicine, thereby improving the success rate and research value of clinical trials.
[0086] Example 2 This example 2 provides a method for optimizing the inclusion and exclusion conditions based on experimental simulation. The method for optimizing the inclusion and exclusion conditions based on experimental simulation includes the following steps: S1, parse the target text to generate an initial set of inclusion and exclusion conditions.
[0087] S2, based on the initial set of inclusion and exclusion conditions, retrieve matching object data from at least one database to obtain an initial simulation queue and baseline feature distribution, wherein the baseline feature distribution includes at least two of the following: age distribution, gender distribution, disease stage distribution, and biomarker expression distribution.
[0088] The specific implementation methods of S1 and S2 can be referred to the specific implementation methods of S10 and S20 in Embodiment 1.
[0089] S3. Based on the baseline feature distribution, the propensity score matching algorithm is used to perform nearest neighbor matching on the experimental group and control group obtained from the initial simulation queue division to obtain a baseline balanced simulation experimental queue.
[0090] The core of the Propensity Score Matching (PSM) algorithm is to eliminate confounding factors. It constructs a comprehensive propensity score index to quantify the probability of an individual being assigned to the experimental group. Based on this score, the experimental group is matched with the control group, ensuring that the two groups have no significant differences in the distribution of key covariates. This simulates the baseline balance of a randomized controlled trial and improves the objectivity of subsequent effect evaluation.
[0091] Specifically, the propensity score is the probability value of an object being assigned to the experimental group, calculated by a logistic regression model based on the object's key covariates. The value ranges from [0, 1]. Essentially, it compresses multiple key covariates into a comprehensive index, achieving dimensionality reduction from multi-dimensional covariates to single-dimensional scores, thereby simplifying the matching operation.
[0092] Among them, key covariates are baseline characteristics (such as age, disease stage, and biomarker expression) that are selected from the baseline characteristic distribution and are directly related to the pre-specified key outcomes (effect, survival, and adverse reactions) of the clinical trial, and may have different distributions between the experimental and control groups. These variables are potential confounding factors that lead to differences in effects between groups and serve as the core basis for propensity score calculation and matching.
[0093] Nearest neighbor matching is used to find the person in the control group with the closest propensity score for each subject in the experimental group. 1:1 matching is the most commonly used matching ratio, which can maximize the sample utilization rate.
[0094] Baseline balance means that there is no statistically significant difference in the distribution of key covariates between the matched experimental group and the control group. It is a prerequisite for ensuring the reliability of the effect evaluation results. In other words, only when the baseline is balanced can it be considered that the difference in effect between the groups observed later is caused by the treatment regimen rather than the difference in baseline characteristics.
[0095] In one specific implementation, S3 includes the following steps: S31, according to the experimental grouping criteria, the objects in the initial simulation queue are divided into the first object corresponding to the initial experimental group and the second object corresponding to the initial control group.
[0096] S32, based on key covariates related to the pre-defined key outcome, calculate the propensity score for each first subject and each second subject using a logistic regression model, wherein the pre-defined key outcome includes at least one of effect response, survival, and adverse reaction, and the key covariates include at least two of age, sex, disease stage, and biomarker expression.
[0097] S33, Based on the propensity score, the propensity score matching algorithm is used to perform a 1:1 nearest neighbor matching on the first object and the second object to determine the matching first object and the second object. The matching parameters of the propensity score matching algorithm include the matching caliper value.
[0098] S34, a reference experimental group is formed based on all matching first subjects, and a reference control group is formed based on all matching second subjects.
[0099] S35. Based on the distribution data of the first subject in the reference experimental group and the second subject in the reference control group for each key covariate, determine the inter-group comparison P value corresponding to each key covariate through statistical tests.
[0100] S36. If the inter-group comparison P-values for all key covariates are greater than or equal to the first preset threshold, then the current reference experimental group is determined as the target experimental group, the current reference control group is determined as the target control group, and a baseline-balanced simulated experimental cohort is obtained by combining the target experimental group and the target control group. Otherwise, the matching parameters of the propensity score matching algorithm are adjusted, and the execution step S33 is returned using the adjusted matching parameters.
[0101] Based on the trial grouping criteria, and according to information such as treatment regimens and molecular characteristics, all subjects in the initial simulation cohort were divided into a group receiving the trial intervention (first subject) and a group receiving the control intervention (second subject), providing a basis for subsequent matching.
[0102] Logistic regression models can predict the probability of a binary dependent variable based on multiple key covariates. For example, they can predict whether an individual belongs to the experimental group; a value of 1 is assigned to the initial experimental group, and a value of 0 is assigned to the initial control group. Based on this, the key covariates are integrated into a single propensity score, achieving dimensionality reduction from multi-dimensional covariates. Then, by substituting the key covariates of each individual into the trained logistic regression model, the probability value of being assigned to the experimental group, i.e., the propensity score, is output.
[0103] The matching clamp value is the maximum allowed difference in propensity scores between the experimental and control groups, typically ranging from 0.02 to 0.2 times the standard deviation of the propensity scores. Using propensity scores as the matching criterion, the system finds the second subject in the control group whose score is closest to the first subject in the experimental group. The clamp value limits the score difference to ensure that the matched subjects have high similarity in baseline characteristics. Specifically, the matching ratio is first determined to be 1:1, meaning one first subject in the experimental group is matched with one second subject in the control group, and an initial matching clamp value is set, such as 0.2 times the standard deviation of the propensity scores. Then, for each first subject in the initial experimental group, the absolute difference between their propensity score and the propensity scores of all second subjects in the initial control group is calculated. The second subject with the smallest difference that is less than or equal to the matching clamp value is selected and paired with that first subject. Paired second subjects do not participate in subsequent matching to avoid duplicate matching. Finally, all successfully matched pairs are output, and unmatched pairs are discarded.
[0104] Statistical tests were used to determine whether there were significant differences in the distribution of key covariates between the matched experimental and control groups, in order to quantify the baseline balancing effect. Specifically, the distribution of each key covariate in the experimental and control groups was statistically analyzed. For continuous variables, the mean and standard deviation were calculated; for categorical variables, the frequency and percentage were calculated. Then, statistical tests were selected: for continuous covariates, the independent samples t-test (if conforming to a normal distribution) or the Wilcoxon rank-sum test (if not conforming to a normal distribution) was used; for categorical covariates, the chi-square test (if the sample size is sufficient) or Fisher's exact test (if the frequency of a certain category is less than 5) was used. For each key covariate, the p-value was calculated to reflect the probability that there was no difference in the distribution of covariates between the two groups.
[0105] By comparing the inter-group p-values corresponding to all key covariates with the first preset threshold, it is determined whether the baseline has reached the balance requirement. If it has not, the matching parameters are adjusted and rematched until the balance condition is met, resulting in a baseline-balanced simulation test queue to ensure the reliability of the final simulation test queue. The specific value of the first preset threshold can be set by the implementer according to the actual situation. For example, in this embodiment, the first preset threshold is set to 0.05 according to industry-standard settings.
[0106] It should be noted that adjusting the matching parameters of the propensity score matching algorithm refers to reducing the matching caliper value by a preset step size. The specific value of the preset step size can be set by the implementer according to the actual situation. For example, in this embodiment, the preset step size is set to 0.02.
[0107] It should be noted that when the value of the matching caliper is adjusted to the lower limit of the preset range, that is, 0.02 times the propensity score standard deviation, the iterative adjustment is stopped. The baseline balanced simulation test queue is obtained by combining the reference experimental group and the reference control group at the time of stopping, so as to avoid excessive loss of sample size due to infinite iteration and ensure the reliability and statistical power of the simulation test queue.
[0108] As described above, by integrating multiple key covariates through propensity score matching algorithms, the experimental and control groups are highly similar in baseline characteristics, effectively eliminating confounding factors and improving the objectivity and reliability of subsequent effect simulation results. Through a closed-loop process of matching, testing, and adjustment, unbalanced cohorts can be re-matched through parameter optimization, ensuring that the final output simulated trial cohort meets the clinical trial design requirements and guaranteeing the baseline balance quality of the simulated trial cohort. By simulating the baseline balance of randomized controlled trials, the potential effects of treatment regimens can be assessed without conducting actual subject recruitment and randomization, significantly reducing the upfront costs and time risks of clinical trials.
[0109] In one specific implementation, key covariates are obtained through the following steps: combining clinical treatment guidelines for the target disease, conclusions of historical clinical trials, and the experience of domain experts, pre-defined features that are not clearly associated with the pre-defined key outcome are eliminated from all baseline features, resulting in several reference features.
[0110] Statistical methods were used to examine the association between each reference feature and the pre-specified key outcome.
[0111] Based on a preset correlation threshold, several key covariates are selected from all reference features.
[0112] For example, in clinical trials for non-small cell lung cancer, the clinical consensus holds that EGFR mutation status and PD-L1 expression level are strongly correlated with the efficacy of immunotherapy and can be directly included as reference features; while the subject's occupation, dietary habits, etc. are not related to the efficacy and are directly excluded.
[0113] For continuous reference features, if the corresponding preset critical outcome type is a binary outcome (e.g., effective / ineffective), Pearson correlation analysis can be used to obtain the association between the reference feature and the preset critical outcome, and a threshold for the absolute value of the correlation coefficient can be set as the preset association threshold. For categorical reference features, if the corresponding preset critical outcome type is a binary outcome (e.g., effective / ineffective), chi-square test / Fisher exact test can be used to obtain the association between the reference feature and the preset critical outcome, and a p-value threshold can be set as the preset association threshold. For continuous / categorical reference features, if the corresponding preset critical outcome type is a survival outcome (e.g., survival time), univariate Cox regression analysis can be used to obtain the association between the reference feature and the preset critical outcome, and a hazard ratio (HR) threshold and a p-value threshold can be set as preset association thresholds. Specific threshold values can be set by the implementer based on the actual situation.
[0114] The final key covariates should include at least two of the following: age, sex, disease stage, and biomarker expression, to provide core input for subsequent propensity score calculation.
[0115] It should be noted that multicollinearity can be further examined on the reference features selected through a preset correlation threshold, and redundant features can be removed. Specifically, the variance inflation factor (VIF) among these features is calculated. If VIF < 10, it indicates no significant multicollinearity, and the corresponding reference feature is identified as a key covariate. If VIF ≥ 10, it indicates strong multicollinearity, and one of the reference features is randomly removed, with the remaining reference features identified as key covariates.
[0116] S4. Perform effect simulation calculations on the simulated trial cohort according to the preset prediction model to obtain the simulation results characterizing the prediction effect. The simulation results include at least the hazard ratio and the corresponding 95% confidence interval, the Kaplan-Meier survival curve, and the Log-rank test p-value.
[0117] The preset prediction model is a survival analysis model pre-trained based on historical clinical trial subject data. In this embodiment, it is a Cox proportional hazards model or a random survival forest model, used to predict the risk of endpoint events (such as disease progression or death) occurring during the simulated follow-up period. Those skilled in the art will know that the structure and training methods of the Cox proportional hazards model and random survival forest model in the prior art fall within the protection scope of this invention, and will not be described in detail here.
[0118] The hazard ratio (HR) is used to quantify the difference in the risk of an event between the experimental group and the control group. An HR < 1 indicates that the risk in the experimental group is lower than that in the control group, meaning the experimental protocol is superior. An HR > 1 indicates the opposite, and an HR = 1 indicates that there is no difference in risk between the two groups. The 95% confidence interval reflects the statistical reliability of the hazard ratio. If the interval does not include 1, it means that the difference in risk between the groups is statistically significant.
[0119] The Kaplan-Meier product limit estimation method is a commonly used nonparametric method in survival analysis. It is used to estimate the trend of survival probability changes over time for different groups of subjects, and finally generate an intuitive survival curve.
[0120] The Log-rank test is a statistical test used to compare the differences in survival curves between two or more groups. It determines whether the difference in survival between groups is statistically significant by calculating the p-value. A p-value < 0.05 is generally considered to be significant.
[0121] Therefore, this embodiment relies on a preset prediction model, inputs baseline balanced simulated trial cohort data, and predicts the effect difference between the experimental group and the control group through quantitative modeling and statistical analysis; combined with multi-dimensional indicators such as hazard ratio, survival curve, and Log-rank test, a complete effect evaluation system is formed to achieve prospective simulation of clinical trial outcomes.
[0122] In one specific implementation, S4 includes the following steps: S41, inputting the baseline characteristic data of each object in the simulated trial cohort into a preset prediction model to obtain the risk prediction value of each object during the simulated follow-up period, wherein the preset prediction model is a Cox proportional hazards model or a random survival forest model pre-trained based on historical clinical trial object data.
[0123] S42, based on all the predicted risk values, calculate the risk ratio of the target experimental group relative to the target control group and the 95% confidence interval of the risk ratio.
[0124] S43. Based on all the risk prediction values, a structured dataset containing the object simulation observation time and event state is generated through statistical simulation methods.
[0125] S44. The Kaplan-Meier product limit estimation method is applied to process the structured dataset, and the Kaplan-Meier survival curves of the target experimental group and the target control group are plotted.
[0126] S45. Perform a Log-rank test on the structured dataset to obtain the Log-rank test p-value.
[0127] The pre-set prediction model has learned the correlation between baseline characteristics and event risk in historical clinical trials. Therefore, by inputting the baseline characteristic data of each subject in the simulated trial cohort, the model can output the risk prediction value of each subject during the simulated follow-up period (e.g., 12 months, 36 months), which represents the probability value of the endpoint event occurring at a certain follow-up time point during the simulated follow-up period and reflects the effect-related risk level at the individual level.
[0128] All predicted risk values for the target experimental group and the target control group were extracted, and the mean predicted risk values for both groups were calculated. Using the mean predicted risk value of the target control group as a benchmark, the ratio of the mean predicted risk value of the target experimental group to the mean predicted risk value of the control group was calculated to obtain the hazard ratio. Statistical methods (such as the normal approximation method based on logarithmic transformation) were then used to estimate the 95% confidence interval of the hazard ratio to clarify the statistical fluctuation range of the hazard ratio.
[0129] Statistical simulation is a method that simulates the observation time (from enrollment to the occurrence of the endpoint event or the end of follow-up) and event status (1 for the occurrence of the endpoint event, 0 for the absence of the endpoint event, i.e., censored data) of subjects during the follow-up period based on individual risk prediction results. Based on all risk prediction values, a structured dataset containing simulated observation time and event status of subjects is generated through statistical simulation for subsequent survival curve plotting and testing.
[0130] The Kaplan-Meier product limit estimation method calculates survival probabilities at each time point to characterize the survival trend of the target population, and can intuitively show the survival differences between different groups during the follow-up period. The Log-rank test compares the difference between the actual number of events and the expected number of events at all follow-up time points between the two groups to determine whether the difference in the survival curves of the two groups is caused by random factors. Those skilled in the art will know that the Kaplan-Meier product limit estimation method and the Log-rank test method and their application methods in the prior art fall within the protection scope of this invention, and will not be described in detail here.
[0131] The above-mentioned methods improve the accuracy and reliability of prediction results by using Cox proportional hazards models or random survival forest models pre-trained on historical data; quantitatively characterize the difference in effects between the experimental and control groups by calculating hazard ratios and 95% confidence intervals; and construct a complete data chain required for survival analysis by generating a time-event structured dataset through statistical simulation, providing a data foundation for Kaplan-Meier curve plotting and Log-rank tests. This ensures that the effect simulation results are supported by both quantitative data and statistical verification, thereby enhancing the scientific rigor of the experimental simulation conclusions.
[0132] In one specific implementation, S43 includes the following steps: S431, based on the risk prediction values corresponding to all objects, using the inverse transformation sampling method, to generate simulated event occurrence times for each object.
[0133] S432, generate simulated censoring time according to the preset censoring distribution.
[0134] S433, Based on the simulated event occurrence time and simulated censoring time, determine the simulated observation time and event status flag for each object, and construct a structured dataset for survival analysis.
[0135] The inverse transformation sampling method is a random sampling method based on probability distribution. Its core logic is to use the inverse function of the cumulative distribution function to transform uniformly distributed random numbers into sampled values that conform to the target distribution. In this embodiment, the target distribution is the event occurrence time distribution of the object. This distribution is determined by the risk prediction value. The higher the risk of the object, the shorter the simulated event occurrence time; the lower the risk of the object, the longer the simulated event occurrence time.
[0136] Specifically, key parameters such as the simulated follow-up period and censoring rate were first set to ensure that the simulated scenario closely resembled real clinical trials. An inverse transformation sampling method was then used to determine the event occurrence risk function h for each subject based on the predicted risk value. i (t), for example, the event occurrence risk function of the Cox proportional hazards model is h. i (t)=h0(t)×e^(X i ×β), where i is the index identifier of the object, used to distinguish different objects in the simulation test queue, i=1,2,...,Q, and Q is the total number of objects. X i X is a vector composed of the baseline features of the i-th object. For example, if "age, gender, disease stage, and PD-L1 expression" are selected as baseline features, then X i =[age i ,gender i Disease staging i PD-L1 expression ih0(t) is a preset baseline risk function, representing the risk level when all baseline features X are equal. i When the value is 0, the risk of the event occurring at time t is the baseline risk in the Cox proportional hazards model, which can be obtained by fitting historical clinical trial data. β is the regression coefficient vector of the Cox proportional hazards model, used to quantify the degree and direction of the influence of different baseline characteristics on the risk of event occurrence. It can be the optimal solution obtained by training the Cox proportional hazards model using historical clinical trial data. Specifically, the dimension of β is related to X. i Completely consistent, each β component corresponds to the influence weight of a baseline feature, and if the β component corresponding to a certain baseline feature is... j >0 indicates that this baseline characteristic increases the risk of events occurring in the object; if β j <0 indicates that the baseline feature reduces the risk of events occurring in the object; if β j =0 indicates that the baseline feature has no significant impact on the risk of the event, where j=1,2,...,K, and K is the total number of categories of the baseline feature.
[0137] Then, by integrating the event occurrence risk function, we obtain the cumulative risk function H. i (t)=∫0 t h i (u)du, then through formula S i (t)=e^(-H i (t) calculate the survival function, where S i (t) represents the probability that the event has not occurred for the object at time t. Then, a random number u following a uniform distribution in the interval (0, 1) is generated for each object. i Let the survival function S i (t) equals the random number u i The time t of the simulated event is obtained by solving the inverse function. event,i The corresponding formula is: t event,i =S i -1 (u i ), where S i -1 For the survival function S i The inverse function of (t). Finally, the simulated event time t for each object is output. event,i This creates a table that maps "object ID to simulated event time".
[0138] In real clinical trials, some subjects may not be able to observe the occurrence of the endpoint event due to reasons such as loss to follow-up, follow-up cutoff, or withdrawal from the trial. This type of data is called censored data. In this embodiment, the censoring time is simulated by preset censoring distribution (such as exponential distribution or uniform distribution) to restore the data characteristics of real clinical trials and avoid bias in survival analysis results caused by ignoring censoring.
[0139] Specifically, based on the design of the target clinical trial, a suitable censoring distribution is selected and parameters are determined. For example, if it is a trial with a fixed follow-up period, the censoring time follows a uniform distribution U(0, T). max ), where T max The maximum follow-up time is preset; for random censoring trials, the censoring time follows an exponential distribution exp(λ), where λ is a censoring rate parameter set based on historical data. Based on the selected censoring distribution, a simulated censoring time t is generated for each subject. cens,i It also outputs a "Object ID - Simulated Censorship Time" lookup table to ensure that the distribution characteristics of censorship time are consistent with those of real clinical trials.
[0140] By comparing the simulated event occurrence time with the simulated censoring time for each subject, the final simulated observation time and event status flags are determined, thus constructing a structured dataset for survival analysis. The simulated observation time is the smaller of the simulated event occurrence time and the simulated censoring time. The event status flags are used to identify the subject's follow-up outcome: if the simulated event occurrence time is less than or equal to the simulated censoring time, "1" indicates the endpoint event has occurred; if the simulated event occurrence time is greater than the simulated censoring time, "0" indicates censoring (the endpoint event did not occur). Finally, the subject ID, group (experimental group / control group), simulated observation time, and event status flags are integrated to form a structured dataset for subsequent survival curve plotting and validation.
[0141] The above-mentioned method uses inverse transformation sampling to generate simulated event occurrence times based on individual risk prediction values, making the distribution of event occurrence times highly correlated with the actual risk level of the subjects. By introducing a pre-set censoring distribution to simulate censoring time, the generated dataset can restore the censoring characteristics of real clinical trials, avoiding bias in survival analysis results caused by ignoring censored data. Furthermore, by determining the simulated observation time and event status markers through clear time comparison rules, the realism and reliability of the effect simulation are improved.
[0142] S5. If the initial simulation queue, experimental simulation results, and baseline feature distribution meet the preset optimization conditions, then the current initial set of admission conditions is determined as the target set of admission conditions. Otherwise, the initial set of admission conditions is adjusted according to the initial simulation queue, experimental simulation results, and baseline feature distribution, and the adjusted set of admission conditions is used to replace the initial set of admission conditions. Then, return to step S2.
[0143] Among them, based on multi-dimensional quantitative judgment criteria, the initial simulation queue, experimental simulation results and baseline feature distribution generated under the current set of inclusion and exclusion conditions are comprehensively evaluated. If the preset optimization conditions are met, the set of inclusion and exclusion conditions is determined as the final solution; if not, the inclusion and exclusion conditions are adjusted in reverse based on the evaluation results. Through iterative retrieval and simulation, the inclusion and exclusion conditions are accurately optimized to ensure that the final solution takes into account both experimental feasibility and significant effect.
[0144] In one specific implementation, the preset optimization conditions include queue size conditions, baseline balance conditions, and effect prediction conditions.
[0145] The queue size condition is: the total number of objects in the current initial simulation queue is greater than or equal to the second preset threshold.
[0146] The baseline balance condition is: the current inter-group comparison P-values are all greater than or equal to the third preset threshold.
[0147] The predictive conditions are as follows: the current hazard ratio is less than the fourth preset threshold, and the corresponding 95% confidence interval does not contain the preset value, and the current Log-rank test p-value is less than the fifth preset threshold, and after the preset time point, the difference between the cumulative survival probabilities of the target experimental group and the target control group corresponding to the current Kaplan-Meier survival curve is greater than or equal to the sixth preset threshold.
[0148] The second preset threshold can be set according to the incidence rate of the target disease and the statistical requirements of the clinical trial design. For example, it is usually set to 200 for tumor clinical trials and can be appropriately reduced to 50 for rare disease trials.
[0149] The third preset threshold is the statistically accepted standard of 0.05; if the p-value for intergroup comparisons of all key covariates is ≥0.05, it indicates that there is no statistically significant difference between groups and the baseline has been balanced.
[0150] The fourth preset threshold is usually set to 1.0; a risk ratio < 1.0 indicates that the experimental group has a lower risk, and the smaller the value, the more significant the effect advantage.
[0151] The preset value is 1.0; if both the upper and lower limits of the 95% confidence interval are <1.0, it indicates that the conclusion that the hazard ratio is <1.0 is statistically significant.
[0152] The fifth preset threshold is usually set to 0.05; a p-value < 0.05 indicates that the difference in survival trends between the two groups is not caused by random factors.
[0153] The preset time point can be set to 1 year, 2 years or 3 years; the sixth preset threshold can be set to 10%, 15% or other values according to clinical needs.
[0154] When all three criteria—cohort size, baseline balance, and effect prediction—are met, it indicates that the population selected by the current inclusion / exclusion criteria is both statistically feasible and demonstrates the effectiveness of the trial protocol. Therefore, the current initial inclusion / exclusion criteria set is directly determined as the target inclusion / exclusion criteria set, serving as the final clinical trial inclusion / exclusion protocol output. Simultaneously, the corresponding simulated trial cohort, baseline characteristic distribution, and effect simulation results are archived as supporting evidence for the protocol.
[0155] For specific conditions that were not met, the direction for adjusting the inclusion and exclusion criteria was deduced in reverse to ensure that the adjusted criteria could specifically address the current problems. Specifically, if the cohort size was insufficient, some non-core restrictions in the inclusion and exclusion criteria were relaxed (e.g., the age range was expanded from 18-65 years to 18-70 years) to broaden the search scope. If the baseline was unbalanced, the scope of key covariate screening was optimized (e.g., adding biomarkers strongly correlated with the effect) to improve the targeting of the match. If the effect prediction did not meet the target, the inclusion and exclusion criteria were tightened, and high-benefit populations were screened (e.g., only those with positive biomarkers were included) to strengthen the difference in effect between groups.
[0156] The above-mentioned multi-dimensional quantitative judgment criteria ensure that the optimization of inclusion and exclusion conditions no longer relies on human experience, but is based on objective statistical data and clinical indicators, thereby improving the scientific rigor and reliability of the protocol. By setting cohort size conditions, the final inclusion and exclusion protocol guarantees the statistical power of the clinical trial, avoiding trial failure due to insufficient sample size. The verification of baseline balancing conditions ensures that the difference in efficacy between the experimental and control groups is attributed to the treatment protocol, rather than confounding baseline characteristics, ensuring the accuracy of efficacy evaluation results. The multi-sub-criteria constraint of efficacy prediction conditions ensures that the final protocol possesses both statistical significance and clinical practical value, guaranteeing the potential benefit of the clinical trial. The closed-loop iterative adjustment mechanism allows for targeted optimization of inclusion and exclusion conditions, significantly reducing the design risks and costs of clinical trials and improving research and development efficiency.
[0157] In one specific implementation, before returning to the execution step S2, S5 further includes the following step: saving the currently adjusted set of alpha and alpha conditions, the corresponding baseline feature distribution, the corresponding initial simulation queue, and the corresponding experimental simulation results as a record of one optimization iteration.
[0158] The interactive interface displays records of at least two optimization iterations, where the baseline feature distribution, initial simulation queue, and experimental simulation results are displayed in the form of distribution histograms, pie charts, or box plots.
[0159] Each round of adjustment of the inclusion and exclusion conditions corresponds to a unique data chain. Saving this chain enables full traceability of the iteration process, making it convenient for researchers to trace back the adjustment logic and verify the adjustment effect. At the same time, it provides a basis for multi-round data comparison and avoids iteration interruption or repeated experiments due to data loss.
[0160] The interactive interface is a visual interface for researchers to interact with the optimization system, such as a desktop application interface or a web interface. It supports functions such as data viewing, parameter setting, retrieval of iteration records, and chart interaction. Compared with plain text data, visual charts can more intuitively present the differences between multiple iterations, such as changes in cohort size, differences in baseline distribution balance, and differences in effect prediction trends. This reduces the data interpretation cost for researchers, helps them quickly identify the optimal adjustment direction, and improves decision-making efficiency.
[0161] As described above, by parsing the target text of natural language to generate a structured initial set of inclusion and exclusion conditions, and then retrieving matching object data from the database to construct an initial simulation cohort and analyze the baseline feature distribution, a representative study population can be obtained without conducting real-subject recruitment in the early stages of clinical trials, reducing the time and manpower costs of trial design. Based on the baseline feature distribution and using a propensity score matching algorithm to achieve baseline balance between the experimental and control groups, the confounding interference of key covariates is effectively eliminated, allowing the differences in subsequent effect simulation results to be accurately attributed to the treatment plan itself, improving the objectivity and reliability of effect assessment. By outputting multi-dimensional effect simulation results of hazard ratio, 95% confidence interval, Kaplan-Meier survival curve, and Log-rank test p-value through a pre-set prediction model, the potential benefit of the trial under the current inclusion and exclusion conditions can be comprehensively and intuitively predicted. Thus, through a closed-loop optimization process of simulation, evaluation, adjustment, and iteration, the inclusion and exclusion conditions are dynamically adjusted based on a comprehensive judgment of the initial simulation cohort, trial simulation results, and baseline feature distribution, so that the final set of target inclusion and exclusion conditions can take into account cohort size, baseline balance, and effect significance, improving the scientific design and success rate of clinical trials.
[0162] Example 3 This example 3 provides a method for joint optimization of inclusion and exclusion conditions, as shown in Figure 3. The method for joint optimization of inclusion and exclusion conditions includes the following steps: S100, parsing the target text to generate an initial set of inclusion and exclusion conditions.
[0163] S200: Based on the initial set of sorting and grading conditions, retrieve matching object data from at least one database to obtain the initial simulation queue and baseline feature distribution.
[0164] S300, based on the baseline feature distribution, uses the propensity score matching algorithm to perform nearest neighbor matching on the experimental group and control group obtained from the initial simulation queue division, and obtains a baseline balanced simulation experimental queue.
[0165] S400: Based on the preset prediction model, the simulation calculation of the effect of the simulation test queue is performed to obtain the test simulation results that characterize the prediction effect.
[0166] S500 performs pre-processing on the object data related to the target domain corresponding to the initial simulation queue to obtain the panoramic object feature matrix.
[0167] S600 analyzes the feature matrix of the panoramic object to obtain the contribution of each feature variable to the prediction of the preset key outcome.
[0168] S700: If the initial simulation queue, experimental simulation results, baseline feature distribution, and list of key feature variables generated according to contribution satisfy the joint optimization conditions, then the current initial set of admission and ranking conditions is determined as the target set of admission and ranking conditions. Otherwise, the initial set of admission and ranking conditions is adjusted according to the initial simulation queue, experimental simulation results, baseline feature distribution, and list of key feature variables, and the adjusted set of admission and ranking conditions is used to replace the initial set of admission and ranking conditions. Then, the process returns to step S300.
[0169] Among them, the core of the inclusion-exclusion condition optimization method corresponding to Example 1 lies in quantitatively verifying the difference in effect between the experimental group and the control group under inclusion-exclusion conditions through propensity score matching and effect simulation calculation, thus solving the problem that the potential benefit of the experiment cannot be predicted after the inclusion of key features. On the other hand, the core of the inclusion-exclusion condition optimization method corresponding to Example 2 lies in accurately locating the key feature variables that affect the preset key outcome through multi-omics data mining and feature contribution analysis, thus solving the problem that inclusion-exclusion conditions are driven by experience and lack molecular-level support.
[0170] Therefore, in order to further improve the scientific nature and pertinence of the inclusion and exclusion conditions, this embodiment combines the inclusion and exclusion condition optimization methods corresponding to Embodiment 1 and Embodiment 2 above, integrates the effect simulation verification capability and the key feature mining capability, and realizes the inclusion and exclusion condition design with precise feature positioning, effect simulation verification, and iterative optimization closed loop.
[0171] The specific implementation methods for each step can be found in Embodiment 1 and Embodiment 2.
[0172] The joint optimization conditions include queue size conditions, baseline balancing conditions, effect prediction conditions, feature correlation conditions, and distribution rationality conditions. Among them, the queue size condition is: the total number of objects in the current initial simulation queue is greater than or equal to a second preset threshold.
[0173] The baseline balance condition is: the current inter-group comparison P-values are all greater than or equal to the third preset threshold.
[0174] The predictive conditions are as follows: the current hazard ratio is less than the fourth preset threshold, and the corresponding 95% confidence interval does not contain the preset value, and the current Log-rank test p-value is less than the fifth preset threshold, and after the preset time point, the difference between the cumulative survival probabilities of the target experimental group and the target control group corresponding to the current Kaplan-Meier survival curve is greater than or equal to the sixth preset threshold.
[0175] The feature association condition is: the proportion of the number of feature variables in the list of key feature variables that rank in the top N in terms of contribution and are included in the current set of inclusion and exclusion conditions is greater than or equal to the third preset threshold, where N is a positive integer.
[0176] The reasonable distribution condition is that, in at least one preset key dimension, the statistical distribution difference between the baseline feature distribution and the reference population distribution in the target field is less than the fourth preset threshold.
[0177] If the joint optimization conditions are met, the candidate list of admission and ranking conditions is determined as the target list of admission and ranking conditions; if not, the admission and ranking conditions are adjusted according to the initial simulation queue, experimental simulation results, baseline feature distribution and key feature variable list, and the process returns to step S300 to re-execute the queue construction and feature mining process until all joint optimization conditions are met.
[0178] The above-mentioned approach, by integrating the inclusion and exclusion condition optimization methods based on trial simulation and the inclusion and exclusion condition optimization methods based on key factor mining, not only solves the limitations of designing inclusion and exclusion conditions based on experience, but also avoids the problem of not being able to predict the trial effect after the inclusion of key features, thus achieving precise design of inclusion and exclusion conditions. Through the dual filtering of feature mining and effect verification, the selected set of target inclusion and exclusion conditions takes into account both the relevance of key feature variables and the significance of trial benefits, effectively improving the effect discrimination and success rate of clinical trials.
[0179] Example 4 of the present invention provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the key factor mining-based optimization method for the above embodiments.
[0180] Embodiment 5 of the present invention provides an electronic device, which includes a processor and a non-transitory computer-readable storage medium as described in Embodiment 4 of the present invention.
[0181] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for optimizing inclusion and exclusion conditions based on key factor mining, characterized in that, The method includes the following steps: S10, parsing the target text to generate an initial set of inclusion and exclusion criteria, wherein the target text includes trial grouping criteria, inclusion criteria, and exclusion criteria; S20, retrieving matching object data from at least one database based on the initial set of inclusion and exclusion criteria to obtain an initial simulation cohort and baseline feature distribution, wherein the baseline feature distribution includes at least two of age distribution, gender distribution, disease stage distribution, and biomarker expression distribution; S30, performing pre-processing on object data related to the target domain corresponding to the initial simulation cohort to obtain a panoramic object feature matrix; S40, processing the panoramic object feature matrix... The matrix is analyzed to obtain the contribution of each feature variable to the prediction of the preset key outcome, wherein the preset key outcome includes at least one of effect response, survival time, and adverse reaction; S50, if the initial simulation queue, the baseline feature distribution, and the list of key feature variables generated according to the contribution are satisfied with the preset optimization conditions, then the current initial inclusion and exclusion condition set is determined as the target inclusion and exclusion condition set; otherwise, the initial inclusion and exclusion condition set is adjusted according to the initial simulation queue, the baseline feature distribution, and the list of key feature variables, and the adjusted inclusion and exclusion condition set is used to replace the initial inclusion and exclusion condition set, and the process returns to step S20.
2. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 1, characterized in that, S10 includes the following steps: S101, receiving target text input by the user via file upload or manual input, wherein the target text is in natural language form; S102, parsing the target text using a preset large language model to generate an initial arrangement condition set including several initial arrangement conditions, wherein the preset large language model is a large language model fine-tuned with medical domain knowledge and optimized with professional prompts, each initial arrangement condition includes a logical node and an entity node, the logical node is used to define the AND / OR logical relationship between conditions, and the entity node includes entity type, attribute name, operator and corresponding value.
3. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 1, characterized in that, S20 includes the following steps: S201, converting the initial set of inclusion and exclusion conditions into a query statement compatible with at least one selected database; S202, executing the query and performing de-privacy operations according to the query statement, and integrating to obtain an initial simulation queue; S203, extracting the baseline feature distribution corresponding to the initial simulation queue.
4. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 1, characterized in that, The object data includes at least two of the following: clinical phenotype data, genomics data, proteomics data, and single-cell omics data corresponding to each object. S30 includes the following steps: S301, for any object corresponding to the object data, preprocessing is performed on the object data of the current object for each preset type to obtain intermediate data of the current object for each preset type. The preprocessing includes batch effect correction, normalization, and quality filtering. The preset types include at least two of clinical phenotype, genomics, proteomics, and single-cell omics. S302, feature extraction is performed on the intermediate data of the current object for each preset type to obtain several feature variable data corresponding to the current object. The clinical phenotype-related feature variables include at least three of the following: age, sex, disease stage, previous treatment history, ECOG score, and laboratory test results. The genomics-related feature variables include at least two of the following: gene mutation status, tumor mutation burden, gene fusion, copy number variation, and pathway enrichment score. The proteomics-related feature variables include at least one of differentially expressed proteins and phosphorylation modification levels. The single-cell omics-related feature variables include at least one of the following: cell subpopulation ratio and differentially expressed gene module score. S303, using objects as rows and feature variables as columns, the panoramic object feature matrix is constructed based on several feature variable data corresponding to each object.
5. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 4, characterized in that, S40 includes the following steps: S401, using the preset key outcome as the prediction target, the preset interpretable machine learning model is trained using the panoramic object feature matrix to obtain the trained interpretable machine learning model. S402, based on the trained interpretable machine learning model, attribution analysis is performed on each feature variable in the panoramic object feature matrix using a model interpretability framework to obtain the contribution of each feature variable to predicting the preset key outcome.
6. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 5, characterized in that, S40 further includes the following steps: S403, sorting all feature variables in descending order of contribution to obtain a sorting result; S404, selecting feature variables whose ranking is lower than a preset ranking threshold from the sorting result to form the list of key feature variables.
7. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 1, characterized in that, The preset optimization conditions include queue size conditions, feature association conditions, and distribution rationality conditions; wherein, the queue size condition is: the total number of objects in the initial simulated queue is greater than or equal to a second preset threshold; the feature association condition is: the proportion of the number of feature variables in the list of key feature variables whose contribution ranks in the top N and are included in the current inclusion and exclusion condition set is greater than or equal to a third preset threshold, where N is a positive integer; the distribution rationality condition is: on at least one preset key dimension, the statistical distribution difference between the baseline feature distribution and the reference population distribution in the target field is less than a fourth preset threshold.
8. The method for optimizing inclusion and exclusion conditions based on key factor mining according to claim 1, characterized in that, Before returning to step S20, step S50 further includes the following steps: saving the currently adjusted set of inclusion and exclusion conditions, the corresponding baseline feature distribution, the corresponding initial simulation queue, and the corresponding list of key feature variables as a record of one optimization iteration; displaying the records of at least two optimization iterations on the interactive interface, wherein the baseline feature distribution, the initial simulation queue, and the list of key feature variables are displayed on the interactive interface in the form of a distribution histogram, pie chart, or box plot.
9. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the key factor mining-based optimization method for the categorization conditions as described in any one of claims 1-8.
10. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 9.