Clinical research-based intelligent entry and exit method, equipment, medium and product

Through the intelligent in-order and arrangement method based on large models, the problem of low automation in clinical research is solved, and the full-link automation derivation from research plans to patient lists is realized, supporting the unified mapping of multi-source heterogeneous data and the dynamic iteration of medical knowledge is improved, and screening efficiency and accuracy are improved.

CN120544762APending Publication Date: 2025-08-26上海和今信息科技有限公司

Patent Information

Application Number
CN202511023709.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing technology has a low degree of automation in clinical research, and the data mapping and code debugging process still require manual intervention, which cannot achieve end-to-end intelligence, and lacks the ability to generalize different research types and multi-source heterogeneous data.

Method used

Using a large model-based intelligent in-ordering method, we can realize the full-link automated derivation and dynamic update from the research plan to the patient list by determining research design elements, generating inference rules and discriminative codes, including the extraction of research design elements, the generation of inference rules and the dynamic generation of discriminative codes, supporting the unified mapping of multi-source heterogeneous data and the dynamic iteration of medical knowledge.

Benefits of technology

It realizes an end-to-end intelligent process without manual intervention, improves the system's generalization ability in different scenarios, solves the problems of rigid type adaptation and heterogeneous data integration in traditional technologies, and improves screening efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544762A_ABST
    Figure CN120544762A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of computers, and discloses an intelligent entry and exit method and device based on clinical research, a medium and a product. The method is realized based on a large model, and comprises the following steps: determining research design elements according to a received research scheme; determining an inference rule according to the research design elements; according to the inference rule, generating a discrimination code of each research design element; and according to the discrimination code, determining a patient list conforming to the research scheme. The technical problems that in the prior art, the automation degree is low, manual intervention is still needed in the links of data mapping, code debugging and the like, real end-to-end intelligence cannot be achieved, and generalization adaptability to different research types, multi-source heterogeneous data and medical knowledge updating is lacked can be at least solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an intelligent admission and discharge method, device, medium and product based on clinical research. Background Art

[0002] In clinical research, the ability to accurately screen out subjects who meet the research plan directly determines the quality and efficiency of the research. In traditional methods, subject screening requires close collaboration between clinical researchers and technical personnel. Clinical researchers have a medical background and are responsible for defining the retrieval logic for the required patient data based on medical needs; technical personnel have technical expertise and are responsible for handling technical aspects such as storage, retrieval, and analysis of medical data. Specifically, this includes writing SQL code based on the retrieval logic provided by clinical researchers to obtain the required patient data from medical databases such as electronic medical records and medical insurance data. The above-mentioned data extraction process that relies on manual collaboration is often inefficient and error-prone due to difficulties in cross-disciplinary communication, difficulties in adapting multi-source heterogeneous data, and limitations in manual rule design. It is difficult to meet the growing demand for efficient and accurate screening in digital clinical research.

[0003] To address these challenges, some technical solutions have begun to try to introduce rule engines or standardized templates to simplify the screening process, improve efficiency and reduce human errors.

[0004] However, the inventors found that these technical solutions have at least the following technical problems: the degree of automation is low, manual intervention is still required in aspects such as data mapping and code debugging, and true end-to-end intelligence cannot be achieved. In addition, there is a lack of generalized adaptability to different research types, multi-source heterogeneous data, and medical knowledge updates. Summary of the Invention

[0005] One purpose of this application is to provide an intelligent inclusion and exclusion method, device, medium and product based on clinical research, at least to solve the technical problems of low degree of automation in related technologies, manual intervention in data mapping and code debugging, inability to achieve true end-to-end intelligence, and lack of generalized adaptability to different research types, multi-source heterogeneous data, and medical knowledge updates.

[0006] To achieve the above objectives, some embodiments of the present application provide the following aspects: In a first aspect, some embodiments of the present application also provide an intelligent inclusion and exclusion method based on clinical research, which is implemented based on a large model and includes: determining research design elements based on a received research plan; determining inference rules based on the research design elements; generating a discrimination code for each research design element based on the inference rules; and determining a list of patients who meet the research plan based on the discrimination code.

[0007] In a second aspect, some embodiments of the present application further provide an electronic device comprising: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, cause the processor to perform the steps of the method described above.

[0008] In a third aspect, some embodiments of the present application further provide a computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method described above.

[0009] In a fourth aspect, some embodiments of the present application further provide a computer program product, comprising a computer program / instruction, which implements the steps of the above-described method when executed by a processor.

[0010] Compared with related technologies, the embodiments of the present application provide an intelligent, automated, accurate and efficient inclusion and exclusion solution. This solution can fundamentally solve the problem that traditional technologies cannot achieve true intelligence and lack of generalization adaptability through an end-to-end process of dynamically deriving research design elements, inference rules and discrimination codes based on research plans: specifically, the method can automatically identify the required research design elements based on the received research plan. This process can avoid the type adaptation rigidity caused by traditional methods relying on fixed rules or manual experience, and thus naturally have the ability to adapt to different research types; furthermore, the inference rules generated based on research design elements can uniformly map multi-source heterogeneous data to standardized discrimination logic, thereby solving the integration problem of heterogeneous data caused by structural differences; moreover, the dynamic generation mechanism of discrimination codes enables the system to automatically adjust research design elements and inference rules as medical knowledge is updated, which can avoid the knowledge solidification problem of traditional static rule bases. This full-link automated derivation and dynamic update mechanism, from inputting research plans to outputting patient lists, can not only realize an end-to-end intelligent process without human intervention, but also comprehensively improve the system's generalization capabilities in different scenarios through adaptive analysis of research types, unified mapping of multi-source data, and dynamic iteration of knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0012] Figure 1 An exemplary flow chart of an intelligent admission and exclusion method based on clinical research provided in some embodiments of the present application; Figure 2An exemplary schematic diagram of an intelligent admission and exclusion method based on clinical research provided in some embodiments of the present application; Figure 3 An exemplary structural diagram of an electronic device provided for some embodiments of the present application. DETAILED DESCRIPTION

[0013] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0014] The following terms are used in this document.

[0015] 1. Large models: These are AI models with a large number of parameters (typically millions to trillions) that can handle complex natural language understanding and generation tasks. In this application, large models are used to extract key study design elements from study proposal text and assist in generating inference rules.

[0016] 2. Clinical research: A scientific research method used to evaluate the safety and effectiveness of medical interventions, drugs, or treatments, including randomized controlled trials, simulated randomized controlled trials, cohort studies, cross-sectional studies, etc.

[0017] 3. Data dictionary: This is metadata that records database structure information. In this application, the data dictionary includes at least: the names of all forms in the database, the names of the variables in each form, the variable labels used to explain the meaning of each variable, the variable data types, and the corresponding encoding rules.

[0018] 4. Research design elements: These refer to the smallest units in a research design that can be implemented and used for data mapping. They are divided into categories such as inclusion criteria, exclusion criteria, intervention measures, control methods, clinical outcomes, and covariates (confounding factors). For example, in a research protocol, “Inclusion criteria: (1) Age older than 18 years; (2) Suffering from type 2 diabetes; (3) Not receiving hypoglycemic treatment…” Each of the above inclusion criteria can be considered as one or more research design elements. Generally speaking, a clinical research protocol can be decomposed into more than ten or even dozens of research design elements in a structured manner.

[0019] 5. Inference rules: In this application, it refers to the factor variable inference rules, which are presented in the form of natural language text or pseudocode, and are mainly used to explain how to infer research design elements based on the variables (variable labels) in the database. A research design element can have multiple feasible inference rules. For example, for the research design element "suffering from diabetes", when there are variables such as "disease diagnosis" and "fasting blood glucose value" in the database, the corresponding inference rule can be: "The disease diagnosis contains the keyword <diabetes>", or it can be "fasting blood glucose value is greater than 7mmol / L".

[0020] 6. Decision code: This is a piece of code that can actually be executed and can be written in SQL or other languages ​​that can perform database operations. Its function is to determine whether the information of a specific patient in the database meets the requirements of the specified study design elements based on the specified inference rules.

[0021] 7. Workflow: It is a set of logical processes containing multiple nodes (such as steps and functions) that are used to control how each node handles user input, function call logic, response decisions, and other operations.

[0022] 8. Thought Chain: This is a technology that enhances the reasoning ability of AI systems. By simulating the step-by-step thinking process of humans when solving problems, it enhances the model's understanding and solution efficiency of complex problems and scenarios.

[0023] 9. RAG Knowledge Base: Retrieval-augmented generative knowledge base is a mechanism that combines the advantages of retrieval technology and generative models. It allows AI systems to generate answers based on their own training data and obtain relevant information from external knowledge sources to support output content.

[0024] 10. Semantic Embedding: This is a technique that converts text into numerical vectors. Through this conversion, machines can understand the meaning of words and the semantic relationships between words.

[0025] 11. Multi-source heterogeneous data refers to a collection of data from different data sources with different structures and formats.

[0026] First embodiment The first embodiment of this application relates to an intelligent admission and exclusion method based on clinical research. Figure 1 and Figure 2 As shown, the method is implemented based on a large model and may include the following steps: Step S101, determining the research design elements according to the received research plan; Step S102, determining inference rules based on the research design elements; Step S103, generating a discriminant code for each research design element according to the inference rule; Step S104: Determine a list of patients who meet the research plan based on the identification code.

[0027] Exemplarily, the method can be applied to a large-scale model-based clinical research patient screening and data extraction system. Taking into account the operating mechanism of the large model and the strict requirements of clinical research on the accuracy of the results, the embodiments of this application can use a combination of workflow and thinking chain to systematically implement each step. Among them, the workflow ensures the efficiency of data processing and rule execution by orderly connecting multiple functional nodes; the thinking chain simulates the human reasoning process and assists the large model in disassembling and verifying complex medical logic layer by layer, thereby improving the screening efficiency while ensuring the medical accuracy and reliability of the results. The following is a detailed description of each of the above steps.

[0028] Regarding step S101, illustratively, the research plan can be presented in the form of an electronic text. After the system receives the electronic text of the research plan, the macro model can extract all research design elements based on the content of the electronic text of the research plan. The electronic text format can be, but is not limited to, .text. The research plan can be, but is not limited to, a clinical trial plan, an observational study design plan, a project brief, or other research plan. The research design elements are key framework items that define core content such as the target population, intervention measures, observation indicators, and influencing factors of the clinical study.

[0029] For example, the research design elements may include multiple dimensions: in terms of inclusion and exclusion criteria, the target population that meets the research conditions is defined through specific requirements such as age range, disease diagnosis conditions, and comorbidity restrictions; intervention measures / exposure factors detail the specific operations of the experimental group and the control group, such as drug dosage, treatment cycle, exposure environment, etc.; clinical outcomes include primary endpoints and secondary endpoints, which specify indicators to be observed and measured in the study, such as the degree of improvement of disease symptoms, survival time, etc.; key covariates involve important factors that may affect the research results, such as the patient's gender, BMI, smoking history, etc.

[0030] For step S102, for example, inference rules can be determined based on the study design elements. For example, for the inclusion criteria of "age ≥ 18 years and < 65 years," the inference rules must specify the age calculation method (based on the difference between the enrollment date and the birth date) and the handling of boundary values ​​(including 18 years old and excluding 65 years old). For "using a certain drug for treatment for more than 3 months," the inference rules must determine the drug name matching rules and the start and end times for calculating the treatment duration.

[0031] For step S103, the inference rules can be converted into computer-readable and executable judgment code using a programming language or specific data processing tools. For example, "age ≥ 18 years and < 65 years" can be written as an SQL query statement, R code, or Python code; "using a certain medication for treatment for more than 3 months" can be converted into code logic, which can be used to retrieve medication records in the electronic medical record, calculate the medication usage time, and make a judgment.

[0032] Regarding step S104, illustratively, based on the discriminant code, data retrieval and screening operations can be performed in a medical database (e.g., an electronic medical record system or a laboratory database). The system can compare patient data against the discriminant code conditions one by one, screening patients who meet all study design requirements to form a list of patients that meet the study plan. This list of patients can be the direct subject of subsequent clinical research and statistical analysis of the data.

[0033] It is not difficult to find that compared with the relevant technologies, the embodiments of the present application provide an intelligent, automated, accurate and efficient inclusion and exclusion solution. This solution can fundamentally solve the problem that traditional technologies cannot achieve true intelligence and insufficient generalization and adaptability through an end-to-end process of dynamically deducing research design elements, inference rules and discrimination codes based on the research plan: Specifically, the method can automatically identify the required research design elements based on the received research plan. This process can avoid the rigidity of type adaptation caused by traditional methods relying on fixed rules or manual experience, and thus naturally have the ability to adapt to different research types; furthermore, the inference rules generated based on the research design elements can uniformly map multi-source heterogeneous data to standardized discrimination logic, thereby solving the integration problem caused by structural differences in heterogeneous data; moreover, the dynamic generation mechanism of the discrimination code enables the system to automatically adjust the research design elements and inference rules as medical knowledge is updated, which can avoid the problem of knowledge solidification in the traditional static rule base. This full-link automated derivation and dynamic update mechanism, from inputting research plans to outputting patient lists, can not only realize an end-to-end intelligent process without human intervention, but also comprehensively improve the system's generalization capabilities in different scenarios through adaptive analysis of research types, unified mapping of multi-source data, and dynamic iteration of knowledge.

[0034] Second embodiment The second embodiment of the present application relates to an intelligent inclusion and exclusion method based on clinical research. The second embodiment is an improvement on the first embodiment. The specific improvement is that: in this embodiment, a specific implementation method for determining research design elements based on a received research plan is provided.

[0035] Specifically, in some embodiments, determining the research design elements according to the received research plan, i.e., step S101, may include the following steps: Step S1011, determining the element extraction prompt words according to the research plan; Step S1012: extract prompt words based on the elements, and output a structured list of research design elements through the large model.

[0036] For step S1011, exemplarily, after the system receives the electronic text of the research plan, the implementation process of determining the element extraction prompt words according to the research plan is a process of converting unstructured text into structured instructions by combining the research type characteristics and industry specifications through a large model. Specifically, the large model can be used to perform semantic analysis on the electronic text of the research plan, extract research design elements such as research objectives, inclusion and exclusion criteria, and intervention measures, and combine clinical research field knowledge to convert natural language descriptions into more specific directive statements. For example, from "adult patients with type 2 diabetes who are not using insulin treatment", detailed prompts such as "extract type 2 diabetes diagnostic criteria" and "clarify the age definition of adult patients" are disassembled to form element extraction prompt words to assist the large model in extracting elements.

[0037] Regarding step S1012, for example, the element extraction prompt can be used as input, leveraging the information extraction and structuring capabilities of the large model to extract key information (such as age range, intervention dosage, etc.) from the study protocol. This information is then organized into a structured list containing elements such as inclusion criteria, intervention measures, and clinical outcomes in a pre-set standardized format (e.g., JSON, table). This structured list, with its unified data format and presentation, eliminates ambiguity and ambiguity in the original content corresponding to the study protocol, facilitating the provision of a standardized data foundation for research. Taking a clinical study protocol for type 2 diabetes as an example, when the element extraction prompt "Extract type 2 diabetes diagnostic criteria and basis for determining non-insulin use" is input, the large model can extract key information from the study protocol and structure it: "Type 2 diabetes diagnostic criteria" is mapped to the International Classification of Diseases code "ICD-10 code E11.x," enabling direct retrieval of the diagnostic criteria using medical record codes; and "basis for determining non-insulin use" is converted to "no insulin medication record in the electronic medical order system within the past three months," facilitating screening of eligible patients from real-world data. In the structured list formed, the inclusion criteria can be clearly presented as "age ≥18 years old and ICD-10 code E11.x, no insulin medication record in the past 3 months." It can be seen that the standardized format can eliminate vague expressions such as "no insulin use" in the original text of the research plan, which is conducive to providing a quantifiable and verifiable data basis for subject screening.

[0038] Optionally, in some embodiments, determining the element extraction prompt word according to the research plan, that is, step S1011 may include: Step S10111, identifying the research type of the research plan; Step S10112, obtaining international standard guidelines and / or user-predefined prompt word templates corresponding to the research type; Step S10113 : determining the element extraction prompt words according to the research plan, the research type, the international standard guidelines, and / or the user-predefined prompt word template.

[0039] For step S10111, the macro model can, for example, perform semantic parsing on the electronic text of the study protocol, extracting key features from the core aspects of the study design to determine the study type. For example, the macro model can first identify whether the study protocol mentions "random assignment" or "control group setting." If the study protocol mentions "randomly assigning patients to different treatment groups," the study type is a randomized controlled trial (RCT); if the study protocol emphasizes "long-term follow-up of disease progression in a specific population," the study type is a cohort study; and if the study protocol focuses on "surveying the disease status of a population at a specific point in time," the study type is a cross-sectional study. For example, if the study protocol describes "randomly assigning patients to a drug group A or a placebo group in a 1:1 ratio, with six-month follow-up to observe blood glucose changes," the macro model can identify the study type as an RCT by capturing keywords such as "random assignment" and "control intervention," ensuring that the study type identification conforms to standard classification logic for clinical research.

[0040] For example, the study type may include a randomized controlled trial, a cross-sectional study, a cohort study, a diagnostic study, etc. In some embodiments, the study type may also be specified by the user according to personalized needs.

[0041] For step S10112, exemplarily, based on the identified research type, the built-in knowledge graph association mechanism can be automatically triggered in the large model prompt word generation logic. This automatically associates the international standard guidelines and / or user-defined prompt word templates corresponding to the research type. For example, for RCT type, the CONSORT statement can be automatically retrieved; for cohort studies, the STROBE statement can be retrieved. If there is a user-predefined template (such as a tumor research design template developed within the company), the large model can synchronously load the preset instructions in the template (such as "extract specific gene mutation detection standards"), integrate these standard requirements with the original content of the research plan, and form an element extraction prompt word that includes industry standards and customized requirements, to ensure that the subsequently extracted research design elements comply with both international common standards and meet the needs of personalized research scenarios.

[0042] The international standard guidelines may include but are not limited to: SPIRIT statement, CONSORT statement, STROBE statement, etc.

[0043] The user-defined prompt word templates are sets of instructions pre-set by the user for specific research types based on specific research needs, industry regulations, or institutional standards. These templates can specify the study design elements to be extracted and the presentation requirements for different study types. For example, in an oncology clinical trial template, users can customize instructions such as "Extract PD-L1 expression detection methods" and "Define imaging assessment criteria for disease progression."

[0044] For step S10113, for example, the electronic text of the research plan can be parsed through the big model to extract the core content such as inclusion and exclusion criteria, intervention measures, etc.; combined with the identified research type, the standard requirements for research design in the corresponding international standard guidelines can be called; and / or personalized instructions for this type in the user's predefined template can be loaded. Subsequently, the big model can semantically fuse the key information of the original text of the research plan, the standard requirements of the research type, and the user-customized instructions, and convert them into structured prompt statements. For example, the description of the "random grouping of diabetic patients" is integrated with the CONSORT specifications of the RCT type and the requirements of "extracting sample size calculation details" in the user template to generate the element extraction prompt words "Extract the randomization method, sample size estimation basis, and user-specified blood glucose monitoring frequency of the RCT study according to the CONSORT standards", ensuring that the prompt words have both the original intention of the plan, industry standards, and customized needs.

[0045] Optionally, in some embodiments, extracting prompt words based on the elements and outputting a structured list of research design elements through the large model, that is, step S1012 may include: Step S10121, determining target items of research design elements according to the PICO framework of the research type; Step S10122: Based on the element extraction prompt word, query the large model item by item for the research design elements of the target item in the research plan; Step S10123: forming the structured list according to the retrieved research design elements.

[0046] Regarding step S10121, the PICO framework, a classic paradigm for clinical study design, exemplifies how it divides core research content into categories such as study subjects, interventions, control groups, and outcomes. The division of study design elements within the PICO framework needs to be dynamically adjusted based on the specific study type. In this step, using the PICO framework as a guide, key areas of focus for study design are identified and the scope of extracted elements is defined.

[0047] For example, for a clinical study of a new glucose-lowering drug, the PICO framework, based on this study type, identifies the target items as: "Patients with type 2 diabetes (study subjects)", "How the new glucose-lowering drug is used (intervention)", "Traditional glucose-lowering drugs as a control (control regimen)", and "Changes in blood glucose control indicators (outcome indicators)". Another example is the PICO framework identifying the target items as: "Age ≥ 18 years", "Patients with type 2 diabetes diagnosed according to WHO criteria" (study subjects), "Combined severe hepatic and renal insufficiency", "Pregnant patients" (exclusion criteria), "Experimental group uses drug A, control group uses placebo" (intervention / exposure factor), "Primary endpoint is change from baseline in fasting blood glucose after 12 weeks of treatment", "Secondary endpoint includes overall survival within 24 weeks" (clinical outcome), "BMI", and "Previous history of malignant tumors" (key covariates).

[0048] Regarding step S10122, illustratively, when extracting study design elements based on element extraction prompts and target items, the system may first semantically match the element extraction prompts, which include the study protocol, the study type, and the international standard guidelines and / or user-defined prompt templates, with each target item (e.g., inclusion criteria, intervention measures, etc.). For example, for a clinical trial of a type 2 diabetes drug, if the element extraction prompt "Type 2 diabetes must meet WHO diagnostic criteria" successfully matches the target element "Disease Diagnostic Criteria" under "Inclusion Criteria," the system may generate a structured question based on the key information in the prompt, such as "According to the WHO diagnostic criteria, what are the specific blood glucose indicators and testing requirements for patients with type 2 diabetes as defined in the study protocol?" By executing this "match-question" process for each successfully matched target element, the system can guide the large model to accurately locate the corresponding content within the electronic text of the study protocol, ensuring that the direction of element extraction is compliant with standards and targeted.

[0049] For step S10123, illustratively, after the large model completes the extraction of each target element, the system can standardize and integrate the scattered information. The large model can automatically verify the consistency of the extracted content (such as excluding statements that contradict the main diagnostic criteria) to ensure the accuracy of the information. Subsequently, the system can classify and organize all elements in a structured format such as JSON, summarize the specific content under categories such as "inclusion criteria", "exclusion criteria", and "intervention measures", and form a complete structured list. Taking the type 2 diabetes study as an example, the final list will clearly present the diagnostic criteria of "fasting blood glucose ≥7.0mmol / L and glycated hemoglobin ≥6.5%", the intervention details of "drug A taken orally once a day for two weeks", and other information. This structured integration not only makes it easier for researchers to intuitively obtain core content, but also provides a standardized data foundation for subsequent patient screening and research analysis based on real-world data.

[0050] For example, when extracting study design elements from a clinical trial protocol for a type 2 diabetes drug based on secondary prompts, the system first semantically matches the secondary prompts, including the protocol text, the study type "randomized controlled trial (RCT)," the loaded CONSORT guidelines, and the user-defined "need to clarify diagnosis and exclusion criteria" template, with target items such as "inclusion criteria" and "intervention measures." When the system finds the "disease diagnostic criteria" target element under "inclusion criteria," based on the secondary prompt's directive that "type 2 diabetes must meet the WHO diagnostic criteria," it asks the macromodel, "According to the WHO diagnostic criteria, what are the specific blood glucose indicators and testing requirements for patients with type 2 diabetes as defined in the protocol?" The macromodel extracts the phrase "fasting blood glucose ≥ 7.0 mmol / L and glycated hemoglobin ≥ 6.5%" from the electronic text and automatically verifies consistency with the diagnostic criteria in other sections of the protocol to avoid inconsistencies (e.g., eliminating conflicting statements such as "diagnosis relies solely on postprandial blood glucose."). After completing the "match-question-extract-verify" process for all target elements (such as intervention drug doses and primary clinical outcomes), the large model integrates the scattered information into a structured list in JSON format, for example: { "Inclusion criteria": {"Disease diagnosis criteria":"Fasting blood glucose ≥ 7.0 mmol / L and glycosylated hemoglobin ≥ 6.5%","Age requirement":"18-65 years old"}, "Exclusion criteria":{"eGFR<60","pregnancy status"}, "Intervention":{"Drug A taken orally once daily for two weeks"}, … "Clinical indicators": {"The patient has tumor volume detection values ​​before treatment and 3 months after treatment"} … } It is not difficult to find that in the embodiment of the present application, by determining the element extraction prompt words according to the research plan, the original text of the research plan, the research type specification and the user-defined template and other information are integrated into clear instructions, which provides a precise extraction direction for the large model; on this basis, according to the element extraction prompt words, the structured list of research design elements is output through the large model, and the information processing capability of the large model is used to convert the unstructured research plan into a standardized and systematic data structure. Therefore, the combination of these two steps makes the extraction process of research design elements more efficient and accurate, can avoid omissions and errors that may occur in manual extraction, and can quickly and completely sort out the core content of the research plan, thereby providing a reliable data basis for subsequent research analysis, data processing and research results evaluation, etc., and significantly improving the efficiency and quality of research design element processing.

[0051] Third embodiment The third embodiment of the present application relates to an intelligent inclusion and exclusion method based on clinical research. The third embodiment is an improvement on the second embodiment. The specific improvement is that: in this embodiment, a specific implementation method for forming the structured list based on the queried research design elements is provided.

[0052] Specifically, forming the structured list based on the queried research design elements, i.e., step S10123, may include the following steps: Step S3A, performing a conversion operation on the retrieved research design elements to generate specific conditions for data interpretation; Step S3B: forming the structured list according to the specific conditions.

[0053] Regarding step S3A, for example, because original study protocols are often described in natural language, they contain ambiguous terms such as "stable condition" or unstructured medical terms such as "type 2 diabetes." These can easily lead to ambiguity when directly applied to real-world data (e.g., electronic medical records and laboratory test reports). Therefore, in this step, the retrieved study design elements are converted: abstract concepts are quantified into specific numerical standards (e.g., "stable condition" is refined into "vital signs within normal range for 72 consecutive hours"), and medical terms are mapped to standard codes (e.g., replacing "type 2 diabetes" with ICD-10 code E11.9). This transforms the study design elements into specific conditions that can be directly compared with data fields in electronic medical records, such as blood pressure values ​​and medication records. This process eliminates the ambiguity inherent in natural language and facilitates accurate interpretation for subsequent data screening.

[0054] Regarding step S3B, the standardized specific conditions can be integrated into a structured list. For example, specific conditions such as "age ≥ 18 years," "ICD-10 code E11.9," and "normal vital signs for 72 consecutive hours" can be categorized and arranged by "inclusion criteria" and "exclusion criteria," and presented in a table or JSON format. This structured list not only clearly maps study design elements to real-world data fields (e.g., "ICD-10 code E11.9" corresponds to the "Disease Diagnosis Code" field in the electronic medical record), but also ensures that subsequent data retrieval tools (such as SQL queries) can accurately screen eligible patients, improving the efficiency and reliability of research data collection.

[0055] Optionally, in some embodiments, the conversion operation may include at least one of the following: Decompose the compound type inclusion condition into atomic conditions; Translate intervention plan descriptions into actionable medical behaviors; The calculation requirements of outcome indicators are converted into the existence requirements of original test data at preset time nodes.

[0056] Specifically, composite inclusion criteria are comprehensive screening criteria composed of multiple related elements in the research plan (such as "early-stage breast cancer is confirmed to be triple-negative breast cancer based on pathological examination after surgery"). Because they contain multi-dimensional limiting factors and are complex to express, they are prone to verification difficulties due to semantic ambiguity when directly used for electronic medical record screening. When breaking them down, it is necessary to first analyze the logical structure and split them into indivisible basic judgment units: taking breast cancer cases as an example, they can be deconstructed into "diagnosed as breast cancer" (matching the medical record diagnosis field or ICD code), "early breast cancer" (judged based on TNM stage or tumor size), and "postoperative pathological confirmation of triple-negative" (verified by negative results for estrogen receptor, progesterone receptor, and human epidermal growth factor receptor 2 in the pathology report). Each atomic condition corresponds to a specific data field in the electronic medical record or test report, avoiding the complexity of judgment caused by multiple factors and achieving accurate screening of research subjects.

[0057] Specifically, the description of the intervention measures in the research plan (such as "adopting a new therapy" and "implementing Plan A") is expressed in abstract and general language, lacking specific details such as drug names, dosages, and frequencies, and cannot be directly matched with the medication records and doctor's order data in the electronic medical record. During conversion, the vague concepts need to be refined into actionable medical behaviors: for example, "performing Plan A treatment" is clarified as "drug A is taken orally once a day, 50 mg each time, for two weeks." By supplementing elements such as drug dosage form, dosage, and course of treatment, the intervention measures are transformed from an abstract state to a concrete action. This conversion can accurately match specific records in real-world data, provide a clear basis for judging the selection of patients who meet the intervention conditions, and improve the accuracy of data collection.

[0058] Specifically, some outcome measures must be calculated (such as "change in tumor volume over three months"), which rely on comparative calculations of data at different time points and cannot be obtained directly from the original data. During conversion, the computational requirements must be broken down into a definition of data completeness: Taking tumor volume change as an example, it is necessary to ensure that the original data contains tumor volume measurements before and three months after treatment. By clarifying the time points for data collection and the type of indicator, the existence requirement of "patients must have tumor volume measurement records before and three months after treatment" is established. This conversion can avoid computational difficulties caused by missing data and ensure that subsequent analysis can accurately derive outcome indicator results based on complete data.

[0059] Optionally, in some embodiments, after determining the research design elements according to the received research plan, the method may further include: Step S201, sample training step: using historical research plans and corresponding research design elements as training samples, optimizing the large model through supervised fine-tuning or low-rank fine-tuning; Step S202, dynamic calibration step: in response to identifying a fuzzy description item, perform the following operations: search the medical knowledge base to obtain relevant clinical guidelines; generate quantifiable rules based on the clinical guidelines to replace the fuzzy description item; and output the quantifiable rules as a new research design element.

[0060] For step S201, illustratively, the sample training step is intended to optimize the performance of the large model in the task of extracting research design elements through historical data. Specifically, the historical research plan and its corresponding confirmed research design elements are used as training samples, and supervised fine-tuning or low-rank fine-tuning techniques are used to adjust the parameters of the large model in a targeted manner. By constructing specific training tasks in the scenario of extracting research design elements, the model is guided to deeply learn the semantic patterns, data features and logical relationships in medical texts, thereby establishing a knowledge representation system adapted to this field. After systematic training, the large model can more accurately identify and analyze various research design elements based on the learned rules and patterns when processing actual research plans, significantly improve the accuracy and consistency of the extraction results, and provide a reliable data foundation for subsequent research.

[0061] Regarding step S202, a dynamic calibration step can, for example, address semantic uncertainty caused by ambiguous terms in the study plan. When the large model identifies clinical terms lacking clear quantitative criteria, such as "stable condition" or "symptom improvement," it triggers the Retrieval Augmented Generation (RAG) technique. First, it leverages authoritative literature, such as international medical guidelines and clinical consensus documents, stored in the local medical knowledge base to provide a standardized basis for calibration. Second, after vectorizing the fuzzy terms, it uses semantic analysis and knowledge graph matching to retrieve relevant authoritative terms from the knowledge base. Finally, based on the search results and in line with research needs, the fuzzy terms are converted into quantifiable operational rules. For example, "recent poor blood sugar control" can be clarified as "HbA1c > 8% for the past three months." This process not only eliminates inconsistent standards caused by subjective interpretation but also ensures the scientific and authoritative nature of the study criteria. It transforms previously ambiguous terms into actionable and verifiable study design elements, effectively improving the standardization of study designs and data availability.

[0062] It should be noted that this embodiment may also be an improvement based on the first embodiment.

[0063] It is not difficult to find that in the embodiment of the present application, by performing a conversion operation on the queried research design elements, the abstract and vague original expressions can be converted into specific conditions that can be directly interpreted based on real-world data, so that data screening and matching have a clear and accurate basis, effectively avoiding data misjudgment caused by unclear semantics; therefore, these specific conditions are further integrated to form a structured list, thereby realizing the systematic and standardized presentation of research design elements. In this way, it can ensure the seamless connection between research design elements and real-world data, improve the accuracy and efficiency of data extraction, and provide a clear and orderly data foundation for subsequent data analysis, research results evaluation and other work, significantly improving the reliability and operability of the research.

[0064] Fourth embodiment The fourth embodiment of the present application relates to an intelligent inclusion and exclusion method based on clinical research. The fourth embodiment is an improvement on the first embodiment. Specifically, the improvement is that: in this embodiment, a specific implementation method for determining inference rules based on the research design elements is provided.

[0065] Specifically, in some embodiments, determining the inference rules based on the research design elements, i.e., step S102, may include: Step S1021, semantically matching the research design elements with the data dictionary to generate element-variable mapping relationships; Step S1022: Determine an inference rule based on the element-variable mapping relationship.

[0066] For step S1021, illustratively, this step aims to establish an association between the research design elements and the data dictionary, and to generate an element-variable mapping relationship by semantically matching the research design elements with the data dictionary. Since there are semantic differences between the expression of the research design elements and the variables in the real-world data, this step uses a large model to parse the semantics of the elements, and can combine direct matching and indirect reasoning to find the correspondence between the research design elements and the data dictionary variables. For example, for the research design element "patients with diabetes", by parsing the medical meaning of "patients with diabetes", it can be directly associated with clear diagnostic variables such as "diagnosis: past medical history", and auxiliary diagnostic variables such as "laboratory test: fasting blood glucose value" can be indirectly inferred through medical logic to obtain the element-variable mapping relationship.

[0067] For step S1022, illustratively, the data dictionary may contain detailed information such as form name, variable name, variable label / description, variable type, and coding rules. In this step, specific data screening logic may be formulated based on the variables and their attributes corresponding to each research design element. For example, when the research design element "age ≥ 18 years old" is mapped to the "age at consultation" variable, combined with the setting of "age at consultation" as a numerical variable in the data dictionary, the inference rule is determined to be "screening patients whose age at consultation field value is greater than or equal to 18 from the electronic medical record." By converting the element-variable mapping relationship into clear judgment conditions, the research design elements are made operational and can be directly applied to the screening and analysis of real-world data.

[0068] Optionally, in some embodiments, semantically matching the research design elements with a data dictionary to generate element-variable mapping relationships, that is, step S1021 may further include: Step S10211, parsing the semantic information of the research design elements through the large model; Step S10212: semantically match the semantic information with the variable labels in the data dictionary, and output a matching list of elements and variable names including the variable name, variable label, and the form to which it belongs.

[0069] For step S10211, illustratively, the powerful semantic understanding ability of the big model is used to deeply analyze the semantic information of the research design elements. Research design elements in the medical field often contain complex professional terms and implicit logic. For example, "adopting an intensive blood sugar lowering program" may involve multiple information such as drug type, dosage, and frequency. The big model uses natural language processing technology to identify key semantics in elements, split complex expressions, and explore potential medical associations. For example, for "poor recent blood sugar control", the big model can parse out its dual semantics involving blood sugar indicators and time ranges, providing an accurate semantic basis for subsequent matching with the data dictionary.

[0070] For step S10212, illustratively, two strategies can be used during the matching process: direct matching and indirect reasoning. Direct matching can be used to directly map semantically clear elements to variable labels through string similarity calculation or regular expressions. For example, the variable label corresponding to "age ≥ 18 years old" is directly matched to "age at medical consultation" (belonging to the form: basic information); indirect reasoning can be based on the medical knowledge graph and implement semantic extension matching based on the medically associated potential variables of the elements. For example, variable labels such as "glycated hemoglobin" can be indirectly inferred from the element "patients with diabetes". The generated matching list can be presented in a structured form, clearly marking the source form of each variable (such as "laboratory test" and "diagnosis record") and the matching type, to ensure that the correspondence between the research design elements and the data dictionary variables is intuitive and traceable, providing an accurate data association basis for the formulation of subsequent inference rules.

[0071] Optionally, in some embodiments, determining the inference rule according to the factor-variable mapping relationship, that is, step S1022, may further include: Step S10221′: construct and fill a prompt word template for each generated element-variable mapping relationship; wherein the prompt word template is injected with the following information: the medical definition of the target element, the matching variable name, the variable label and type, and the encoding rules of the data dictionary; Step S10222', input the filled prompt words into the large model, and output the inference rules in pseudocode form in combination with the data dictionary; the inference rules include at least one of the following: threshold judgment of numerical variables, keyword matching of text variables, and code value mapping of categorical variables.

[0072] For step S10221', for example, a prompt word template needs to be constructed and populated for each element-variable mapping relationship to generate semantic instructions that can guide the large model to output inference rules. Specifically, the medical definition of the target element can be first extracted, for example, clarifying the specific diagnostic criteria for "patient has diabetes." The matching variable name, variable label, and type can also be obtained, such as the text variable "past medical history." The encoding rules in the data dictionary can also be retrieved, such as the ICD-10 code for diabetes is E11.9. Furthermore, this information can be injected into the preset prompt word template. For a single-variable mapping, for example, the mapping relationship between "patient has diabetes" and "past medical history," a prompt word is generated: "How to infer [patient has diabetes] based on the [past medical history] variable? Output the inference rule in pseudocode form." For multivariate complex inference, such as requiring simultaneous reference to the "admission diagnosis" and "discharge diagnosis" variables, a prompt word containing multiple variable names can be constructed, such as "How to infer [patient has diabetes] based on the [admission diagnosis] and [discharge diagnosis] variables? Output the inference rule in pseudocode form." By filling in key information, we can ensure that the prompt words can accurately convey the inference requirements, so that the large model can deeply understand the logical relationship between factors and variables, and thus generate accurate inference rules.

[0073] For step S10222', for example, the populated prompt word can be input into the large model and combined with the data dictionary to generate pseudo-code inference rules. The large model uses the data dictionary as a knowledge base and outputs corresponding judgment logic based on the variable type (numeric, text, categorical, etc.) and encoding rules in the prompt word. For example: For the text variables "admission diagnosis" and "discharge diagnosis", generate "'diabetes' in admission diagnosis | 'diabetes' in discharge diagnosis" and make judgments through keyword matching; For numerical variables (such as blood glucose), combine the threshold requirements of the data dictionary to generate "fasting blood glucose value>=7.0mmol / L"; For categorical variables (such as disease codes), "ICD-10 code == E11.9" is generated through coding dictionary mapping.

[0074] The final output pseudocode form of the inference rules covers the variable judgment logic and corresponding form information, ensuring that the inference rules can be directly used for the screening and verification of electronic medical record data. For example, "logical type: whether suffering from diabetes == TRUE" can be used to directly match the Boolean variables in the data dictionary.

[0075] Optionally, in some embodiments, during the element-variable label matching and / or inference rule generation process, the method may further include: calling a RAG knowledge base to perform the following operations: Step S301: construct a medical knowledge base to store disease guidelines, clinical literature, medical coding dictionaries, and historically confirmed factor-variable mapping results; Step S302: In response to the semantically ambiguous element description, the knowledge base is searched to obtain relevant medical definitions and threshold standards, and the variable matching logic is modified; if a matching result with the same element in the historical archive is retrieved, the confirmed variable label is directly called; Step S303, when generating inference rules in pseudocode form, search the knowledge base to obtain clinical diagnostic standards and historical rules: if there are identical generation requirements, directly output the historical archived rules; if there are similar requirements, inject the historical rules as reference context into the large model prompt words.

[0076] Regarding step S301, illustratively, this operation aims to build an integrated medical knowledge storage system, which structurally stores information such as disease guidelines (such as diabetes diagnosis and treatment standards), clinical literature, medical coding dictionaries (such as the ICD-10 coding library), and historically confirmed factor-variable mapping results. For example, the knowledge base will record the medical definition of "fasting blood glucose ≥7.0mmol / L is the diagnostic standard for diabetes," as well as historical matching cases of "patients with diabetes" corresponding to variables such as "fasting blood glucose value" and "glycated hemoglobin." By building this knowledge base, authoritative knowledge support and historical experience reference can be provided for factor-variable matching and inference rule generation, ensuring the accuracy and consistency of subsequent operations.

[0077] For example, in step S302, when encountering a semantically ambiguous element description (such as "abnormal blood sugar control"), the system will search the medical knowledge base to enhance the matching logic: first, it obtains relevant medical definitions and threshold standards (such as "fasting blood sugar > 7mmol / L is the diagnostic threshold for diabetes"), converting the ambiguous description into a clear variable matching condition; if there are historical matching results for the same element in the knowledge base (such as the previously confirmed "abnormal blood sugar control" corresponding to the "fasting blood sugar value" variable), the confirmed variable label is directly called to avoid repeated reasoning. For example, when processing "recent poor blood sugar control", the quantitative standard of "glycated hemoglobin > 8% in the past three months" stored in the knowledge base is retrieved, and the matching logic can be modified to accurately associate this element with the "glycated hemoglobin" variable, improving the accuracy and efficiency of matching.

[0078] For step S303, illustratively, when generating inference rules in pseudo-code form, the system can search the knowledge base to obtain clinical diagnostic standards and historical rules: if the rule to be generated is exactly the same as the historical archive requirements in the knowledge base (such as also inferring diabetes based on the "admission diagnosis" variable), the historically confirmed rule (such as "'diabetes' in admission diagnosis") is directly output; if the requirements are similar (such as inferring diabetes based on the "discharge diagnosis"), the historical rules are injected into the large model prompt words as a reference context to assist in generating more reliable rules. For example, when generating the rule "inferring anemia based on the 'blood routine' variable", the clinical standard "hemoglobin <110g / L is the diagnostic standard for anemia" and the historical rule "hemoglobin value <110g / L" in the knowledge base are retrieved, and numerical judgment pseudo-code can be directly generated based on this to ensure that the rules comply with medical standards and are consistent with historical experience.

[0079] Optionally, in some embodiments, the feature-variable label matching process may further include a semantic embedding optimization step: Step S401: Pre-train semantic embedding for all variable labels in the data dictionary to generate high-dimensional vector representations; Step S402: Input the target research element description into the same embedding model and convert it into an element semantic vector; Step S403: Calculate the similarity between the element semantic vector and each variable label vector, and select variable labels with similarity exceeding a preset threshold to add to the matching candidate set; Step S404 : performing subsequent element-variable mapping relationship generation based on the matching candidate set.

[0080] Regarding step S401, illustratively, this step uses a pre-trained model to perform semantic embedding on variable labels in the data dictionary (e.g., "eGFR," "serum creatinine"), converting each label into a vector representation in a high-dimensional space. This vector not only captures the literal meaning of the label but also captures the underlying medical semantic associations (e.g., the clinical correlation between "eGFR" and "chronic kidney disease"). For example, after processing using a pre-trained model specifically for the medical field, the vector for "serum creatinine" will be closer to the vector for "renal function indicators" in high-dimensional space, thus overcoming the limitations of character matching and laying the foundation for precise semantic matching.

[0081] For step S402, for example, the study design elements (e.g., "chronic kidney disease") are input into the same embedding model as the variable labels to generate corresponding element semantic vectors. This process converts the textual study elements into vector space representations with the same dimensionality as the variable labels, ensuring that the two are comparable in the same semantic space. For example, the semantic vector for "chronic kidney disease" will cluster in high-dimensional space with vectors for related concepts such as "renal impairment" and "glomerular filtration rate," thereby capturing the implicit medical association between the element description and the variable label, even if the two do not directly overlap in terms of characters (e.g., "chronic kidney disease" and "eGFR").

[0082] For step S403, variable labels with similarities above a preset threshold can be included in the matching candidate set. For example, when processing the "chronic kidney disease" element, the model will calculate the similarity between its vector and the variable label vectors of "eGFR," "serum creatinine," and "urine protein." Since these variables are key indicators for kidney disease diagnosis, the similarity between their vectors and the element vectors will exceed a threshold (such as 0.7), and thus they will be selected into the candidate set. This semantic-based screening mechanism can avoid omissions caused by "character matching" (such as missing variables with no character overlap but semantically related), improving the comprehensiveness of the match.

[0083] Regarding step S404, illustratively, after obtaining a candidate set of semantically related variable labels, the system can further combine strategies such as direct matching and indirect reasoning to generate the final feature-variable mapping relationship. For example, the "eGFR" and "serum creatinine" variables in the candidate set can be linked to the diagnostic criteria for "chronic kidney disease" (e.g., eGFR < 60 mL / min / 1.73 m²) through the medical knowledge graph, thereby forming a mapping relationship of "chronic kidney disease → laboratory test: eGFR." This step utilizes the semantically embedded optimized candidate set to ensure that the mapping relationship is based on both the structured information of the data dictionary and the semantic logic of the medical field, thereby improving the accuracy and generalization of feature-variable matching.

[0084] It should be noted that this embodiment may also be an improvement based on the second embodiment and / or the third embodiment.

[0085] It is not difficult to find that in the embodiment of the present application, by semantically matching the research design elements with the data dictionary to generate an element-variable mapping relationship, a bridge can be built between the research design text and the real-world data, so that the abstract research elements are accurately mapped to the specific variables in the data dictionary, and then the inference rules are determined based on the mapping relationship, so that the research elements are converted into executable data screening logic. Therefore, this method not only eliminates the ambiguity of natural language descriptions, but also uses the standardized structure of the data dictionary to ensure the accuracy and operability of the inference rules, and ultimately achieves accurate mapping from research plans to clinical data, providing a reliable logical basis for research object screening and data analysis.

[0086] Fifth embodiment

[0087] The fifth embodiment of this application relates to an intelligent inclusion and exclusion method based on clinical research. The fifth embodiment is an improvement on the first embodiment. Specifically, the improvement is that this embodiment provides a specific implementation method for generating a discriminant code for each study design element based on the inference rule.

[0088] Specifically, in some embodiments, generating a discriminant code for each research design element according to the inference rule, that is, step S103 may include the following steps: Step S1031: extracting structured metadata of the variables involved in the inference rule from the data dictionary; the metadata includes: the form name to which the variable belongs, the variable name, the variable data type, and the variable value range definition; Step S1032, generating code generation prompt words according to the inference rule, the structured metadata and the target language; Step S1033: convert the code generation prompt words into discriminant codes through the large model.

[0089] Regarding step S1031, illustratively, before generating the discriminant code, the system extracts structured metadata for the variables involved in the inference rule from the data dictionary. This structured metadata serves as the basis for generating the discriminant code. Specifically, for each inference rule, such as "Fasting Blood Glucose Value > 7 mmol / L," the system retrieves detailed attributes of the corresponding variable from the data dictionary. These include the name of the form to which the variable belongs (e.g., the "LABTEST" laboratory test form), the variable name (e.g., "FBG"), the data type (e.g., "numeric"), and the range definition (e.g., "mmol / L," "normal reference range," etc.). For example, when the inference rule involves the variable "Fasting Blood Glucose Value," the extracted metadata would be "Table Name: LABTEST | Variable Name: FBG | Data Type: Numeric | Unit: mmol / L." This information accurately reflects the variable's storage structure and semantic definition in the database, thereby avoiding code generation errors caused by ambiguous variable attributes and ensuring that the subsequently generated SQL and other discriminant code accurately matches the database structure.

[0090] For step S1032, code generation prompts are generated, illustratively, based on the inference rule, structured metadata, and target language. This step integrates the inference rule logic (e.g., "Fasting blood glucose value > 7 mmol / L"), variable metadata (e.g., form and data type), and user-specified code language (e.g., SQL) into a pre-set template. For example, a prompt might be "Given the database structure [Fasting blood glucose value | Table name: LABTEST | Variable name: FBG | Data type: numeric | Unit: mmol / L], using the variable [Laboratory test: Fasting blood glucose value], generate SQL code to express the inference rule [Fasting blood glucose value > 7 mmol / L]." In this way, the prompt provides a complete code generation context for the large model, ensuring that the generated discriminant code conforms to the database structure and the grammatical rules of the specified language.

[0091] For step S1033, the code generation prompt is illustratively input into the macromodel, which then converts it into executable discriminant code. Based on the variable metadata (e.g., numeric data type) and inference rules (e.g., threshold determination) in the prompt, the macromodel generates code that conforms to the target language syntax. For example, for the prompt, the macromodel generates the SQL code "SELECT * FROM LABTEST WHEREFBG > 7.0," which can be directly executed in the database to filter out records with fasting blood glucose values ​​greater than 7 mmol / L. The generated discriminant code not only meets the structured requirements of the data dictionary but also accurately implements the logic of the inference rules, providing an automated tool for subsequently screening patients eligible for the study protocol from the medical database.

[0092] Optionally, in some embodiments, after generating the inference rules, the method may further include the step of a code verification feedback loop: Step S501: Generate a simulated data set based on the data dictionary structure and fill it with medically compliant sample data according to variable type; Step S502, running the generated pseudo-code rules on the simulation data set to capture execution errors or abnormal outputs in real time; Step S503, when an execution error is detected, the error log, data snapshot and execution environment are injected into the code generation prompt word to re-trigger the inference rule generation step; and the verification is repeated until the pseudocode rule runs stably in the simulation environment.

[0093] For step S501, for example, a simulated data set needs to be generated based on the data dictionary structure and filled with medically compliant sample data according to variable type. In specific operations, the system can generate simulated data that conforms to medical logic based on metadata such as the variable's form, data type (e.g., numeric, text), and range definition (e.g., blood glucose range, ICD coding rules) recorded in the data dictionary. For example, a numerical sample within the range of 5.0-12.0 mmol / L is generated for the "fasting blood glucose value" variable, and a text record containing keywords such as "diabetes" and "hypertension" is filled for the "disease diagnosis" variable.

[0094] For step S502, the generated pseudocode rules are exemplarily run on a simulated data set to capture execution errors or abnormal outputs in real time. By applying the pseudocode (such as SQL queries and Python conditional statements) converted from the inference rules to the simulated data, the system can verify the logical correctness and compatibility of the code. For example, when running the SQL code "SELECT * FROM PATIENTSWHERE fasting blood glucose value > 7.0", if the unit of the "fasting blood glucose value" variable in the simulated data is marked as "mg / dL" instead of "mmol / L", an execution error of unit mismatch will be triggered; if the threshold judgment logic in the code is written in reverse (such as "<7.0"), the output will not meet the expected screening results.

[0095] Regarding step S503, illustratively, when an execution error is detected, the error log, data snapshot, and execution environment can be injected into the code generation prompt to re-trigger the inference rule generation step. For example, when a "fasting blood glucose value unit mismatch" error is detected, the system can integrate information such as the error type (unit conversion exception), the actual unit in the simulation data (mg / dL), and the code execution environment (SQL syntax requirements) into the code generation prompt, such as "Current code execution failed due to unit mismatch. The unit of 'fasting blood glucose value' in the simulation data is mg / dL. It is necessary to convert 7.0mmol / L to 126mg / dL according to medical standards and regenerate the SQL rule." By feeding this error information back to the macro model, the system can automatically correct the inference rules and regenerate the code until the pseudocode runs stably in the simulation environment, ensuring that the resulting discriminant code accurately adapts to the structure and logic of real medical data.

[0096] It should be noted that this embodiment may also be an improvement based on any one or more of the second to fourth embodiments.

[0097] It is not difficult to find that in the embodiment of the present application, by extracting the structured metadata of the variables involved in the inference rules from the data dictionary, the form, data type and value range definition of the variables can be clarified, providing accurate basic information for code generation, and then combining the inference rules, structured metadata and target language to generate code generation prompt words, so that the prompt words contain complete code generation elements. Therefore, when the prompt words are converted into discriminant codes through the large model, it can be ensured that the generated code conforms to the data dictionary structure, inference logic and target language grammar, thereby realizing the accurate conversion from research design elements to executable code, which is conducive to the efficient screening of patient data that meets the research plan.

[0098] Sixth embodiment The sixth embodiment of this application relates to an intelligent inclusion and exclusion method based on clinical research. The sixth embodiment is an improvement on the first embodiment. Specifically, the improvement is that: in this embodiment, a specific implementation method for determining a list of patients who meet the research plan based on the discrimination code is provided.

[0099] Specifically, in some embodiments, determining the list of patients who meet the research plan based on the discrimination code, that is, step S104, may include: Step S1041, executing the discrimination code to obtain discrimination results of each research design element; Step S1042: Determine a list of patients who meet the research plan based on the identification results.

[0100] For step S1041, for example, by executing the discrimination code, the patient data in the medical database can be automatically screened to obtain the discrimination results of each research design element. The discrimination code is generated based on the research design elements and the data dictionary, and contains precise logical judgment conditions (such as numerical thresholds, text matching rules). For example, when executing the SQL discrimination code "SELECT * FROM PATIENTSWHERE fasting blood glucose value > 7.0 AND age >= 18", the system will traverse the patient data table and judge the "fasting blood glucose value" and "age" fields in each record, marking the records that meet the conditions (blood glucose exceeds the standard and age meets the standard) as "true" and those that do not meet the conditions as "false", thereby generating a discrimination result set for each element, providing a quantitative basis for subsequent screening.

[0101] For example, step S1042 comprehensively evaluates all study design elements to determine the final list of patients who meet the study protocol. This step logically integrates the judgment results of each element (e.g., combining conditions using "and" and "or" relationships). For example, only patients who simultaneously meet all the conditions, such as "fasting blood glucose level > 7.0," "age >= 18," and "no severe cardiovascular or cerebrovascular disease (corresponding to the relevant diagnostic variables in the data dictionary)," are included in the list. The system filters patient records with a "true" judgment result to output a complete and accurate list of patients who meet the study protocol.

[0102] Optionally, in some embodiments, executing the discrimination code to obtain the discrimination results of each research design element, that is, step S1041, may further include: Step S10411, executing the discrimination code in the target database to obtain an original discrimination result set; Step S10412: generating a discrimination result matrix based on the original discrimination result set; wherein: the row index of the discrimination result matrix corresponds to the patient identifier, the column index corresponds to the study design element, and the matrix cell value stores the discrimination result of each patient-element combination; Step S10413: Using the discrimination result matrix as the discrimination result of each research design element.

[0103] For step S10411, the system can, illustratively, connect to the target database via a secure channel and sequentially execute the discriminant code (e.g., SQL query, Python script) corresponding to each study design element. For example, for the two elements "age ≥ 18 years" and "fasting blood glucose level > 7.0 mmol / L," the corresponding code statements are executed, extracting patient records from the database and performing conditional judgments. For each patient, the code generates a Boolean value (TRUE / FALSE) or a fuzzy probability value (0-1) to indicate whether the element condition is met. Missing data is marked as NA, ultimately forming the original set of patient-element discrimination results.

[0104] For step S10412, the discrimination result matrix is ​​illustratively indexed by patient identifiers as rows and study design elements as columns. Each matrix cell stores the discrimination result for the corresponding patient under that element. For example, the row with patient_id = 1 in the matrix would sequentially record the discrimination result (TRUE / FALSE / NA) for each element, such as "age condition" and "blood glucose condition." When a certain element corresponds to multiple discrimination codes (e.g., inferring the same element through different variables), the maximum value can be used (using OR logic for Boolean values ​​and the maximum value for probability values) to ensure comprehensiveness of the results.

[0105] For step S10413, the discrimination result matrix is ​​illustratively output as the discrimination result for each study design element. This discrimination result matrix can be shown in Table 1, visually presenting the matching relationship between all patients and study design elements. For example, the matrix can directly check whether the patient with patient_id=2 meets all inclusion criteria (e.g., all columns are TRUE), or which elements have missing data (displayed as NA) for patient_id=i. This structured output not only provides a clear basis for determining the final patient list in step S1042, but also allows researchers to quickly locate eligible target populations through the matrix. It also facilitates subsequent statistical analysis of the discrimination results (e.g., the compliance rate of each element, the data missing rate, etc.), improving the efficiency and accuracy of research data processing.

[0106] Table 1. Discrimination result matrix

[0107] Optionally, in some embodiments, determining a list of patients who meet the research protocol based on the discrimination result, that is, step S1042 may further include: Step S10421, determining a summary function according to the research plan; Step S10422, inputting the discrimination result matrix into the summary function to calculate the inclusion and exclusion conclusion of each patient; Step S10423: Determine the list of patients who meet the research plan based on the inclusion and exclusion conclusions.

[0108] For step S0421, illustratively, a summary function can be determined based on the specific requirements of the study protocol. This function is used to synthesize the judgment results of each study design element. Different types of elements in the study protocol (such as inclusion criteria, exclusion criteria, intervention conditions, etc.) can be summarized using different logical rules. For example, inclusion criteria generally require "all conditions to be met simultaneously" and can be defined as the Boolean logic "res1&res2&…&resj"; exclusion criteria require "no exclusion conditions to be met simultaneously" and correspond to "!res1&!res2&…&!resj". If the study involves probability-weighted judgment, weight allocation rules can also be defined (such as a weight of 0.8 for the primary endpoint and 0.5 for the secondary endpoint) to ensure that the summary function conforms to the scientific logic and screening rigor of the study.

[0109] For step S0422, the discrimination result matrix can be input into a summary function to calculate each patient's inclusion or exclusion conclusion. Each row in the matrix corresponds to a patient identifier, and each column corresponds to a factor discrimination result (Boolean value or probability value). The summary function operates on the row data according to a preset logic. Taking Boolean logic as an example, the inclusion criteria integrate the factor results through an "AND" operation (e.g., resPin = res1 & res2), and the exclusion criteria through a "NOT AND" operation (resPout = !res1 & !res2), resulting in the final inclusion or exclusion conclusion RES = resPin & resPout. If probability weighting is used, assuming the inclusion criterion weight is 0.6 and the exclusion criterion weight is 0.4, then RES = resPin ^ 0.6 × resPout ^ 0.4. When RES ≥ the threshold L, the condition is considered met.

[0110] For step S0423, illustratively, a list of patient IDs that meet the inclusion criteria is output based on the inclusion and exclusion conclusions. The system can traverse the discrimination result matrix, filter out patient records whose inclusion and exclusion conclusions are "TRUE" (Boolean logic) or "RES ≥ L" (probability weighting), and extract the corresponding patient identifiers. For example, when the RES value of patient i is 0.92 and the threshold L is set to 0.8, their ID will be included in the list. The final output patient ID list can be directly connected to the electronic medical record system or research database, providing a precise target population for subsequent recruitment and data collection. It also supports researchers to trace the patient's complete medical records through ID to ensure that the enrolled patients strictly meet all requirements of the research protocol.

[0111] It should be noted that this embodiment may also be an improvement based on any one or more of the second to fifth embodiments.

[0112] It is not difficult to find that in the embodiment of the present application, since the medical data is automatically screened by executing the discrimination code, the discrimination results of each research design element can be obtained quickly and accurately, and the compliance of each patient under a single element can be clarified. Then, based on these discrimination results, a comprehensive analysis is performed using the summary function to determine the list of patients who meet the research plan. Therefore, the entire process from research design elements to specific patient screening can be automated, which can avoid the inefficiency and subjective errors of manual screening, and can ensure that the screened patients accurately meet the research requirements based on the standardized logic of the data dictionary and inference rules, providing reliable and efficient target population data support for clinical research.

[0113] Seventh embodiment The seventh embodiment of the present application relates to an intelligent inclusion and exclusion method based on clinical research. The seventh embodiment is an improvement on the first embodiment. Specifically, the improvement is that, after generating the patient list, the seventh embodiment also includes an expert review feedback loop step.

[0114] Specifically, in some embodiments, after determining the list of patients who meet the research protocol based on the discrimination code, the method may further include: Step S601 , extracting time-invariant information and time-series medical data from a database based on the patient list, integrating them along a time axis to generate a structured personal narrative report; Step S602, receiving the admission and exclusion annotation data generated by the medical expert based on the narrative report; Step S603: updating the element-variable mapping relationship, the inference rule, the discrimination code and the large model according to the entry and exit annotation data.

[0115] For step S601, for example, based on the patient list, time-invariant information (such as age, gender, and underlying medical history) and time-series medical data (such as test results, medication records, and treatment history) can be extracted from the database and integrated along a timeline to generate a structured personal narrative report. Specifically, the system can use the patient ID as an index to traverse all forms in the electronic medical record system, extracting all data relevant to the study, such as the patient's blood sugar levels and medication adjustments at different time points. By annotating these discrete data points with timestamps, these data points are linked together into a continuous chain of medical events, forming a complete medical history portrait for each patient. For example, the report might display information such as "The patient was diagnosed with diabetes in January 2023 and started taking metformin. In March, poor blood sugar control led to the switch to insulin therapy. In June, abnormal renal function indicators developed," providing a comprehensive and coherent clinical perspective for expert review.

[0116] For step S602, medical experts review the structured report and, in combination with clinical experience, determine whether the patient truly meets the study protocol. If inaccurate inclusion or omission is discovered, the specific error type can be identified: omissions in inference rules (e.g., excluding chronic kidney disease based solely on the diagnosis name containing 'renal insufficiency,' while ignoring "eGFR < 55"), missing content matches (e.g., failing to identify the product name "Tuozi" as corresponding to "Ixetine"), and ignoring time requirements (e.g., failing to limit the time range for "evidence of folate deficiency six months prior to enrollment"). Experts can select the erroneous elements through the interface and add annotations (e.g., "eGFR < 55 should be added as an exclusion criterion"). This creates an annotated inclusion and exclusion dataset containing patient IDs, erroneous elements, and suggested corrections, providing clear guidance for subsequent model optimization.

[0117] Regarding step S603, illustratively, the factor-variable mapping relationship, the inference rules, the discriminant code, and the large model are updated based on the inclusion and exclusion annotation data. For example, based on the inclusion and exclusion annotation data, when an expert provides feedback that "eGFR < 55 should be used as an exclusion indicator for chronic kidney disease," the system will sequentially optimize each module: this suggestion is fed back to the factor-variable mapping module to supplement the mapping relationship between "chronic kidney disease" and "eGFR"; the exclusion condition "eGFR < 55" is added to the inference rule generation module; and SQL discriminant code containing eGFR threshold judgment is automatically generated. Finally, using the final inclusion list confirmed by the experts as GroundTruth, the large model is fine-tuned or reinforced through supervised learning. Through repeated training and optimization of model parameters, the large model masters the new screening logic. Through this closed-loop feedback mechanism, the accuracy of factor-variable matching, inference rule construction, and discriminant code generation is continuously improved, gradually achieving a high degree of consistency between the automated level of patient screening and the expert annotation results, and achieving precise and intelligent execution of the research plan.

[0118] It should be noted that this embodiment may also be an improvement based on any one or more of the second to sixth embodiments.

[0119] It is not difficult to find that in the embodiment of the present application, since medical data can be extracted and integrated according to the patient list to generate a structured narrative report, it can provide medical experts with comprehensive and time-series patient clinical information, and then support experts to verify the correctness of the enrollment based on the report and mark errors, clarify the problem type and correction direction, and synchronize the marked data to each module for iterative optimization, thereby realizing a closed-loop feedback from data generation, expert review to model improvement. This method not only ensures the accuracy of the enrolled patients through manual verification, but also can transform expert knowledge into a basis for model optimization, thereby continuously improving the accuracy of factor mapping and rule generation, so that the system can continuously enhance the automation capability and reliability of patient screening during iteration.

[0120] Eighth embodiment The eighth embodiment of the present application relates to an intelligent admission and exclusion method based on clinical research. The eighth embodiment is an improvement on the first embodiment, specifically in that it also includes a historical archiving and knowledge reuse mechanism.

[0121] Specifically, in some embodiments, the method further includes: Step S701 , storing the verified research design elements, inference rules, discriminant codes, and associated metadata as a structured historical archive; Step S702, in response to the input of a new research plan, performs the following operations: encode the current research elements and inference rules into semantic vectors; retrieve the Top-K similar archived entries in the historical knowledge base; if there is a completely matching entry, directly reuse the inference rules and discrimination codes; if there is a partially matching entry, perform an automatic adaptation operation to generate adapted inference rules and discrimination codes.

[0122] For step S701, for example, verified study design elements, inference rules, discriminant codes, and associated metadata can be stored as a structured historical archive, forming a reusable knowledge base. The archive includes study design elements (e.g., inclusion and exclusion criteria) within the PICO framework, pseudocode-based inference rules, expert-verified discriminant codes (e.g., SQL queries), and metadata such as study type, data dictionary, and audit records, stored in a structured format such as JSON. For example, for a type 2 diabetes study, the archive would record the inclusion criterion "fasting blood glucose ≥ 7 mmol / L," the inference rule "FBG > 7.0 OR HbA1c > 6.5%," and the corresponding SQL code "SELECT patient_id FROM LABTEST WHEREF BG > 7.0;." This archiving mechanism provides historical references for new research, avoiding duplication of effort.

[0123] Regarding step S702, when a new study proposal is entered, the system first encodes the current study elements and inference rules into semantic vectors, then searches the historical knowledge base for the top-K similar archived entries. If a perfect match exists (e.g., for the same disease, study type, and consistent data dictionary), the verified inference rule and discriminant code can be directly reused. If a partial match exists (e.g., for variable naming differences or threshold adjustments), the following automatic adaptation operation can be performed: variable mappings and threshold parameters from the historical rules are extracted, and information about the differences, such as the variable naming changes in the new data dictionary and the threshold adjustments for the new proposal, is incorporated into the macro model's prompts. For example, if a new study requires "fasting blood glucose ≥ 7.2 mmol / L" as an inclusion criterion, and the retrieved historical archived rule is "fasting blood glucose ≥ 7 mmol / L," after retrieving a partial match, the historical SQL code "FBG>7.0" can be extracted and the threshold difference "7.2" can be incorporated into the prompt. The macro model automatically generates the adapted code "FBG>7.2" based on the historical context, improving the efficiency and accuracy of rule generation and enabling efficient reuse of historical knowledge.

[0124] It should be noted that this embodiment may also be an improvement based on any one or more of the second to seventh embodiments.

[0125] It is not difficult to find that in the embodiment of the present application, a traceable knowledge asset library is constructed by storing verified research design elements, inference rules, discrimination codes and associated metadata as a structured historical archive. When a new research plan is input, the historical knowledge base can be retrieved through semantic vectors, and similar entries can be quickly located and verified discrimination codes can be reused or historical rules can be extracted as context. Therefore, this method can not only avoid the waste of resources in repeatedly developing discrimination codes, but also use historical experience to improve the accuracy and reliability of new rule generation. Especially in partial matching scenarios, the large model automatically adapts to variable differences and threshold updates, thereby achieving standardization and intelligence of research plan execution, significantly shortening the research preparation cycle and reducing human errors.

[0126] Ninth embodiment The ninth embodiment of the present application relates to an intelligent inclusion and exclusion method based on clinical research. The ninth embodiment is an improvement on the first embodiment. Specifically, the ninth embodiment includes a time logic enhancement step for research plans with time constraints.

[0127] Specifically, in some embodiments, the temporal logic strengthening step may include: Step S801, parsing the research plan to identify index events and determine the start time of the individual patient study; Step S802, expanding the linked time variable field in the data dictionary to dynamically associate the medical variable with the timestamp; Step S803, converting the time-related description in the research plan into a time window relative to the start time, and generating an inference rule in pseudo-code form; Step S804: Output the database query code including the time link to ensure that the screening condition corresponds to the start time and the hook time variable.

[0128] Regarding step S801, illustratively, by parsing the study protocol to identify the index event, the starting time T0 of the individual patient study can be determined. In clinical studies, the index event is typically a key medical event with a clear time stamp, such as the date of first diagnosis, first medication date, or surgery date. For example, for the time condition of "no insulin treatment within 6 months before enrollment," the index event "enrollment" must first be identified as T0, and then the time condition must be converted into a relative time range based on T0.

[0129] For step S802, as shown in Table 2, a new "Link Time Variable" field can be added to the existing data dictionary fields (e.g., table name, variable name) to specify the timestamp corresponding to the medical variable (e.g., the "fasting blood glucose" variable is associated with the "test_date" time field). If a variable does not have an explicit time field, it is associated with the standard timestamp of the table to which it belongs by default (e.g., the diagnosis record is associated with "visit_date"). This extension allows previously independent medical variables to have a time dimension attribute. For example, the "ICD10 code" variable can determine the diagnosis time through "visit_date", providing a data foundation for subsequent time condition screening.

[0130] Table 2 Hook Time Variable Examples

[0131] For step S803, exemplarily, when processing time-related descriptions in a research plan, the system can convert the time-related descriptions into a relative time window based on the starting time T0 and generate inference rules in the form of pseudocode. For example, "There is a record of fasting blood glucose ≥ 7 in the 6 months before enrollment" can be converted into a logical expression of "T0−180 ≤ test date < T0 and fasting blood glucose ≥ 7", where T0 is the starting time determined by step S801 (such as the enrollment date). For complex time conditions (such as "The detection time of A is earlier than the detection time of B"), it can be achieved through cross-table time variable comparison, such as associating timestamp fields in different forms for logical judgment. The generated pseudocode rules will combine time constraints with medical index conditions, for example, it is reflected as "WHERE test_date BETWEEN T0-180 AND T0 AND FBG>=7.0" in SQL pseudocode. Ensure that the screening logic meets both the index threshold requirements and the time window requirements to achieve automated processing of time-sensitive research conditions.

[0132] For step S804, exemplarily, convert the pseudocode rules into an executable database language (such as SQL), and generate an accurate time constraint query by integrating the T0 time reference and the extended hook time variables in the data dictionary. For example, for the condition of "Not receiving insulin treatment in the 6 months before enrollment", the generated SQL code will associate the "medication_date" time field in the medication table with T0 and filter out records where "medication_date < T0-180", thereby excluding patients who have taken insulin within the time window. This code with time joins can accurately match the requirements of clinical research for time sensitivity and can avoid mis-screening of patients due to time logic oversights.

[0133] It should be noted that this embodiment can also be an improvement based on any one or more of the second to eighth embodiments.

[0134] It is not difficult to find that in the embodiment of the present application, by parsing the research plan to identify the index event and determine the starting time T0, a unified reference origin is established for the time logic. By extending the data dictionary to link the time variable field, the medical variable has a time dimension attribute, and then the time description can be converted into a relative time window to generate a pseudocode rule, and the query code containing the time connection is output. Therefore, this method realizes the precise mapping from the time conditions of the research plan to the database screening logic, which can not only ensure that the screening conditions strictly meet the time sensitivity requirements of the clinical research, but also avoid the misscreening or missed screening of patients due to omissions in the time logic through the dynamic association of time variables and the conversion of relative time windows, and ultimately provides accurate time dimension screening capabilities for clinical studies that require strict time control (such as simulated RCTs and longitudinal cohort studies).

[0135] Tenth embodiment The tenth embodiment of the present application relates to an intelligent inclusion and exclusion method based on clinical research. The tenth embodiment is an improvement on the first embodiment. Specifically, in this embodiment, the method uses a distributed server architecture to achieve privacy protection.

[0136] Specifically, in some embodiments, privacy protection is implemented using a distributed server architecture, including the following steps: Step S901: deploy a first server on the public network; the first server is used to perform large model reasoning and knowledge base management, and output encrypted discrimination codes; Step S902: deploying a second server on the intranet; the second server connects to the medical database and executes the discrimination code to generate a ranking list; Step S903: the first server transmits the identification code to the second server in a unidirectional manner through a compliant encrypted channel; In step S904, the second server only responds to the authorization request of the first server and does not transmit the original medical data externally.

[0137] Regarding step S901, the first server can, for example, run an NLP model to parse the research proposal, extract research design elements from the PICO framework, and generate discriminant code (e.g., SQL) based on local knowledge bases such as data dictionaries and medical guidelines. To ensure data privacy, the generated code is encrypted to prevent the leakage of sensitive information during transmission. The advantage of public network deployment is that it supports elastic cloud scalability, dynamically allocating resources based on large model training needs or concurrent request volume, and facilitating remote maintenance of model versions and knowledge base updates.

[0138] Regarding step S902, illustratively, the second server connects to the hospital's electronic medical record database, medical insurance database, etc. through a secure interface, supporting database software such as Oracle and MySQL. After receiving the encrypted code from the first server, the second server performs query operations in the local medical database, such as screening patient records that meet the research criteria and generating a list of patients to be included in the study. The intranet deployment adopts a strict security design, such as implementing access control such as dynamic tokens through a bastion host to ensure that only authorized requests can access. At the same time, patient data that needs to be sent out (such as personal narrative reports) is automatically desensitized to hide sensitive fields such as name and ID number.

[0139] Regarding step S903, illustratively, the identification code can be transmitted unidirectionally from the first server to the second server via a compliant encrypted channel. To meet medical data privacy requirements, the transmission channel uses two-way authenticated encryption technology (such as TLS 1.3) to ensure that the code is not intercepted or tampered with during transmission over the public network. The unidirectional transmission mechanism restricts data flow to only from the first server to the second server, preventing sensitive information in the medical database from being transmitted back to the public network. For example, the second server will not return the original data to the first server, but only the code execution results (such as a list of enrolled patient IDs), thereby eliminating the path for private data to be leaked.

[0140] Regarding step S904, for example, the second server can be specified to only respond to authorization requests from the first server, strengthening network access control. The second server's firewall rules can strictly limit the acceptance of encrypted requests only from the first server's IP address. Access from other sources (such as external hacker attacks or unauthorized internal requests) will be denied. The authorization mechanism, combined with dynamic token or API key verification, ensures that each code execution request is authenticated. For example, the first server must include a time-sensitive key to trigger code execution on the second server. This prevents unauthorized parties from forging requests to access the medical database, ensuring data access security and compliance at the architectural level.

[0141] It should be noted that this embodiment may also be an improvement based on any one or more of the second to ninth embodiments.

[0142] It is not difficult to find that in the embodiment of the present application, since the first server is deployed on the public network to perform large-model reasoning and knowledge base management and output encrypted judgment codes, and the second server is deployed on the intranet to connect to the medical database and execute code to generate a ranking list, the one-way transmission of code from the first server to the second server can be achieved through a compliant encrypted channel, and the second server is limited to only responding to the authorization request of the first server. Therefore, this distributed architecture not only utilizes the elastic computing power of the public network server to achieve large-model reasoning and code generation, but also ensures the physical isolation of medical data through the intranet server, and at the same time uses encrypted transmission and authorization mechanisms to prevent privacy leakage, thereby meeting the needs of automated screening in clinical research while achieving the dual goals of data processing and security protection.

[0143] Eleventh embodiment The eleventh embodiment of the present application relates to a specific application example of an intelligent admission and exclusion method based on clinical research.

[0144] The following is a summary of the workflow of a breast cancer simulation RCT study protocol, structured according to the logical steps and key content: 1. Extraction of research design elements 1.1 Identification of research types Input: Research protocol text (including the statement “conduct a simulated RCT study”).

[0145] Action: Identify the study type as "simulated RCT" through the large model and load the CONSORT guidelines as the prompt word context.

[0146] Output: Secondary prompt words with research type label, clearly indicating the guidelines followed by the research design.

[0147] 1.2 Extraction of research design elements Input: Contains the prompt word of the research type.

[0148] Operation: According to the requirements of simulated RCT, extract the inclusion criteria, exclusion criteria, intervention measures, clinical outcomes and other elements one by one through the large model, for example: Inclusion criteria: women aged 18-70 years, post-operative triple-negative breast cancer, ECOG score 0-1, etc.

[0149] Exclusion criteria: neoadjuvant therapy, bilateral breast cancer, metastatic lesions, etc.

[0150] Interventions: The first treatment drugs of the intervention group (PCb regimen) and the control group (CEF-T regimen).

[0151] Output: List of study design elements in standardized JSON format.

[0152] 2. Feature-variable matching and inference rule generation 2.1 Factor - Variable Label Matching Input: List of research design elements, data dictionary.

[0153] Operation: The large model achieves direct matching (such as "female" → "gender") and indirect reasoning (such as "poor blood sugar control" → "fasting blood sugar") through semantic parsing.

[0154] Output: A list of feature-variable name matches, associating medical features with specific variables in the data dictionary.

[0155] 2.2 Inference Rule Generation Input: match list, data dictionary, encoding dictionary.

[0156] Operation: Generate prompt words for each feature-variable pair (such as "How to infer breast cancer based on [diagnosis name]?"). The large model combines with the data dictionary to generate pseudocode rules, for example: Text type: "'Breast cancer' in diagnosis name | 'Breast cancer' in diagnosis name".

[0157] Categorical variable: "ICD10 code IN('C50')".

[0158] Output: Inference rules and corresponding variables in pseudocode form.

[0159] 3. Discriminant code generation 3.1 Code Generation Process Input: inference rules, SQL language, data dictionary.

[0160] operate: Database structure extraction: Get the table name, field name, type, etc. of the variable from the data dictionary (for example, "cTNM staging" corresponds to the table name "rxa_zd_zkzd").

[0161] Prompt word construction: integrating rules, structure and language. Example: "Generate SQL code to screen patients with cTNM stage I to IIC".

[0162] Large model code generation: For example, using CASE WHEN statements to determine stage values ​​and group and filter patient IDs.

[0163] Output: SQL discrimination code corresponding to each factor, such as the code for screening early breast cancer: SELECT patient_id, MAX(CASE WHEN "ZD-03-004" IN ('I', 'IA', ..., 'IIC') THEN 1 ELSE 0 END) AS res4FROM rxa_zd_zkzd GROUP BY patient_id; 4. Identify code execution and result collection Input: identification code set, database port.

[0164] Action: Execute the code in the database to generate Boolean judgment results (TRUE / FALSE / NA) for each patient-element.

[0165] Output: discriminant result matrix, rows are patient IDs, columns are feature results (e.g. patient_1 is FALSE under the "early breast cancer" feature).

[0166] 5. Confirmation of the list of enrolled patients Input: Discriminant result matrix.

[0167] operate: Define aggregation functions: combine elements using Boolean logic, for example: Inclusion criteria: resPin=res1&res2&…&res11 (all met simultaneously).

[0168] Exclusion criteria: resPout=!res12&…&!res20 (none of them are met).

[0169] Calculate the inclusion and exclusion conclusion: RES=resPin&resPout&resIC&resO, and screen patients with RES=TRUE.

[0170] Output: A list of patient IDs that meet the criteria (e.g. [2,3,5,...,1000]).

[0171] 6. Expert review and feedback optimization (optional) 6.1 Patient Narrative Data Generation Input: group list, database port.

[0172] Operation: Integrate patient demographic, diagnosis, pathology, and other data by timeline to generate a structured report (such as patient_1's diagnosis time, examination results, etc.).

[0173] Output: Personal narrative data with timestamp.

[0174] 6.2 Expert Review and Error Marking Action: Experts review narrative data and mark error types (e.g., incomplete inference rules, missing keyword matches), for example: Patient 1 was not identified as having breast cancer due to the diagnosis of “breast cancer”, so the keyword “breast cancer” needs to be added.

[0175] Output: Annotated dataset (patient ID, error elements, correction suggestions).

[0176] 6.3 Model Iteration operate: Rule modification: Updated feature-variable matching (e.g., adding the keyword “breast cancer” to “diagnosed with breast cancer”).

[0177] Code update: Regenerate the discrimination code to include the revised rules.

[0178] Fine-tuning large models: Using expert annotations as groundtruth to optimize model reasoning capabilities.

[0179] 7. Historical archiving and knowledge reuse (optional) 7.1 Historical Archive Construction Content: Stores research elements, inference rules, code, audit records, etc. The example JSON archive contains project ID, study type, and revision annotation (such as "Diagnosis adds the keyword 'breast cancer'").

[0180] 7.2RAG Retrieval and Reuse Operation: New studies (such as the cervical cancer simulation RCT) are retrieved from historical archives through semantic vectors. Codes are directly reused for fully matching elements (such as "18-70 years old"), and similar elements (such as "diagnosed with cervical cancer") are adapted after adjusting keywords.

[0181] 8. Time logic enhancement (optional) 8.1 Definition of Starting Event Input: Research protocol (including "postoperative pathological examination").

[0182] Operation: Identify the starting event as "breast cancer surgery" and define the starting time T0.

[0183] Output: Extended research elements (including T0).

[0184] 8.2 Data Dictionary Extension Action: Add a linked time field to the variable (e.g., "Pathology Report Date" as the timestamp of "Histological Type").

[0185] 8.3 Temporal Conditional Embedding Operation: Convert the time description into a relative window (e.g., "blood test within 14 days after surgery" → T0 ≤ test date ≤ T0+14), and generate SQL code with time constraints (e.g., associate the T0 form to filter patients).

[0186] In summary, this application is driven by a large model to achieve a fully automated closed loop from research plan text to patient screening. It combines data dictionaries, historical archives and expert feedback to ensure the accuracy and reusability of the screening logic. It is especially suitable for clinical research scenarios that require strict time control and privacy protection.

[0187] It can be seen that this application has at least the following technical effects: 1. Fully automated patient screening, no programming background required: Leveraging a large language model, the system intelligently parses study proposal text, automatically extracts structured study design elements, and generates discriminant code that can be run directly in the database. Medical researchers, without programming skills, simply submit a study proposal described in natural language to directly obtain a list of eligible patients, effectively eliminating barriers to cross-disciplinary collaboration.

[0188] 2. Unified processing of multi-source heterogeneous data: Data dictionaries and RAG technology enable semantic matching, automatically mapping variable names, table structures, and encoding rules across different databases. This supports the integration of heterogeneous databases across institutions and multi-center studies, significantly reducing the labor cost of manual database adaptation and significantly improving the utilization efficiency of multi-source data.

[0189] 3. Dynamic medical knowledge integration and continuous optimization: Local knowledge bases can integrate medical guidelines, literature, and historical archives, leveraging RAG technology to retrieve the latest knowledge in real time. Furthermore, expert review and feedback mechanisms can drive model iteration and updates. This ensures that inclusion and exclusion rules remain synchronized with medical frontiers and researchers' specific requirements, preventing rule degradation or misinterpretation of research objectives.

[0190] 4. Improved efficiency and accuracy: Combining technologies such as broad-spectrum medical knowledge retrieval, automated code generation, feedback verification mechanisms, and precise processing of time-sensitive conditions, this system enables patient screening in large multi-center databases to shift from the traditional manual docking and repeated communication model to an intelligent human-computer interaction model, balancing efficiency and accuracy.

[0191] 5. Strict Data Privacy and Security: A distributed deployment architecture separates the networked large model from the intranet database, ensuring data security through encrypted transmission. The large model never accesses real patient data, and sensitive information is completely isolated within the intranet, strictly adhering to data privacy and security requirements.

[0192] The step division of the above various methods is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0193] In addition, some embodiments of the present application further provide an electronic device. The electronic device may be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device may also be various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0194] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, enable the processor to perform the steps of the method provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing some of the necessary operations. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0195] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means, with the bus connection being used as an example in the figure.

[0196] Input device 1103 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. Examples include a touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, and other input devices. Output device 1104 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). The display device may include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0197] To provide user interaction, the electronic device may be a computer. The computer includes a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse) through which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback), and input from the user may be received in any form (e.g., voice input or tactile input).

[0198] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium. When executed by a processor, the computer program / instruction implements the steps of the method provided in any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments, or it may exist independently and not be incorporated into the device. The computer-readable medium carries one or more computer-readable instructions.

[0199] The memory 1102 can be used as a non-transitory computer-readable storage medium to store non-transitory software programs, non-transitory computer executable programs, and modules. The processor 1101 executes the non-transitory software programs, instructions, and modules stored in the memory 1102 to execute various functional applications and data processing of the server, thereby implementing the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.

[0200] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely located relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0201] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. Computer-readable media may be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.

[0202] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, compact discs, digital versatile discs or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0203] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0204] In the above embodiments, all or part of the steps or functions of the present invention may be implemented using software, hardware, firmware, or any combination thereof. For example, implementation may be achieved using a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present application may be implemented using hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.

[0205] The computer program product provided in the embodiments of the present application includes one or more computer programs / instructions that, when executed by a processor, fully or partially produce the processes or functions described in accordance with the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).

[0206] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-specific system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0207] The scope of this application is defined by the appended claims rather than the foregoing description and is therefore intended to encompass within this application all changes that come within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to which they relate. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim may also be implemented by one unit or device through software or hardware. Words such as "first" and "second" are only used to distinguish the description and do not indicate any particular order, nor should they be understood as indicating or implying relative importance.

[0208] The above descriptions are merely specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art may easily propose variations or substitutions within the technical scope disclosed in the present application, and such variations or substitutions shall be encompassed within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims, and the above descriptions shall be regarded as exemplary and non-limiting.

Claims

1. An intelligent admission and exclusion method based on clinical research, characterized in that: The method is implemented based on a large model and includes: Determine the research design elements based on the received research proposal; Determine inference rules based on the research design elements described; According to the inference rules, a discriminant code for each research design element is generated; According to the identification code, the list of patients who meet the research plan is determined.

2. The method according to claim 1, characterized in that According to the received research plan, the research design elements include: Identify the research type of the research protocol being described; Obtain international standard guidelines and / or user-defined prompt word templates corresponding to the research type; Determining element extraction prompt words according to the research plan, the research type, the international standard guidelines, and / or a user-defined prompt word template; Prompt words are extracted based on the elements, and a structured list of research design elements is output through the large model.

3. The method according to claim 2, characterized in that The extracting prompt words according to the elements and outputting a structured list of research design elements through the large model includes: Determine the target items of the research design elements according to the PICO framework of the research type; Extracting prompt words based on the elements, querying the large model item by item for the research design elements of the target items in the research plan; The structured list is formed according to the queried research design elements.

4. The method according to claim 3, characterized in that The structured list formed based on the queried research design elements includes: Performing a conversion operation on the retrieved research design elements to generate specific conditions that can be interpreted by data; the conversion operation includes at least one of the following: decomposing complex conditions into atomic conditions; converting the description of the intervention plan into an actionable medical behavior; converting the calculation requirements of the outcome indicator into the existence requirements of the original test data at a preset time node; The structured list is formed according to the specific conditions.

5. The method according to claim 1, characterized in that Determining the inference rules based on the research design elements includes: Semantically matching the research design elements with the data dictionary to generate element-variable mapping relationships; An inference rule is determined based on the element-variable mapping relationship.

6. The method according to claim 5, characterized in that The semantic matching of the research design elements with the data dictionary to generate element-variable mapping relationships includes: Parsing semantic information of the research design elements through a large model; The semantic information is semantically matched with the variable labels in the data dictionary, and a matching list of elements and variable names including the variable name, variable label, and the form to which it belongs is output.

7. The method according to claim 5, characterized in that Determining the inference rule according to the element-variable mapping relationship includes: Constructing and filling a prompt word template for each generated element-variable mapping relationship; wherein the prompt word template is injected with the following information: the medical definition of the target element, the matching variable name, the variable label and type, and the encoding rules of the data dictionary; The filled prompt words are input into the large model, and the inference rules in pseudocode form are output in combination with the data dictionary; the inference rules include at least one of the following: threshold judgment of numerical variables, keyword matching of text variables, and code value mapping of categorical variables.

8. The method according to claim 1, characterized in that Generating the discriminant code of each research design element according to the inference rule includes: Extracting structured metadata of the variables involved in the inference rule from a data dictionary; the metadata includes: the form name to which the variable belongs, the variable name, the variable data type, and the variable value range definition; generating code generation prompt words according to the inference rule, the structured metadata, and the target language; The code generation prompt words are converted into discriminant codes through the large model.

9. The method according to claim 1, characterized in that The patient list determined to be in compliance with the research plan according to the discrimination code includes: Executing the discrimination code to obtain discrimination results of each research design element; According to the discrimination results, a list of patients who meet the research plan is determined.

10. The method according to claim 9, characterized in that The executing of the discrimination code to obtain the discrimination results of each research design element includes: Execute the discrimination code in the target database to obtain an original discrimination result set; Generate a discrimination result matrix based on the original discrimination result set; wherein: the row index of the discrimination result matrix corresponds to the patient identifier, the column index corresponds to the study design element, and the matrix cell value stores the discrimination result of each patient-element combination; The discrimination result matrix is ​​used as the discrimination result of each research design element.

11. The method according to claim 9, characterized in that According to the discrimination results, the list of patients who meet the research plan includes: Determine the summary function based on the research plan; Input the discrimination results into the summary function to calculate the inclusion and exclusion conclusion of each patient; According to the inclusion and exclusion conclusions, the list of patients who meet the research plan was determined.

12. The method according to claim 1, characterized in that After determining the list of patients who meet the research plan according to the discrimination code, the method further includes: Based on the patient list, time-invariant information and time-series medical data are extracted from the database and integrated along the timeline to generate a structured personal narrative report; receiving the admission and exclusion annotation data generated by the medical expert based on the narrative report; Based on the entry and exit annotation data, the element-variable mapping relationship, the inference rule, the discrimination code and the large model are updated.

13. The method according to claim 1, wherein The method further comprises: Store validated study design elements, inference rules, discriminant codes, and associated metadata as a structured historical archive; In response to the input of a new research plan, the following operations are performed: the current research elements and inference rules are encoded into semantic vectors; the Top-K similar archived entries are retrieved in the historical knowledge base; if there is a fully matching entry, the inference rules and discrimination codes are directly reused; if there is a partially matching entry, an automatic adaptation operation is performed to generate adapted inference rules and discrimination codes.

14. The method according to claim 1, wherein The method further comprises: Parse the study protocol to identify the index event and determine the start time of the individual patient study; Expand the linked time variable field in the data dictionary to dynamically associate medical variables with timestamps; Convert the time-related descriptions in the research plan into time windows relative to the start time, and generate inference rules in pseudocode form; Output the database query code containing the time link, ensuring that the filter conditions correspond to the start time and hook time variables.

15. The method according to claim 1, wherein The method uses a distributed server architecture to achieve privacy protection: Deploy a first server on the public network; the first server is used to perform large-model reasoning and knowledge base management, and output encrypted discriminant codes; Deploy a second server on the intranet; the second server connects to the medical database and executes the discrimination code to generate a ranking list; The first server unidirectionally transmits the identification code to the second server through a compliant encrypted channel; The second server only responds to the authorization request of the first server and does not transmit the original medical data to the outside.

16. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method according to any one of claims 1 to 15.

17. A computer readable medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

18. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

Citation Information

Patent Citations

  • Identification of candidates for clinical trials

    CN105940427A

  • Method for checking testee recruitment conditions for clinical researches

    CN106815360A

  • Clinical test patient matching method

    CN110223784A

  • Clinical test item recommendation method and device, electronic equipment and storage medium

    CN115878893A

  • Method and device for automatically formulating entering and ranking standards, electronic equipment and storage medium

    CN118888069A

Cited By

  • Clinical research information batch extraction system and extraction method

    CN122117190A