Patient enrollment model generation method, patient enrollment method, equipment and media

By generating a large model for target patient enrollment and using multiple models to process clinical data, the problems of low patient enrollment efficiency and poor quality were solved, achieving efficient and accurate enrollment processing.

CN117747037BActive Publication Date: 2026-04-03BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Current technologies suffer from low patient enrollment efficiency and poor quality, which can easily lead to misjudgments.

Method used

By acquiring a training sample set, semantic analysis and inclusion condition prompts are performed on clinical data using the first and second pre-trained models. The error rate is calculated and the confidence level is adjusted to generate a large model for the inclusion of target patients.

Benefits of technology

This improved patient enrollment efficiency, reduced misdiagnosis, and enhanced enrollment quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117747037B_ABST
    Figure CN117747037B_ABST
Patent Text Reader

Abstract

This application discloses a method for generating a large-scale patient enrollment model, a patient enrollment method, equipment, and medium. A first pre-trained model and a second pre-trained model process clinical data from a first training sample to obtain first and second predicted enrollment information for the first training sample. Then, using the first and second predicted enrollment information and the actual enrollment information of the first training sample, the error rates of the first and second pre-trained models are calculated. The preset confidence level is then adjusted based on the error rates until the first and second pre-trained models respectively meet their corresponding preset conditions, resulting in a target large-scale patient enrollment model and target confidence level for patient enrollment. This method can improve patient enrollment efficiency, avoid misjudgments, and improve enrollment quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer science, specifically to a method for generating a large patient enrollment model, a patient enrollment method, equipment, and media. Background Technology

[0002] Patient enrollment has a significant impact on disease research. However, current methods often rely on existing patient data and past experience for patient enrollment, which not only leads to low enrollment efficiency but also increases the likelihood of misjudgment, resulting in poor patient enrollment quality. Summary of the Invention

[0003] To address the aforementioned issues, this application proposes a method for generating a large-scale patient enrollment model, comprising:

[0004] Obtain a training sample set, wherein each first training sample in the training sample set includes clinical data of the sample patients and actual enrollment information;

[0005] The clinical data are processed by the first pre-trained model and the second pre-trained model respectively to determine the first predicted enrollment information and the second predicted enrollment information for each of the first training samples.

[0006] The first pre-trained model is used to perform semantic analysis on the clinical data to determine the first predicted enrollment information of the first training sample; the second pre-trained model is used to determine the second predicted enrollment information of the first training sample through preset enrollment condition prompts.

[0007] Based on the first predicted entry information and the second predicted entry information of the first training sample, as well as the corresponding real entry information, the error rates of the first pre-trained model and the second pre-trained model are calculated respectively.

[0008] The preset credibility of the first pre-trained model and the second pre-trained model is adjusted according to the error rate until the first pre-trained model and the second pre-trained model respectively meet the corresponding preset conditions, so as to obtain the target patient enrollment large model and the target credibility; the target patient enrollment large model is composed of the first pre-trained model and the second pre-trained model that meet the corresponding preset conditions.

[0009] In one example, obtaining the training sample set includes:

[0010] By connecting with the hospital's medical system through relevant network interfaces, the initial medical data of each sample patient in the hospital system can be obtained.

[0011] The initial medical data is subjected to a data quality assessment, and the initial medical data that does not meet the preset quality conditions is filtered out to obtain preprocessed medical data;

[0012] Descriptive statistical analysis was performed on the preprocessed medical data to obtain the clinical data of the sample patients;

[0013] The first training sample is formed based on the patient's clinical data and the corresponding actual enrollment information.

[0014] In one example, the method further includes:

[0015] For each set of initial medical data, generate an association between the initial medical data and the corresponding clinical data;

[0016] The initial medical data, the corresponding clinical data, and the associated relationships are stored in a preset database.

[0017] In one example, the method further includes:

[0018] A second training sample is obtained from a preset database, the second training sample consisting of the initial medical data and the corresponding clinical data;

[0019] A model is generated based on the second training sample to obtain a data preprocessing model for cleaning the initial medical data.

[0020] In one example, adjusting the preset confidence levels of the first pre-trained model and the second pre-trained model according to the error rate until the first pre-trained model and the second pre-trained model respectively meet the corresponding preset conditions to obtain the target patient enrollment large model includes:

[0021] The preset confidence level is adjusted based on the error rate;

[0022] For each of the first training samples, the joint probability of the clinical data is calculated based on the adjusted preset confidence level, the first predicted enrollment information of the first training sample, and the second predicted enrollment information, and the target group of the clinical data is determined based on the joint probability.

[0023] Wherein, the first predicted enrollment information and the second predicted enrollment information are used to represent at least one predicted group of the clinical data and the predicted probability corresponding to the predicted group.

[0024] Based on the target group and the real group of clinical data, the model parameters of the first pre-trained model and the second pre-trained model are adjusted until the first pre-trained model and the second pre-trained model meet the corresponding preset conditions, thereby obtaining the target patient enrollment model.

[0025] In one example, for each of the first training samples, calculating the joint probability of the clinical data based on the adjusted preset confidence level, the first predicted inclusion information of the first training sample, and the second predicted inclusion information, and determining the target group of the clinical data based on the joint probability, includes:

[0026] For each of the aforementioned clinical data, the same predicted group is obtained based on the first predicted enrollment information and the second predicted enrollment information.

[0027] Calculate the joint probability of the same prediction group based on the first and second prediction probabilities of the same prediction group and the adjusted preset confidence level;

[0028] The identical prediction groups are sorted according to the joint probability, and the identical prediction with the highest joint probability is taken as the target group of the clinical data.

[0029] On the other hand, this application also proposes a patient enrollment method, including:

[0030] Obtain clinical data from patients awaiting treatment;

[0031] The clinical data of the patients to be treated are processed using the target patient enrollment model to obtain the first and second prediction group information of the clinical data; the target patient enrollment model is obtained by the patient enrollment model generation method described above.

[0032] The target group of the clinical data is determined based on the information from the first prediction group and the second prediction group, as well as the target confidence level.

[0033] In one example, determining the target group of the clinical data based on the first prediction group information, the second prediction group information, and the target confidence level includes:

[0034] Based on the first and second predicted group information, obtain the same predicted group;

[0035] Calculate the joint probability of the same prediction group based on the first and second prediction probabilities of the same prediction group and the target confidence level;

[0036] The identical prediction groups are sorted according to the joint probability, and the identical prediction with the highest joint probability is taken as the target group of the clinical data.

[0037] On the other hand, this application also proposes an electronic device, comprising:

[0038] At least one processor; and,

[0039] A memory communicatively connected to the at least one processor; wherein,

[0040] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: perform the patient enrollment large model generation method as described in any of the preceding claims, and / or the patient enrollment method as described in any of the preceding claims.

[0041] On the other hand, this application also proposes a non-volatile computer storage medium storing computer-executable instructions, characterized in that the computer-executable instructions are configured as: a patient enrollment large model generation method as described in any of the preceding claims, and / or a patient enrollment method as described in any of the preceding claims.

[0042] The patient enrollment large-scale model training method, patient enrollment method, electronic equipment, and non-volatile computer storage medium proposed in this application can bring the following beneficial effects:

[0043] The first pre-trained model and the second pre-trained model process the clinical data in the first training sample respectively to obtain the first predicted enrollment information and the second predicted enrollment information of the first training sample. Then, the error rate of the first pre-trained model and the second pre-trained model is calculated using the first predicted enrollment information, the second predicted enrollment information and the actual enrollment information of the first training sample. The preset confidence level is then adjusted according to the error rate until the first pre-trained model and the second pre-trained model respectively meet the corresponding preset conditions. This yields the target patient enrollment model and the target confidence level for patient enrollment, which can improve patient enrollment efficiency, avoid misjudgment, and improve enrollment quality. Attached Figure Description

[0044] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0045] Figure 1 This is a flowchart illustrating the method for generating a large model for patient enrollment in an embodiment of this application.

[0046] Figure 2 This is a flowchart illustrating the patient enrollment method in an embodiment of this application;

[0047] Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0049] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0050] like Figure 1 As shown in the embodiments of this application, the method for generating a large patient enrollment model includes:

[0051] S101: Obtain the training sample set.

[0052] Each first training sample in the training sample set consists of clinical data from sample patients and real enrollment information.

[0053] The aforementioned real enrollment information includes the real group corresponding to the clinical data of the sample patients, for example, the real group is community-acquired pneumonia; or for example, adult-acquired pneumonia, etc.

[0054] It is understood that the real groups in the first training sample mentioned above can be manually labeled, and this application embodiment does not impose any specific limitations.

[0055] In some embodiments of this application, step S101 above: obtaining the training sample set may specifically include the following steps:

[0056] 1. Connect with the hospital's medical system through relevant network interfaces to obtain the initial medical data of each sample patient in the hospital system;

[0057] 2. Conduct a data quality assessment on the initial medical data, and filter out the initial medical data that does not meet the preset quality conditions to obtain preprocessed medical data;

[0058] 3. Perform descriptive statistical analysis on the preprocessed medical data to obtain the clinical data of the sample patients;

[0059] 4. Based on the patients' clinical data and the corresponding real enrollment information, the first training sample is formed.

[0060] The aforementioned data quality assessment includes checking for missing values, outliers, and duplicate values, with the corresponding preset quality condition being the absence of missing values, outliers, and duplicate values.

[0061] The above descriptive statistical analysis specifically uses basic data indicators such as mean, median, standard deviation, minimum and maximum values, and refers to distribution characteristics such as skewness, kurtosis and quantiles to provide an intuitive and comprehensive interpretation of the data.

[0062] In real-world business scenarios, the medical data of patients from different hospitals are largely heterogeneous. Therefore, in order to better enroll patients, this embodiment of the application cleans and transforms the medical data of sample patients through data quality assessment and descriptive statistical analysis to obtain unified clinical data and improve the quality of patient enrollment.

[0063] Understandably, during the patient enrollment process, the initial medical data of patients can be preprocessed using the data cleaning module (which is used for data quality assessment and descriptive statistical analysis) to improve the quality of patient enrollment.

[0064] In some embodiments of this application, the patient enrollment large model generation method provided in this application may further include:

[0065] For each initial medical data point, generate the association between the initial medical data and the corresponding clinical data;

[0066] Initial medical data, corresponding clinical data, and related relationships are stored in a pre-defined database.

[0067] Furthermore, the patient enrollment large model generation method provided in this application may also include:

[0068] A second training sample is obtained from a preset database. The second training sample consists of initial medical data and corresponding clinical data.

[0069] The model is trained based on the second training sample to obtain a data preprocessing model for cleaning the initial medical data.

[0070] The data preprocessing model mentioned above is a neural network model.

[0071] In this embodiment, a corpus storage module can be pre-configured to store the initial medical data, corresponding clinical data, and related relationships in a preset database. Furthermore, a second training sample can be obtained from the preset database; this second training sample consists of the initial medical data and corresponding clinical data. Then, the model is trained based on the second training sample to obtain a data preprocessing model, which is then stored. This data preprocessing model can clean the initial medical data to improve data cleaning efficiency, thereby further improving the efficiency of generating the large-scale patient enrollment model and the patient enrollment efficiency.

[0072] S102: The clinical data are processed by the first pre-trained model and the second pre-trained model respectively to determine the first predicted enrollment information and the second predicted enrollment information for each first training sample.

[0073] The first pre-trained model is used to perform semantic analysis on clinical data to determine the first predicted enrollment information of the first training sample; the second pre-trained model is used to determine the second predicted enrollment information of the first training sample through preset enrollment condition prompts.

[0074] The aforementioned first prediction grouping information includes at least one first prediction group and the prediction probability of the first prediction group; similarly, the aforementioned second prediction grouping information includes at least one second prediction group and the prediction probability of the second prediction group.

[0075] In this embodiment, the clinical data of the first training sample is input into the first pre-training model and the second pre-training model respectively, and the enrollment status of the clinical data is predicted from two aspects: semantic analysis and enrollment conditions, so as to obtain the first predicted enrollment information and the second predicted enrollment information of each first training sample.

[0076] In this embodiment, a semantic analysis module and a rule extraction module can be pre-set. The semantic analysis module stores a first pre-trained model, and the rule extraction module stores a second pre-trained model.

[0077] Furthermore, clinical data is input into the initial neural network for model training. The self-learning characteristics and attention mechanism of the initial neural network are used to obtain the relevant knowledge architecture to obtain the first pre-trained model. Based on the first pre-trained model, the clinical data is analyzed and predicted to obtain the set A of possible patient enrollment situations and to give the corresponding result weights, i.e., the first predicted enrollment information.

[0078] In some embodiments of this application, a second pre-trained model can be constructed based on accumulated medical knowledge. This model performs operations such as thelexical analysis, code matching, logical matching, regular expression matching, and contextual semantic verification for the corresponding disease, resulting in a set B of possible patient inclusion cases. The weight of the results is then evenly distributed across each case to obtain the second predicted inclusion information. Since medical knowledge continuously iterates and optimizes over time, this knowledge is also applied as a text reading semantic prompt to the semantic analysis module. The current prediction prompt includes the relevant problem scenario, filtering conditions, and target result. For example, it specifies the location of the prediction within the medical record, the keywords to be included, the logical meaning between the keywords, and whether a numerical type, string, or complex structure is required.

[0079] S103: Calculate the error rates of the first pre-trained model and the second pre-trained model based on the first predicted entry information and the second predicted entry information of the first training sample, as well as the corresponding real entry information.

[0080] In this embodiment of the application, the error rates of the first pre-trained model and the second pre-trained model are calculated based on the number of first training samples, the first predicted entry information and the second predicted entry information of the first training samples, and the corresponding actual entry information.

[0081] For example, suppose that out of 500 patients, 100 are enrolled using the first pre-trained model with 10 manual error corrections, and the other 400 are enrolled using the second pre-trained model with 20 manual error corrections. Then, the error rate of the first pre-trained model is 10 / 100 = 0.1. Similarly, the error rate of the second pre-trained model is 20 / 400 = 0.05.

[0082] It is understandable that "entering the group through the first pre-trained model" means that the first training sample was successfully entered into the group through the first pre-trained model, that is, the predicted group obtained by the first pre-trained model matches the real group. Similarly, "entering the group through the second pre-trained model" means that the first training sample was successfully entered into the group through the second pre-trained model, that is, the predicted group obtained by the second pre-trained model matches the real group.

[0083] The number of manual corrections mentioned above refers to the number of first training samples whose predicted group and the real group do not match, obtained through the first pre-trained model or the second pre-trained model.

[0084] S104: Adjust the preset confidence levels of the first and second pre-trained models based on the error rate until the first and second pre-trained models respectively meet the corresponding preset conditions, thereby obtaining the target patient enrollment model and the target confidence level.

[0085] The target patient enrollment model consists of a first pre-trained model and a second pre-trained model that meet corresponding preset conditions. The aforementioned target confidence level is the adjusted preset confidence level when the first and second pre-trained models respectively meet their corresponding preset conditions.

[0086] In some embodiments of this application, the above-mentioned adjustment of the preset confidence levels of the first and second pre-trained models based on the error rate until the first and second pre-trained models respectively meet the corresponding preset conditions to obtain the target patient enrollment large model can be achieved through the following steps:

[0087] 1. Adjust the preset confidence level based on the error rate;

[0088] 2. For each training sample, calculate the joint probability of the clinical data based on the adjusted preset confidence level, the first predicted enrollment information and the second predicted enrollment information of the training sample, and determine the target group of the clinical data based on the joint probability.

[0089] Wherein, the first predicted enrollment information and the second predicted enrollment information are used to represent at least one predicted group of clinical data and the predicted probability corresponding to the predicted group.

[0090] 3. Based on the target group and the real group of clinical data, adjust the model parameters of the first pre-training model and the second pre-training model until the first pre-training model and the second pre-training model meet the corresponding preset conditions, and obtain the target patient enrollment model and target confidence.

[0091] The aforementioned adjustment of the preset confidence level based on the error rate can refer to subtracting the error rate from the preset confidence level. In the embodiments of this application, the preset confidence levels of both the first pre-trained model and the second pre-trained model are 0.5.

[0092] The aforementioned preset conditions are: reaching a preset number of training iterations, or the error rate being less than a preset threshold, etc., but are not specifically limited in this embodiment.

[0093] For example, taking the simulation to confirm the enrollment of CAP (community-acquired pneumonia) as an example, the method for generating the above-mentioned large-scale patient enrollment model is explained in detail:

[0094] The first step involves assuming 100,000 medical records containing admission and related outpatient data (i.e., the initial medical data of the sample patients). After analysis by the data cleaning module, duplicate and outlier values ​​are removed, leaving 80,000 relevant records. These records are then stored in the corpus storage module and pre-trained to obtain associated knowledge. The second step involves waiting for the ETL frontend to receive new patient data (let's call it 1001) before it enters the core module for decision analysis. First, the second pre-trained model in the rule extraction module directly utilizes existing patterns: admission to the pediatrics or respiratory and critical care medicine department; primary diagnosis on the admission record being pneumonia, pulmonary infection, or respiratory infection; and age between 2 and 18 years (inclusive). Under these conditions, the patient is considered for inclusion. However, according to the rules, two other diseases may also be matched sequentially: CAP-ADULT (adult-acquired pneumonia) and STEMI (acute ST-segment elevation myocardial infarction). For this patient (1001), the probability of each disease being matched is P=0.33, and the initial result confidence is P(confidence|rule)=0.5. Therefore, the joint probability of each possible outcome is P(0.5*0.33)=0.165. Second, the rule extraction module visualizes and encodes the conditions as follows:

[0095] Condition 1: The admitting department is pediatrics or respiratory and critical care medicine.

[0096] Condition 2: The primary diagnosis in the admission record is pneumonia, lung infection, or respiratory tract infection.

[0097] Condition 3: Age range [2, 18)

[0098] Conditional relationships: Condition 1 & Condition 2 & Condition 3

[0099] If the above conditions are met, output Y; otherwise, output N.

[0100] Other statistics: probability of outcome P

[0101] Similarly, other inclusion rules will also perform text prediction according to the first pre-trained model in the semantic analysis module, and record possible results and scores. Assuming the probabilities of predicting two diseases (CAP, community-acquired pneumonia) are P(CAP) = 0.8 and P(CAP-ADULT) = 0.6, and the initial confidence level is P(confidence|module) = 0.5, then the final probabilities are simply P(CAP) = 0.4 and P(CAP-ADULT) = 0.3. The probability matrix after summing the above two modules is as follows:

[0102]

[0103] The third step is to input the probability matrix from the second step into the cross-validation module. The final enrollment result for patient 1001 is the maximum probabilities CAP, the enrollment method is the first pre-trained model, and the probability is P(CAP) = 0.4.

[0104] The fourth step involves inputting the results and statistical information from the third step into the data warehouse module. After users periodically check and correct errors on the interface, feedback results are recorded. When the sample size reaches 500, the error rates of the rule-based and semantic modules are statistically analyzed. Assuming that out of 500 patients, 100 are enrolled through rule-based methods with 10 manual corrections, and the other 400 are enrolled through semantic methods with 20 manual corrections, then the error rate of the rule-based module is 10 / 100 = 0.1. Similarly, the error rate of the semantic module is 20 / 400 = 0.05. These are converted into feedback factors: Factor(rule) = 0.1, Factor(module) = 0.05.

[0105] Fifth, when incremental patient data enters the core module, the final confidence level calculated by the rules and semantic modules is the initial confidence level plus or minus a feedback factor. An increase or decrease in the month-on-month error rate corresponds to an adjustment in the confidence level. This process is repeated iteratively with the data to obtain the target patient enrollment model and the target confidence level.

[0106] The patient enrollment large-scale model generation method provided in this application processes the clinical data in the training samples using a first pre-trained model and a second pre-trained model respectively, to obtain the first predicted enrollment information and the second predicted enrollment information of the first training sample. Then, using the first predicted enrollment information, the second predicted enrollment information, and the actual enrollment information of the training sample, the error rate of the first pre-trained model and the second pre-trained model is calculated. Next, the preset confidence level is adjusted according to the error rate until the first pre-trained model and the second pre-trained model respectively meet the corresponding preset conditions, thereby obtaining the target patient enrollment large-scale model and the target confidence level for patient enrollment. This method can improve patient enrollment efficiency, avoid misjudgment, and improve enrollment quality.

[0107] like Figure 2 As shown, this application also proposes a patient enrollment method, including:

[0108] S201: Obtain clinical data of patients to be treated.

[0109] S202: Using the target patient enrollment model, the clinical data of the patients to be treated are processed to obtain the first and second prediction group information of the clinical data.

[0110] The target patient enrollment model was obtained using the patient enrollment model generation method described in any of the preceding items;

[0111] S203: Determine the target group for clinical data based on the information from the first and second prediction groups, as well as the target confidence level.

[0112] In some embodiments of this application, determining the target group of clinical data based on the first prediction group information, the second prediction group information, and the target confidence level includes:

[0113] Based on the first and second predicted group information, obtain the same predicted group;

[0114] Calculate the joint probability of the same prediction group based on the first and second prediction probabilities of the same prediction group and the target confidence level;

[0115] The same prediction groups are sorted according to their joint probability, and the same prediction group with the highest joint probability is taken as the target group of the clinical data.

[0116] The patient enrollment model generation method provided in this application processes the clinical data of the patients to be treated using the first target model and the second target model in the target patient enrollment model to obtain the first predicted enrollment information and the second predicted enrollment information of the patients to be treated. Then, the target group of the patients to be enrolled is determined by the predicted groups in the first predicted enrollment information and the second predicted enrollment information, the predicted probabilities of each predicted group, and the target confidence. This embodiment uses the first target model and the second target model to process the patients to be treated, which can improve the efficiency of patient enrollment, avoid misjudgment, and improve the quality of enrollment.

[0117] like Figure 3 As shown, this application also proposes an electronic device, comprising:

[0118] At least one processor; and,

[0119] A memory communicatively connected to the at least one processor; wherein,

[0120] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: perform the patient enrollment large model generation method as described in any of the preceding claims, and / or the patient enrollment method as described in any of the preceding claims.

[0121] This application also proposes a non-volatile computer storage medium storing computer-executable instructions configured as: a patient enrollment large model generation method as described in any of the preceding claims, and / or a patient enrollment method as described in any of the preceding claims.

[0122] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0123] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0124] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0125] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0128] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0129] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0130] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0131] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0132] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating a large patient enrollment model, characterized in that, include: Obtain a training sample set, wherein each first training sample in the training sample set includes clinical data of the sample patients and actual enrollment information; The clinical data are processed by the first pre-trained model and the second pre-trained model respectively to determine the first predicted enrollment information and the second predicted enrollment information for each of the first training samples. The first pre-trained model is used to perform semantic analysis on the clinical data to determine the first predicted enrollment information of the first training sample; The second pre-trained model is used to determine the second predicted inclusion information of the first training sample by using preset inclusion condition prompts. Based on the first predicted entry information and the second predicted entry information of the first training sample, as well as the corresponding real entry information, the error rates of the first pre-trained model and the second pre-trained model are calculated respectively. The preset confidence levels of the first pre-trained model and the second pre-trained model are adjusted according to the error rate until the first pre-trained model and the second pre-trained model respectively meet the corresponding preset conditions, thereby obtaining the target patient enrollment model and the target confidence level; the target patient enrollment model is composed of the first pre-trained model and the second pre-trained model that meet the corresponding preset conditions. The step of adjusting the preset confidence levels of the first pre-trained model and the second pre-trained model according to the error rate until the first pre-trained model and the second pre-trained model respectively meet the corresponding preset conditions to obtain a target patient enrollment model includes: adjusting the preset confidence level according to the error rate; for each first training sample, calculating the joint probability of the clinical data according to the adjusted preset confidence level, the first predicted enrollment information and the second predicted enrollment information of the first training sample, and determining the target group of the clinical data according to the joint probability; wherein the first predicted enrollment information and the second predicted enrollment information are used to represent at least one predicted group of the clinical data and the predicted probability corresponding to the predicted group; and adjusting the model parameters of the first pre-trained model and the second pre-trained model according to the target group and the real group of the clinical data until the first pre-trained model and the second pre-trained model meet the corresponding preset conditions to obtain a target patient enrollment model.

2. The method according to claim 1, characterized in that, Obtain the first training sample set, including: By connecting with the hospital's medical system through relevant network interfaces, the initial medical data of each sample patient in the medical system can be obtained. The initial medical data is subjected to a data quality assessment, and the initial medical data that does not meet the preset quality conditions is filtered out to obtain preprocessed medical data; Descriptive statistical analysis was performed on the preprocessed medical data to obtain the clinical data of the sample patients; The first training sample is formed based on the patient's clinical data and the corresponding actual enrollment information.

3. The method according to claim 2, characterized in that, The method further includes: For each set of initial medical data, generate an association between the initial medical data and the corresponding clinical data; The initial medical data, the corresponding clinical data, and the associated relationships are stored in a preset database.

4. The method according to claim 3, characterized in that, The method further includes: A second training sample is obtained from a preset database, the second training sample consisting of the initial medical data and the corresponding clinical data; A model is generated based on the second training sample to obtain a data preprocessing model for cleaning the initial medical data.

5. The method according to claim 1, characterized in that, For each of the first training samples, the joint probability of the clinical data is calculated based on the adjusted preset confidence level, the first predicted inclusion information of the first training sample, and the second predicted inclusion information. The target group of the clinical data is then determined based on the joint probability, including: For each of the aforementioned clinical data, the same predicted group is obtained based on the first predicted enrollment information and the second predicted enrollment information. Calculate the joint probability of the same prediction group based on the first and second prediction probabilities of the same prediction group and the adjusted preset confidence level; The identical prediction groups are sorted according to the joint probability, and the identical prediction with the highest joint probability is taken as the target group of the clinical data.

6. A method for patient enrollment, characterized in that, include: Obtain clinical data from patients awaiting treatment; The clinical data of the patients to be treated are processed using the target patient enrollment model to obtain the first prediction group information and the second prediction group information of the clinical data; the target patient enrollment model is obtained by the patient enrollment model generation method according to any one of claims 1 to 5. The target group of the clinical data is determined based on the information from the first prediction group and the second prediction group, as well as the target confidence level.

7. The method according to claim 6, characterized in that, The step of determining the target group of the clinical data based on the first prediction group information, the second prediction group information, and the target confidence level includes: Based on the first and second predicted group information, obtain the same predicted group; Calculate the joint probability of the same prediction group based on the first and second prediction probabilities of the same prediction group and the target confidence level; The identical prediction groups are sorted according to the joint probability, and the identical prediction with the highest joint probability is taken as the target group of the clinical data.

8. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to: perform the patient enrollment large model generation method as described in any one of claims 1 to 5, and / or the patient enrollment method as described in any one of claims 6 to 7.

9. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are configured as: the patient enrollment large model generation method as described in any one of claims 1 to 5, and / or the patient enrollment method as described in any one of claims 6 to 7.

Citation Information

Patent Citations

  • Disease prediction method and device, equipment and storage medium

    CN115910355A

  • Systems and methods for automatically identifying a candidate patient for enrollment in a clinical trial

    US20220084633A1