A system and method for assisting in formulating inclusion and exclusion criteria

By designing an inclusion standard assisted formulation system, using deep learning classification models to build a knowledge graph, the problem of inclusion and exclusion of irregular standards in clinical trials is solved, the formulation efficiency and scientificity are improved, and the quality of subjects is ensured.

CN114385825BActive Publication Date: 2025-07-01SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111531659.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-07-01
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The standards for inclusion and exclusion in existing clinical trials are not standardized and unscientific, resulting in the number of qualified subjects not meeting the standards, affecting the implementation of the trial.

Method used

Design a system for assisted formulation of inclusion standard, obtain clinical trial registration data through the data preprocessing module, extract medical entities and relationships, and use deep learning classification models to build and update the knowledge map of inclusion standard, and formulate standard decisions.

Benefits of technology

It improves the efficiency and scientific nature of the formulation of inclusion standards, ensures the quality of subjects in clinical trials, and solves the problems of inefficiency and excessive standards in the traditional formulation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114385825B_ABST
    Figure CN114385825B_ABST
Patent Text Reader

Abstract

The present invention relates to a system for assisting in formulating inclusion and exclusion criteria. A data preprocessing module: obtains clinical trial registration data and extracts medical entities and medical entity relationships; a graph construction and update module: is used to obtain the weights between each medical entity through a constructed deep learning classification model, and construct and update the inclusion and exclusion criteria knowledge graph; an inclusion and exclusion criteria formulation module: is used to formulate inclusion and exclusion criteria decisions according to the inclusion and exclusion criteria knowledge graph, provide assistance for decision-making. Compared with the prior art, the present invention has the advantages of solving the problems of low efficiency and excessive strictness in the traditional formulation of inclusion and exclusion criteria, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of clinical trial auxiliary decision-making, and particularly to a system and method for assisting in formulating inclusion and exclusion criteria. Background Art

[0002] Clinical trials refer to any scientific research that requires human beings (patients or healthy volunteers, also known as "subjects") to participate in the clinical trial process, and are an essential process for researchers to develop new therapies and evaluate their effectiveness and safety. Screening criteria are the main indicators for identifying whether a subject meets a certain clinical trial, and are divided into inclusion criteria and exclusion criteria, mostly in the form of unstructured free text. Clinical trials recruit eligible subjects according to the screening criteria. The traditional clinical trial subject recruitment work is cumbersome. The recruiter needs to manually compare the electronic medical records of patients or the physical examination reports of healthy volunteers one by one according to the screening criteria to screen eligible patients or healthy volunteers to participate in the clinical trial.

[0003] However, at present, the inclusion and exclusion criteria in clinical trial registration data are irregular, unscientific, and not rigorous, resulting in the number of eligible subjects not meeting the standard and the trial cannot be carried out. Summary of the Invention

[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a system and method for assisting in formulating inclusion and exclusion criteria.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A system for assisting in formulating inclusion and exclusion criteria, the system includes:

[0007] A data preprocessing module: obtaining clinical trial registration data, and extracting medical entities and medical entity relationships;

[0008] A graph construction and update module: used to obtain the weights between medical entities through the constructed deep learning classification model, and construct and update the inclusion and exclusion criteria knowledge graph;

[0009] An inclusion and exclusion criteria formulation module: used to formulate inclusion and exclusion criteria decisions according to the inclusion and exclusion criteria knowledge graph and provide auxiliary decision-making assistance.

[0010] The structure of the inclusion and exclusion criteria knowledge graph includes:

[0011] Nodes: used to store medical entities in the inclusion and exclusion criteria;

[0012] Relationships: used to store medical entity relationships between medical entities in the inclusion and exclusion criteria.

[0013] The medical entities include treatment methods, age, and gender.

[0014] The described medical entity relationships include association relationships and weight relationships.

[0015] A method for formulating the inclusion and exclusion criteria auxiliary formulation system as described, the method comprising the following steps:

[0016] Step 1: Screen and obtain clinical trial registration data, extract medical entities and medical entity relationships from the clinical trial registration data, and standardize the concepts and terms of the medical entities.

[0017] Step 2: Construct a knowledge base based on the medical entities and medical entity relationships extracted in Step 1, construct a deep learning classification model through the pre-trained language model bioBERT based on the biomedical corpus and the attention mechanism, and train the weights between each medical entity in the deep learning classification model through the graph construction and update module.

[0018] Step 3: Construct an inclusion and exclusion criteria knowledge graph through the graph construction and update module, use the medical entities, medical entity relationships, and the weights between each medical entity as the nodes, relationships, and relationship attributes of the inclusion and exclusion criteria knowledge graph respectively, and store the nodes, relationships, and relationship attributes of the inclusion and exclusion criteria knowledge graph in the graph database.

[0019] Step 4: Perform evidence-based reasoning on the inclusion and exclusion criteria knowledge graph and give a recommended inclusion and exclusion criteria decision, the content of the inclusion and exclusion criteria decision including treatment method, age, gender, and judgment of pregnancy.

[0020] Step 5: Determine the inclusion and exclusion criteria decision, including adopting the recommended inclusion and exclusion criteria decision according to the clinical trial, adjusting the recommended inclusion and exclusion criteria decision, and adopting the self-determined inclusion and exclusion criteria decision.

[0021] In the described Step 2, the process of training the weights between each medical entity in the deep learning classification model through the graph construction and update module specifically includes the following steps:

[0022] Step 201: Obtain the hidden layer output h of each medical entity in each inclusion and exclusion criteria decision through the pre-trained language model bioBERT based on the biomedical corpus i , and obtain the hidden information u of the hidden layer output h of each medical entity i ; i ;

[0023] Step 202: Obtain the weights of each medical entity in the inclusion and exclusion criteria decision.

[0024] Step 203: Obtain the vector representation of each inclusion and exclusion criteria decision by performing weighted summation on the weights of each medical entity.

[0025] Step 204: Use the vector representation of each inclusion and exclusion criterion as classification features, train a deep learning classification model based on the SoftMax classifier, and update the parameters using the cross-entropy classification loss function.

[0026] In step 201 described above, the hidden layer output h of each medical entity i The hidden information u of i The expression of is:

[0027] u i = tanh(W w h i + b w )

[0028] Among them, W w is the weight matrix, b w is the corresponding bias term, h i is the hidden layer output of the i-th medical entity, and u i is the hidden information of the hidden layer output of the i-th medical entity.

[0029] In step 202 described above, the calculation formula for the weight of each medical entity is:

[0030]

[0031] Among them, α i is the weight of the i-th medical entity in the medical evidence, u i is the hidden information of the hidden layer output of the i-th medical entity, u w is the randomly initialized word-level context vector, and n is the number of medical entities for inclusion and exclusion criterion decision-making.

[0032] In step 203 described above, the vector representation of the inclusion and exclusion criterion decision-making is:

[0033]

[0034] Among them, v i is the vector representation of the i-th inclusion and exclusion criterion decision-making, n is the number of medical entities for inclusion and exclusion criterion decision-making, α i is the weight of the i-th medical entity in the medical evidence, and h i is the hidden layer output of the i-th medical entity.

[0035] In step 204 described above, the expression of the cross-entropy classification loss function is:

[0036]

[0037] Among them, m is the number of relationship categories of the inclusion and exclusion criteria, y j is the true probability distribution of the j-th relationship category, Represents the predicted probability distribution of the j-th relationship category.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] The present invention extracts entities and their relationships in the inclusion and exclusion criteria based on authoritative and reliable medical evidence, uses a deep learning classification model to learn the weights between entities, and constructs an inclusion and exclusion criteria knowledge graph based on entity, relationship, and their weight data; the inclusion and exclusion criteria knowledge graph is based on graph theory and stored in a graph database in the form of a graph structure composed of nodes and relationships; in combination with the inclusion and exclusion criteria of the completed clinical trial registration data, it assists doctors in quickly and effectively formulating the inclusion and exclusion criteria for clinical trials; continuously improves the ability to assist in formulating the inclusion and exclusion criteria knowledge graph; solves the problems of low efficiency and excessive strictness in the traditional formulation of inclusion and exclusion criteria. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flowchart of the method for assisting in formulating inclusion and exclusion criteria provided by an embodiment of the present invention.

[0041] Figure 2 It is a module structure diagram of the system for assisting in formulating inclusion and exclusion criteria provided by an embodiment of the present invention.

[0042] Figure 3 It is a storage structure diagram of the inclusion and exclusion criteria knowledge graph of the method for assisting in formulating inclusion and exclusion criteria provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] Embodiment

[0045] The present invention provides a system and method for assisting in formulating inclusion and exclusion criteria, relying on an inclusion and exclusion criteria knowledge graph constructed based on authoritative and reliable medical evidence, and assisting doctors in quickly and effectively formulating inclusion and exclusion criteria according to the existing clinical trial registration data.

[0046] As Figure 1 shown, a method for assisting in formulating inclusion and exclusion criteria includes the following steps:

[0047] Step 1: Obtain clinical trial registration data, and extract medical entities and medical entity relationships therefrom;

[0048] Step 2: Construct a knowledge base according to the extracted medical entities and medical entity relationships, construct a deep learning classification model based on the pre-trained language model bioBERT on the biomedical corpus and the attention mechanism, and perform named entity work through the graph construction and update module;

[0049] Step 3: Through the graph construction and update module, construct the inclusion and exclusion criteria knowledge graph: Use medical entities, medical entity relationships, and weights between medical entities as nodes, relationships, and relationship attributes of the inclusion and exclusion criteria knowledge graph respectively, and store the nodes, relationships, and relationship attributes of the inclusion and exclusion criteria knowledge graph in the graph database;

[0050] Step 4: Conduct evidence-based reasoning on the inclusion and exclusion criteria knowledge graph and give a recommended inclusion and exclusion criteria decision;

[0051] Step 5: The doctor determines the inclusion and exclusion criteria. The inclusion and exclusion criteria decision means that the doctor adopts or adjusts the recommended inclusion and exclusion criteria decision according to the clinical trial, or adopts the self-defined inclusion and exclusion criteria;

[0052] Step 1 specifically includes the following steps:

[0053] Step 101: Medical experts select medical evidence from relevant documents;

[0054] Step 102: Extract medical entities and medical entity relationships, and standardize the concepts and terms of medical entities.

[0055] In Step 2, the process of training the weights between various medical entities in the deep learning classification model through the graph construction and update module specifically includes the following steps:

[0056] Step 201: Obtain the hidden layer output of each medical entity in each inclusion and exclusion criterion through the pre-trained language model bioBERT based on the biomedical corpus. Use h i to represent the hidden layer output of the i-th medical entity, and obtain the hidden information u i of the hidden layer output h i :

[0057] u i =tanh(W w h i +b w )

[0058] where W w is the weight matrix, b w is the corresponding bias term, h i is the hidden layer output of the i-th medical entity, and u i is the hidden information of the hidden layer output of the i-th medical entity;

[0059] Step 202: Obtain the weight of each medical entity in the inclusion and exclusion criteria. The calculation formula for the weight of each medical entity is:

[0060]

[0061] where α iis the weight of the i-th medical entity in the medical evidence, u i is the hidden information output by the hidden layer of the i-th medical entity, u w is the randomly initialized word-level context vector, and n is the number of medical entities in the inclusion and exclusion criteria;

[0062] Step 203: Obtain the vector representation of each inclusion and exclusion criterion through weighted summation. The vector representation of the inclusion and exclusion criterion is:

[0063]

[0064] where, v i is the vector representation of the i-th inclusion and exclusion criterion, n is the number of medical entities in the inclusion and exclusion criteria, and α i is the weight of the i-th medical entity in the medical evidence, and h i is the output of the hidden layer of the i-th medical entity;

[0065] Step 204: Use the vector representation of each inclusion and exclusion criterion as a classification feature, train a deep learning classification model based on the SoftMax classifier, and update the parameters using the cross-entropy classification loss function. The expression of the cross-entropy classification loss function is:

[0066]

[0067] where, m is the number of relationship categories of the inclusion and exclusion criteria, and y j is the true probability distribution of the j-th relationship category, represents the predicted probability distribution of the j-th relationship category.

[0068] In step 4, the inclusion and exclusion criteria decision includes treatment method, age, gender, and pregnancy status.

[0069] As Figure 2 shown in the module structure diagram of an inclusion and exclusion criteria auxiliary formulation system, the system includes:

[0070] Atlas construction and update module: used for the deep learning classification model to learn weights, construct and update the inclusion and exclusion criteria knowledge atlas;

[0071] Inclusion and exclusion criteria formulation module: used to assist doctors in formulating inclusion and exclusion criteria for clinical trials.

[0072] As Figure 3 shown in the storage structure diagram of the inclusion and exclusion criteria knowledge atlas of an inclusion and exclusion criteria auxiliary formulation method, the inclusion and exclusion criteria knowledge atlas based on the graph structure is easy to store, maintain, and query nodes and relationships. Based on the graph theory traversal algorithm, it can simply and efficiently infer and retrieve multiple inclusion and exclusion criteria methods most relevant to the current clinical trial to be formulated for doctors.

[0073] The inclusion and exclusion criteria knowledge atlas structure includes:

[0074] Node: used to store entities in the inclusion and exclusion criteria, where the entities include treatment methods, age, and gender;

[0075] Relationship: used to store the relationships between various entities in the inclusion and exclusion criteria, where the relationships include association relationships and weight relationships.

[0076] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any staff member familiar with the technical field of the present invention can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for assisting in formulating inclusion and exclusion criteria, characterized in that, The method includes the following steps: Step 1: Screen and obtain clinical trial registration data, extract medical entities and medical entity relationships from the clinical trial registration data, and standardize the concepts and terms of the medical entities; Step 2: Construct a knowledge base based on the medical entities and medical entity relationships extracted in Step 1. Build a deep learning classification model through the pre-trained language model bioBERT and attention mechanism based on the biomedical corpus, and train the weights between each medical entity in the deep learning classification model through the graph construction and update module; Step 3: Construct an inclusion and exclusion criteria knowledge graph through the graph construction and update module. Use the medical entities, medical entity relationships, and the weights between each medical entity as the nodes, relationships, and relationship attributes of the inclusion and exclusion criteria knowledge graph respectively, and store the nodes, relationships, and relationship attributes of the inclusion and exclusion criteria knowledge graph in the graph database; Step 4: Perform evidence-based reasoning on the inclusion and exclusion criteria knowledge graph and give a recommended inclusion and exclusion criteria decision. The content of the inclusion and exclusion criteria decision includes treatment methods, age, gender, and determination of pregnancy; Step 5: Determine the inclusion and exclusion criteria decision, including adopting the recommended inclusion and exclusion criteria decision according to the clinical trial, adjusting the recommended inclusion and exclusion criteria decision, and adopting a self-formulated inclusion and exclusion criteria decision; The inclusion and exclusion criteria auxiliary formulation system includes: Data preprocessing module: Obtain clinical trial registration data and extract medical entities and medical entity relationships; Graph construction and update module: Used to obtain the weights between each medical entity through the constructed deep learning classification model, and construct and update the inclusion and exclusion criteria knowledge graph; Inclusion and exclusion criteria formulation module: Used to formulate inclusion and exclusion criteria decisions based on the inclusion and exclusion criteria knowledge graph and provide auxiliary decision-making assistance.

2. The method for assisting in formulating inclusion and exclusion criteria according to claim 1, wherein The structure of the inclusion and exclusion criteria knowledge graph described above includes: Nodes: Used to store medical entities in the inclusion and exclusion criteria; Relationships: Used to store medical entity relationships between each medical entity in the inclusion and exclusion criteria.

3. The method for assisting in formulating the inclusion and exclusion criteria according to claim 2, wherein The medical entities described above include treatment methods, age, and gender.

4. The method for assisting in formulating inclusion and exclusion criteria according to claim 2, wherein The medical entity relationships described above include association relationships and weight relationships.

5. The method for assisting in formulating a inclusion and exclusion criterion according to claim 1, wherein, In Step 2 described above, the process of training the weights between each medical entity in the deep learning classification model by the graph construction and update module specifically includes the following steps: Step 201: Obtain the hidden layer output h of each medical entity in each inclusion and exclusion criteria decision through the pre-trained language model bioBERT based on the biomedical corpus i , and obtain the hidden information u of the hidden layer output h i of each medical entity i ; Step 202: Obtain the weights of each medical entity in the inclusion and exclusion criteria decision; Step 203: Obtain the vector representation of each inclusion and exclusion criteria decision by weighted summation of the weights of each medical entity; Step 204: Use the vector representation of each inclusion and exclusion criteria as a classification feature, train the deep learning classification model based on the SoftMax classifier, and update the parameters using the cross-entropy classification loss function.

6. The method for assisting in formulating inclusion and exclusion criteria according to claim 5, characterized in that In the said step 201, the hidden information u i of the hidden layer output h i of each medical entity has the following expression: u i = tanh(W w h i + b w ) Among them, W w is the weight matrix, b w is the corresponding bias term, h i is the hidden layer output of the i-th medical entity, and u i is the hidden information of the hidden layer output of the i-th medical entity.

7. A method for assisting in formulating inclusion and exclusion criteria according to claim 6, characterized in that, In Step 202 described above, the calculation formula for the weight of each medical entity is: where α i is the weight of the i-th medical entity in the medical evidence, u i is the hidden information output by the hidden layer of the i-th medical entity, u w is the randomly initialized word-level context vector, and n is the number of medical entities for the inclusion and exclusion criteria decision.

8. A method for assisting in formulating inclusion and exclusion criteria according to claim 7, characterized in that In Step 203 described above, the vector representation of the inclusion and exclusion criteria decision is: Among them, v i is the vector representation of the i-th inclusion and exclusion criterion decision, n is the number of medical entities for inclusion and exclusion criterion decisions, α i is the weight of the i-th medical entity in the medical evidence, h i is the output of the hidden layer of the i-th medical entity.

9. The method for assisting in formulating inclusion and exclusion criteria according to claim 8, wherein In Step 204 described above, the expression of the cross-entropy classification loss function is: Among them, m is the number of relationship categories of the inclusion and exclusion criteria, and y j is the true probability distribution of the j-th relationship category, represents the predicted probability distribution of the j-th relationship category.