Relevance feedback for improving performance of clustering models that cluster patients with similar profiles together

By combining automatic feature selection and feedback from clinicians, and adjusting patient comparison metrics, the accuracy problem of patient group identification in existing technologies is solved, achieving more efficient patient group identification and clustering.

CN108780661BActive Publication Date: 2026-03-31KONINKLIJKE PHILIPS NV
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-03-08
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and segment patient groups with similarities and differences, especially in clinical trials and medical research, leading to the reliability and accuracy of results being affected by external factors.

Method used

By combining automatic feature selection with relevant feedback from clinicians, the feature set and feature weights of patient comparison metrics are adjusted, and patient clustering is performed using a graphical user interface to improve the accuracy of clustering results.

Benefits of technology

It provides clinician-based relevance feedback, improves the selection of patient groups, enhances the accuracy and reliability of clustering results, and simplifies the workflow for clinicians.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN108780661B_ABST
    Figure CN108780661B_ABST
Patent Text Reader

Abstract

In patient cohort identification, a clustering of patients is performed using a patient comparison metric dependent on a feature set (24) (30). Information about sample patients similar or dissimilar to a query patient according to the clustering is displayed. A user input comparison value comparing a sample patient to the query patient is received. The feature set and / or feature weights are adjusted to generate an adjusted patient comparison metric having improved consistency with the user input comparison value. The clustering is repeated using the adjusted patient comparison metric. A patient cohort is identified from a cluster (34) containing the query patient resulting from the last clustering repetition. Information about sample patients can be shown by simultaneously displaying two or more graphical modal representations (70, 72, 74), each graphical modal representation plotting a sample patient and the query patient for two or more features of a modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following text covers the fields of medicine, electronic clinical decision support (CDS), clinical research, genomics, and related fields. Background Technology

[0002] Many medical tasks benefit from identifying patient groups with relevant similarities. For example, a key initial step in designing clinical trials is identifying patients to be enrolled. To ensure the validity of results, enrolled patients should be similar enough that dissimilar patient outcomes can be reliably attributed to the clinical trial's objective (e.g., a new pharmaceutical drug) rather than to differences in patient outcomes due to external factors such as age, sex, race, or the presence / absence of chronic medical conditions (which are irrelevant to the clinical trial's objective). Identifying suitable patients for clinical trials is challenging because patient outcomes are influenced by numerous relevant factors.

[0003] Group identification can also be performed during the analysis of clinical trial results after enrollment. Within the enrollment, patients with positive results naturally form two groups of interest relative to those with negative results. However, these groups can be further segmented based on similarities and differences within the positive and negative groups to identify and interpret any external factors that may influence the raw data results of the clinical trial.

[0004] Similarity group identification tasks are performed in other types of medical research, such as to assess disease risk factors or to perform “meta-studies” that combine data from many previous studies.

[0005] Other medical tasks include the clinical diagnosis and management of patients. In such tasks, clinicians can benefit from comparing current patients with similar past patients. Similarly, the task of identifying “similar” patients is challenging. No two patients are identical, and group selection tasks require assessing which differences are significant or insignificant.

[0006] The following discloses new and improved systems and methods for solving the above and other problems. Summary of the Invention

[0007] In one aspect of the disclosure, a patient group identification device is disclosed. The computer has a display unit and at least one user input device. The computer communicates with a patient database storing patient data, the patient data including values ​​of features for patients in the patient database. The computer is programmed to perform a patient group identification method comprising: performing an automatic feature selection process on the patient data to select a set of features, and performing automatic clustering of patients in the patient database using a patient comparison metric dependent on the set of features; performing at least one iteration comprising: displaying on the display unit information about one or more sample patients similar to or dissimilar to a query patient according to the automatic clustering; receiving, via the at least one user input device, comparison values ​​input by the user comparing the one or more sample patients to the query patient; adjusting the patient comparison metric to increase the consistency between the comparison values ​​calculated by the patient comparison metric that compare the one or more sample patients to the query patient and the comparison values ​​input by the user, wherein the adjustment comprises adjusting at least one of the set of features and feature weights of the patient comparison metric; and repeating the automatic clustering using the adjusted patient comparison metric. The adjusted patient comparison metric generated from the last iteration is used to identify patient groups for the queried patient.

[0008] In another aspect of the disclosure, a patient group identification device is disclosed. The computer has a display unit and at least one user input device. The computer communicates with a patient database storing patient data, the patient data including values ​​of features for patients in the patient database. The computer is programmed to perform a patient group identification method, the method comprising: simultaneously displaying two or more graphical modal representations on the display unit, wherein each graphical modal representation plots patients in the database for two or more coordinate features of the modality; receiving a selection of a cluster of patients in one graphical modal representation; and, in response to receiving the selection, highlighting patients in the selected cluster in one or more other simultaneously displayed graphical modal representations.

[0009] In another aspect of the disclosure, a patient group identification method is disclosed, which is executed in conjunction with a computer having a display unit and at least one user input device and communicating with a patient database storing patient data, the patient data including values ​​of features for patients in the patient database. The patient group identification method includes the following: performing automatic clustering of patients in the patient database using a patient comparison metric that depends on a set of features; performing at least one iteration, the iteration including: displaying information on the display unit about one or more sample patients that are similar to or dissimilar to a query patient according to the automatic clustering; receiving comparison values ​​input by a user via the at least one user input device, the user-input comparison values ​​comparing the one or more sample patients to the query patient; adjusting at least one of the set of features and feature weights of the patient comparison metric to generate an adjusted patient comparison metric that has improved consistency with the user-input comparison values ​​compared to an unadjusted patient comparison metric; and repeating the automatic clustering using the adjusted patient comparison metric. Patient groups for the query patient are identified as at least a portion of the clusters containing the query patient generated by the automatic clustering repetition of the last iteration.

[0010] One advantage is that it provides relevant feedback from clinicians to improve cohort selection.

[0011] Another advantage is that it provides relevant feedback for group selection based on the overall patient level analysis of clinicians.

[0012] Another advantage is that it provides relevant feedback from clinicians for selecting relevant features without requiring clinicians to perform feature level analysis.

[0013] Another advantage is that it provides a graphical user interface, through which clinicians can visualize the interrelationships between different modalities (clinical, radiological, genomic, demographic, physiological, etc.).

[0014] A given embodiment may not provide any of the aforementioned advantages, but may provide one, two, more or all of the aforementioned advantages, and / or may provide other advantages, as will become apparent to those skilled in the art upon reading and understanding this disclosure. Attached Figure Description

[0015] This invention can take the form of various components and component arrangements, as well as various steps and step arrangements. The accompanying drawings are for illustrative purposes only and should not be construed as limiting the invention.

[0016] Figure 1The illustration shows a patient group identification device.

[0017] Figure 2 The schematic diagram illustrates the composition of Figure 1 The patient group identification device appropriately performs the patient group identification method.

[0018] Figure 3 and Figure 4 Schematic illustration Figure 2 Two exemplary examples of suitable implementations of the method of presentation operation.

[0019] Figure 5 The illustrations depict the visual representation and navigation tools for patient groups as described herein. Detailed Implementation

[0020] This paper recognizes that the complexity of grouping patients can be reduced by selecting an appropriate (reduced) set of patient characteristics to group them. The set of patient characteristics used for group selection should include those relevant to the medical task at hand (e.g., selecting patients for clinical trials, or selecting patients similar to those currently in a clinical diagnosis), and should not include those irrelevant to the medical task. Characteristic selection is important because the number of available patient characteristics is often quite large and can include, for example: demographic data (age, sex, weight, ethnicity, etc.); presence / absence of chronic behavioral conditions (smoking, heavy drinking, consumption of various recreational drugs, etc.); presence / absence of various chronic clinical conditions (hypertension, diabetes, asthma, heart disease, etc.); presence / absence of various acute illnesses (pneumonia or other acute respiratory illnesses, various cancer conditions, etc.); related characteristics (e.g., cancer stage and grade); and so on. The rapidly developing field of genomics is quickly adding to the list of available patient characteristics because gene sequencing can provide a wealth of genomic markers with varying known or suspected associations with various medical conditions. For example, some medical databases contain data with hundreds or more defined features, while the continued expansion of genomic data availability may increase the number of patient features to thousands. Such a large feature space poses a significant challenge to selecting the “best” set of features for a group of patients suited to clinical tasks.

[0021] Many unsupervised (reduced) feature set selection techniques are known. Typical automatic feature selection techniques measure the discriminative power of features and select the most discriminative features. One such technique is Principal Component Analysis (PCA), which selects features to capture the variance of a dataset with a reduced number of features. Other discriminative measures can be employed, such as information gain (IG) per feature or various pairwise feature correlation measures (e.g., selecting features that provide the highest IG, or eliminating features that are strongly correlated with other features).

[0022] Despite their power, unsupervised automated feature set selection techniques have significant limitations when used to select features for identifying patient groups. They can select highly discriminative features that are irrelevant to the clinical task, rather than other features with lower discriminative power but relevant to the medical task. Unsupervised feature set selection techniques also fail to consider the physiological basis for why a particular feature should be demonstrably clinical. For example, consider a suspected clinical condition due to a problem with a specific metabolic pathway. A genomic marker known to be part of that pathway might be relevant in this case, but it might fail to select that marker if PCA or another unsupervised feature selection technique has low overall discriminative power.

[0023] In principle, these problems can be mitigated by manual feature selection performed by clinicians, or by a hybrid approach in which the clinician examines and adjusts the initial set of features generated by unsupervised automatic feature selection (relevance feedback). However, in practice, clinicians may not be able to articulate why a patient is considered similar or dissimilar to a patient of interest (referred to herein as the “query patient”) based on specific features. Based on past experience and overall training, clinicians tend to view patients holistically. Therefore, while clinicians may identify a particular patient as similar or dissimilar to the query patient, they may not be able to precisely articulate which features effectively encompass the similarity or dissimilarity. Furthermore, it may be impractical for skilled clinicians to spend the necessary time sifting through hundreds of available candidate features to identify demonstrable features for a given clinical task.

[0024] The techniques disclosed herein overcome these difficulties by combining unsupervised feature selection with subsequent patient-level relevance feedback provided by clinicians, through examination of automated clustering performed using an automatically selected feature set. In these methods, an initial automated feature set is used to perform unsupervised automated patient clustering to identify patient clusters including the query patient, as well as other clusters. The cluster containing the query patient defines a set of similar patients based on the initial feature set, while other clusters group various less similar patients. Clinicians then examine these clustering results and select similar or dissimilar patients (relevance feedback). The feature set is then automatically adjusted to better suit these clinician selections, and the clustering is repeated using the adjusted feature set. This process can be repeated until the unsupervised automated clustering produces clusters that are (at least substantially) satisfactory to the clinician.

[0025] The method utilizes unsupervised feature set selection to provide an initial approximate selection of a large feature space. Patients are clustered using an initial feature set generated by PCA or another unsupervised feature selection process to identify similar (or dissimilar) patients corresponding to the query patient as measured using that initial feature set. One or more similar (or dissimilar) sample patients are presented to the clinician, and a user interface is provided through which the clinician can provide relevance feedback. For example, a set of similar sample patients {P} can be presented to the clinician. C}, which was identified in the initial clustering as being related to the diagnosed patient (query patient P). Q Similarity. These "similar" sample patients can, for example, be assigned from the cluster to the query patient P. Q Extract from the same cluster, or use a distance metric defined by the initial feature set to extract from the clusters with the shortest distance |P. Q -P C Extraction from a subset of the clusters. Then, clinicians can use a ranking scale of 1...5 to rank patients relative to the queried patient P. Q Similarity or dissimilarity is assigned, where 1 indicates the most similar and 5 indicates the least similar. Subsequently, feature set adjustment is performed to generate an adjusted feature set that more closely aligns with the clinician's similarity ranking of the patients under consideration. The clustering is repeated again, and the clinician is presented with a list of queried patients P. Q Alternatively, clustering of some subsets can be used for similarity ranking. This process can be repeated until the clinician is certain that the query patient P is included. Q Clustering is a suitable grouping for performing the medical task at hand.

[0026] Advantageously, this method for relevance feedback does not require clinicians to evaluate a feature set at an abstract level of the feature space. Instead, clinicians operate in a more familiar environment of comparing and contrasting individual patients, allowing them to leverage past experience and overall training to make relevance feedback decisions. Preferably, the user interface enables clinicians to find each proposed similar patient P under consideration. C Complete medical records, and access to patient P Q Complete medical records are available so that relevant feedback assessments can be conducted using the same information sources that clinicians use to access them.

[0027] refer to Figure 1 The patient group identification device includes a computer with a display unit and at least one user input device. An exemplary computer includes two computers: a server computer 10, which performs computationally intensive operations such as feature selection or clustering; and a user interface computer 12, such as a desktop computer, laptop computer, tablet computer, etc., which includes or is operatively connected to the display unit 14 and at least one user input unit, such as an exemplary keyboard 16 and a mouse 18 (or a trackball, touchpad, touchscreen, or other pointing device). Computers 10 and 12 communicate with a patient database 20, which stores patient data including values ​​for characteristics of patients in the patient database. The patient database 20 may include, for example, one or more of the following: electronic health record (EHR), electronic medical record (EMR), picture archiving and communication system (PACS, for radiological images / data), cardiovascular information system (CVIS), and various combinations thereof. The various components 10, 12, and 20 can be interconnected via various data paths, such as a hospital local area network (LAN), a wireless LAN (WLAN), the Internet, and various combinations thereof.

[0028] Computers 10 and 12 are programmed to perform various processes. An automated feature selection process 22 is executed to select a reduced set of features from a typically larger set of available features contained in or derived from information contained in the patient database 20. The feature selection process 22 may be, for example, a principal component analysis (PCA) feature selection process, an information gain (IG) feature ranking process, a pairwise correlation feature removal process, etc. The automated feature selection process 22 identifies a set 24 of features, typically selecting features with high discriminative power. It should be appreciated that the patient database 20 may (explicitly or implicitly, i.e., derived from other stored information) store dozens, hundreds, or more features for each patient. Some non-restrictive illustrative features include: demographic characteristics (patient age, sex, weight, ethnicity, etc.); features indicating the presence or absence of chronic behavioral conditions (smoking, heavy drinking, consumption of various recreational drugs, etc.); features indicating the presence or absence of various chronic clinical conditions (hypertension, diabetes, asthma, heart disease, etc.); features indicating the presence or absence of various acute diseases (pneumonia or other acute respiratory diseases, various tumor conditions, etc.); condition-specific features, such as cancer stage, cancer grade; genomic features, such as the values ​​of specific genes, various protein expression levels, or other genetic markers; and so on. Therefore, a patient dataset 26 is generated, in which features in the set 24 are annotated or represented for each patient by values ​​extracted from the patient database 20.

[0029] The clustering process 30 performs unsupervised learning to group patients in the patient dataset 26 into cluster sets 32. Typically, the goal is to identify and query patient P. Q Patient groups with similar patients—therefore, the cluster set 32 ​​includes: those containing the query patient P Q Cluster 34 (or, in other words, cluster 34 is the query for patient P) Q The clusters to which the clusters belong (the clusters generated by clustering process 30); and other clusters 36 generated by the clustering process. The clustering process can employ any known clustering method, such as k-means clustering, connectivity-based or hierarchical clustering, centroid-based clustering, expectation-maximization (EM) clustering, etc. The clustering uses a patient comparison metric based on a feature-dependent set 24. For two patients P... i and P j In this article, the abbreviation |P is used. i -P j Write down the values ​​of the patient comparison metric for comparing these two patients. As a non-limiting example, the patient comparison metric could be a distance metric, whose value is smaller for more similar patients. Some suitable distance metrics are Euclidean distances:

[0030]

[0031] Where n = 1, ..., N are the features in the set 24 indexed features, f n,i and f n,j These are for patients P i and P j The value of the nth feature, and w n It is the feature weight of the nth feature in the Euclidean distance of expression (1). As another example, the patient comparison metric can be the squared Euclidean distance, which is the same as in expression (1) except that the square root is omitted. Instead of a distance metric, the patient comparison metric can alternatively be a similarity metric, whose value is larger for more similar patients. These are merely illustrative examples. Generally, the patient comparison metric is preferably functionally dependent on the set 24 of features, wherein the contribution of individual features is controlled by feature weights (e.g., the feature weights w in the illustrative Euclidean distance of expression (1)). n It is also anticipated that patient comparison metrics will be used that do not include adjustable feature weights.

[0032] For the selected clustering process 30, the features of the clustering result 32 depend on the details of the patient comparison metric, particularly the set 24 of features on which the patient comparison metric functionally depends, and the feature weights (if adjustable). The automatic feature selection process 22 selects features based on an assessment of their discriminative power; however, this method is capable of selecting highly discriminative features rather than features with lower discriminative power that are more relevant to the medical task at hand, or features with some physiological basis related to the task at hand.

[0033] exist Figure 1 In the exemplary patient group identification device, these problems are addressed by providing relevance feedback to improve patient comparison metrics, for example, by adjusting the set of features 24 and / or feature weights. For this purpose, a graphical user interface (GUI) process 40 is implemented, for example, on computer 12 in the exemplary embodiment. The GUI process 40 (on display unit 14) presents information about the automatic clustering and query of patient P. Q Information on one or more sample patients that are similar or dissimilar. For example, the sample patients could be those from a list containing the query patient P. Q Similar patient samples were (pseudo)randomly selected from cluster 34. Alternatively, similar patient samples could be selected non-randomly from cluster 34, for example, selecting the patient P closest to the query. Q Patients, as measured by patient comparison metrics. Alternatively, dissimilar samples of patients can be randomly selected from other clusters 36, or patient P can be queried from their centroid distance. QDissimilar sample patients are selected from the furthest remaining clusters, as measured by a patient comparison metric. These sample patients are presented to the user via display component 14, prompting the clinician to provide comparison values ​​between one or more sample patients and the query patient. For example, the clinician may be asked to rank the similarity between the sample patients and the query patient on a scale of 1–5 (or 1–10, etc.). Alternatively, the clinician may be asked to select which of two sample patients is most similar to the query patient. It should be noted that this method does not (at least not directly) require the clinician to assess similarity at the feature level, but rather at the patient level. This leverages the strength of clinicians typically trained to analyze patients based on all available information in the patient record, along with their education and experience. This method avoids requiring the physician to perform feature-level analysis, which is not within the natural scope of a clinician's practice.

[0034] The GUI process 40 receives comparison values ​​from user input via at least one user input device 16, 18, comparing one or more sample patients with the queried patient. This constitutes "relevance feedback". Then, the patient comparison metric adjustment process 42 adjusts the set of features 24 and / or adjusts the feature weights w. n To increase the likelihood of matching one or more sample patients with the query patient P Q The consistency between the comparison value calculated by the patient comparison metric and the comparison value input by the user is compared.

[0035] In one approach, the patient comparison metric adjustment process 42 performs feature set adjustment iterations, each iteration of which is performed as follows. In the first step of the iteration, the feature set 24 is adjusted by adding features to the set or removing features from the set to produce an adjusted set of features. Then, a comparison value is computed using the patient comparison metric and the candidate adjusted set of features, the comparison value comparing one or more sample patients with the query patient P. Q A comparison is made. Based on whether the consistency between the calculated comparison value and the comparison value input by the user increases or decreases, a candidate set of adjusted features is accepted or rejected. If rejected, the candidate set of adjusted features is discarded. If accepted, the candidate set of adjusted features becomes a new (i.e., updated) set 24 of features. This process can be repeated a fixed number of times, or it can be repeated until a certain number of consecutive iterations result in rejection, or some other stopping criteria can be used.

[0036] In another approach, the patient comparison metric adjustment process 42 performs feature weight adjustment iterations, each of which is performed as follows: In the first step of the iteration, the patient comparison metric is adjusted by increasing or decreasing the value of at least one feature weight to generate candidate adjusted patient comparison metrics. Comparison values ​​are calculated using the candidate adjusted patient comparison metrics, which compare one or more sample patients to the query patient. The candidate adjusted patient comparison metrics are accepted or rejected based on whether the consistency between the comparison values ​​and the comparison values ​​input by the user increases or decreases. If accepted, one or more new feature weights are used; if rejected, they are discarded.

[0037] Now for reference Figure 2 , describes the use Figure 1 The process performed by the patient group identification device. In operation 50, a feature selection process 22 is performed to select an (initial) set 24 of features. In operation 52, a clustering process 30 is performed to generate (initial) clusters 32. In operation 54, the clinician is presented with one or more similar and / or dissimilar sample patients, wherein a patient comparison metric is used relative to the query patient P. Q To measure similarity / dissimilarity. More specifically, information about sample patients is presented, preferably in the form of an information request formulated in a way familiar to clinicians, such as ranking the sample patients against the query patient, or a request to identify which of two sample patients is most similar to the query patient. In operation 56, comparison values ​​input by the user are received (e.g., ranking of sample patients, or selection of more similar sample patients from a set of two sample patients). In operation 60, the set of features 24 and / or feature weights w are adjusted. n To increase consistency between the patient comparison metric applied to the sample patients and the comparison value entered by the user. For example, if the user ranks the sample patients as very similar to the query patient, adjustments to the shorter distance between the sample patients and the query patient, as measured by the (adjusted) patient comparison metric, are accepted, while adjustments to increase that distance are rejected. In operation 62, the clustering process 30 is repeated using the adjusted patient comparison metric. The process then returns to operation 54, where similar and / or dissimilar patients are presented to the clinician based on the updated clustering. This loop can be repeated any number of times until, at operation 64, the clinician examines the latest clustering results and concludes that they are satisfactory.

[0038] In the following, some exemplary methods are disclosed for implementing operation 60 as an automatic mapping from an original space to a new space, where a smaller distance is exhibited according to relevant features of a clinical expert (from operation 56). The first exemplary method uses a dimensionality reduction method, while the second exemplary method uses a feature weight adjustment method.

[0039] In the first exemplary method that employs dimensionality reduction, patient data (V) is represented, which contains features F = {f1,..., f m} for patients P = {p1,..., p n}. Next, the distances between patients are calculated to obtain a distance matrix (S m ; size m×m; square, symmetric), and classical multidimensional scaling (MDS) is used to obtain a lower-dimensional projection of this data. In the exemplary method, by specifying the dimensionality from 2 to (m - 1) and calculating the pairwise Euclidean distances between patients p1,... p m to obtain a distance matrix D (2) ,... D (m-1) the MDS analysis is performed. If the doctor in operation 56 believes that a particular patient (group or individual pair) is expected to be more similar, the pairwise distances between all possible pairs in that group are minimized. We identify the K in {2,... (m - 1)} for which this metric is the smallest. Using matrix notation:

[0040]

[0041] And

[0042]

[0043] where the matrix S m is symmetric p≠q; p = {1,..., m}; q = {1,..., m}, and the MDS function takes a distance matrix (size m×m) and the dimensionality (l; l < m). For l within the range {2,... (m - 1)}, the pairwise distances of m points are calculated to obtain a symmetric distance matrix D (l) . The groups of similar patients based on the physician's feedback are represented as G = {g1, g2...}, where g i is a set of patients from P. Then:

[0044]

[0045] And

[0046]

[0047] Here, k is an integer in {2,...,(m - 1)}, which represents placing the patient group in the lowest dimension in the closest G. Principal component analysis (PCA) or other feature reduction algorithms are used to identify the most important k features. These k features are used to cluster new patients in operation 62. Optionally, the physician notification group G is partitioned to obtain cross-validation and prevent overfitting problems.

[0048] A second exemplary method for implementing operation 60 represents the eigenvalues in the new space by adjusting the weights of the importance of these features. By way of illustration, three exemplary patients are as follows:

[0049] Patient P1 has eigenvalues (3, 2, 4, 7)

[0050] Patient P2 has eigenvalues (3, 3, 3, 3) It is used to cluster new patients in operation 62. Optionally, the physician notification group G is partitioned to obtain cross-validation and prevent overfitting problems.

[0051] Patient P3 has eigenvalues (4, 3, 3, 7)

[0052] [End]]In this notation, each patient Pi has features with values (f1, f2, f3, f4) in columns 1 to 4. For illustration, assume the following distances:

[0053] Patient distance D(P1, P3) = 3

[0054] Patient distance D(P1, P2) = 6

[0055] Patient distance D(P2, P3) = 5

[0056] In the initial clustering operation 52, using the Manhattan distance, the first cluster contains patient P1 and patient P3, and patient P2 is in the second cluster. However, in operation 56, the doctor indicates that patient P2 and patient P3 are considered more similar, perhaps because the doctor believes that features f2 and f3 are more important, and thus, the clustering is updated so that P2 and P3 are assigned to the same cluster, while P1 belongs to a separate cluster.

[0057] The centroid of the new cluster is calculated as the average of the eigenvalues in the cluster: Pc = (3.5, 3, 3, 5). Next, the original samples are mapped to the new space, where the distance from two samples to the centroid in the new space is minimized (which can be specified in advance or by the user). To adjust the coordinates to the new space, the original coordinates are multiplied by the adjusted weights (coordinates in the new space) for each feature.

[0058] To solve this problem, a set of linear equations is suitable for use. However, the number of patients n and the number of features m are usually not the same. Therefore, for the selected number of patients p, where p < n, a set of the most varying features to be mapped to the new space is derived. Symbolically represented as:

[0059] w1*f11+w2*f12+...+wp*f1p=d1

[0060] w2*f21+w2*f22+...+wp*f2p=d2 ...

[0062] wp*fp1+wp*fp2+...+wp*f2p=d2

[0063] To do this, the variance for all features is calculated, and the first p distinct features are selected. The new matrix has dimension p×p. For this new matrix, a system of linear equations is solved to find the appropriate weights. Once the weights are determined, the same weights are applied to patients who were not selected by the user—into the new space.

[0064] In the aforementioned example, this would translate to:

[0065] w1*3+w4*3=d1

[0066] w1*4+w4*7=d2

[0067] Here, we assume that w1 and w4 are weights, and that the features in columns 1 and 4 are the features with the greatest variation (for patients P1 and P2).

[0068] The foregoing are merely illustrative examples, and other methods for performing operation 60 are also envisioned. Combinations of adjustments are also envisioned, for example, performing dimensionality reduction (the first illustrative method) followed by weight adjustment (the second illustrative method); or vice versa.

[0069] refer to Figure 3 An illustrative representation of similar sample patients is shown on display component 14 (i.e., Figure 2 Operation 54), which is an exemplary example, involves querying patient P. QThe patient is "John Smith," and two similar patients identified by the last clustering iteration are "Bob Brown" and "Mickey Red." Two relevance feedback responses are requested. The first is a request for the similarity between "Bob Brown" and "John Smith" on a scale of 1-5, where "1" is the most similar and "5" is the least similar. Clinicians can use the mouse pointer to select one of the buttons labeled "1" to "5" to answer this request. The second request is to select which of the two patients, "Bob Brown" and "Mickey Red," is most similar to the queried patient "John Smith." Clinicians can answer this question by using the mouse pointer to select either the "Bob Brown" button or the "Mickey Red" button.

[0070] To meaningfully answer these requests, it should be recognized that clinicians may want to examine medical records or other patient information for the queried patient "John Smith" and for each sample patient "Bob Brown" and "Mickey Red." For this purpose, each reference for one of these patients... Figure 3 The display shows the patient's medical record as a hyperlink (e.g., indicated by emphasizing the patient's name), and the display explains: "Note: You can click on any of the patient names above to view the patient's medical record in a pop-up window." Therefore, in response to a clinician clicking on "John Smith" with the mouse pointer, a pop-up window (not shown) appears, where, preferably using suitable navigation tools, patient record information about John Smith is displayed to allow the clinician to browse John Smith's medical record. Similarly, the same applies if the mouse clicks on the patient names "Bob Brown" or "Mickey Red." Such a pop-up display may include patient characteristic information, but the clinician is able to navigate the entire patient record and is not required to assess patient similarity based on any single patient characteristic or group of patient characteristics. It should be appreciated that, for example, other navigation tool frames could be used instead of pop-ups, on separate display components (if available); Figure 1 The patient record is shown (not shown in the image).

[0071] refer to Figure 4 In other embodiments, information about the sample patient may be displayed in other ways. For example, Figure 4 The illustration shows a visualization tool in which two or more graphical modal representations are simultaneously displayed on display component 14. Figure 4The exemplary example includes three graphical modality representations displayed simultaneously: a graphical modality representation 70 for a genomic modality; a graphical modality representation 72 for a radiological modality; and a graphical modality representation 74 for a clinical modality. Each graphical modality representation 70, 72, 74 plots a waterfall plot of one or more sample patients and query patients (in the exemplary example) for two or more features of the modality. Figure 4 In the example, two sample patients, Bob Brown and Mickey Red, and the query patient, John Smith. Figure 2 In the model, the genomic modality 70 maps the patient based on ER, HER2, and PR genomic marker features. The radiological modality 72 maps the patient based on texture (roughness), volume, and morphological image features. The clinical modality 74 is used for the oncology staging modality and maps the patient based on tumor features such as tumor size (T), lymph node status (N), and metastasis value (M). Figure 4 In the study, clinicians can easily observe that, for the represented features, sample patient Bob Brown looks more similar to query patient John Smith than sample patient Mickey Red.

[0072] refer to Figure 5 , noticed, Figure 4 Visual representations are more universally applicable and can be used to navigate patient databases20 to identify patient groups through interactive graphical visualizations. Figure 5 In the exemplary example, it is shown again Figure 4 The same genomic, radiological, and clinical modalities are represented in 70, 72, and 74. In genomic modal representation 70, the GUI procedure 40 (see...) Figure 1 The selection of a patient cluster has been received (arbitrarily designated as patients {1, 2, 4, 8} by a suitable selection method, such as individually clicking on each patient in the cluster, or, in the illustrative example, by receiving the enclosing circle 80 of the patient cluster {1, 2, 4, 8} via at least one user input device (e.g., mouse 18, or trackball, touchpad, touchscreen, or other indicating device). In each of the other concurrently displayed modal graphical representations 72, 74, in response to selection 80, the patients in the selected patient cluster {1, 2, 4, 8} are also highlighted. In the illustrative example... Figure 5 In this context, the highlighting is accomplished in other modal graphical representations 72 and 74 by removing the display of all other patients so that only patients 1, 2, 4, and 8 are shown. The highlighting can also utilize other methods, such as displaying the selected cluster of patients in red and continuing to display all other patients in black.

[0073] As in Figure 5 As seen in the view, patients 1, 4, and 8 also cluster well in the radiological graphical representation 72, while patient 2 is an outlier in this modal view. In the clinical modal graphical representation 74, only patients 1 and 8 cluster together, while patients 2 and 4 are outliers. Based on these results, clinicians may be able to draw various conclusions. For example, if the queried patient is patient 1, it can be determined that patient 8 is the closest patient to the queried patient 1 (because patient 8 is close to patient 1 in all three views), while patient 2 is shown as the least dissimilar. In another view, a clinician may decide to try changing one or more of representations 72, 74 by updating two or more coordinate features—in response, the GUI process 40 redraws the updated graphical modal representation to depict the patient with the updated features of two or more coordinates for the modality.

[0074] although Figure 4 and Figure 5 The illustrations show visualization examples for genomics, radiology, and clinical modalities; however, various graphical modal representations can more generally include modalities such as clinical, radiology, genomics, demographic, and / or physiological modalities. Typically, users can select which modalities to display from a color palette or a list of available modalities, and can further select which features of each modality to plot.

[0075] The following describes the process by Figure 1 The GUI process 40 uses two or more graphical modal representations (such as...) Figure 4 and Figure 5 The representations 70, 72, and 74) appropriately execute a more detailed illustrative visual representation and navigation process. The method begins by selecting a patient for query, for example, by searching by name or electronic medical record (EMR) number. Basic information for that patient, such as name, age, attending physician, and disease, can be displayed. A user workspace containing graphical modal representations 70, 72, and 74 is displayed. Figure 4 and Figure 5 In an illustrative example, each modality representation is presented as a circle, with modality features of the patient being represented placed at equal intervals around the circle (such as biomarkers associated with a specific disease in the case of genomic modalities). Other patients are automatically incorporated into the visualization, extracted from any available cohorts (e.g., using...). Figure 1The patient group identification device generates a circle to fill the circle. This places the patients of interest within the context of a larger patient group. By default, all modalities are displayed simultaneously, but each modality can be zoomed in for individual examination. Features can be selected to be placed along the perimeter of the circle (i.e., relative to the features drawn). Subsequently, any values ​​associated with the assigned patients are highlighted across all available modalities. These selected patients can then be further analyzed.

[0076] Optionally, as the user selects a patient, a statistical summary is displayed on the screen, highlighting significant characteristics of the selected patient. This summary is dynamically updated as patient selections are updated. The content of the summary can be described based on the nature of the variables: discrete or continuous.

[0077] Given the vast amount of available demographic, pathological, clinical, and genomic features (e.g., 200 or more in some patient databases), navigation tools are provided to support the selection of features such as biomarkers, signatures, prognostic scores, and cohort samples, enabling efficient summarization and visualization of data relevant to specific contexts of interest. Optionally, the GUI tools also allow clinicians to define and save customized selections and easily switch from one context to another.

[0078] exist Figure 5 In the illustrative example, the clinician examines the ER, HER2, and PR receptor status of a selected patient in the context of other patients in a database or selected group, and has the flexibility to examine the same patient in other modal illustrative views 72, 74. The clinician can, for example, use bounding circles 80 to select / highlight a subset of patients of interest from the genomic view, and these patients are highlighted in other views 72, 74. Thus, for example, the T, N, and M stages of selected patients {1, 2, 4, 8} can be observed in the clinical (cancer staging) illustrative view 74 to assess the distribution of T (tumor size), N (nodule status), and M (metastatic status) of the selected patients. (It should be appreciated that if the number of selected patients is greater than the four illustrative examples of selected patients, a more precisely defined distribution may be obtained). Similarly, in the imaging illustrative view 72, MRI features such as volume, wash-in, wash-out characteristics, texture, and morphological features are shown. The clinician can select other modal views (not shown). In this way, clinicians are able to interactively test the correlation between features or feature groups in a selected patient group across different modalities.

[0079] As another example, a more detailed graphical view of the genome is described 70. Genome layers are displayed in circles, as shown in... Figure 5The example shown is a suitable task for assessing breast cancer. For this task, features of interest include ER, PR, and HER2 activity levels, which have demonstrated clinical utility for breast tumor diagnosis and prognosis. (Naturally, other significant genomic features will be selected for mapping other tasks.) When a clinician opens the application, it displays the patient of interest (the queried patient) along with other patients in the selected group (e.g., using...). Figure 1 (Generated by group identification devices). In illustrative Figure 5 The diagram shows three waterfall plots (bars drawn in descending order), each representing one of the ER / PR / HER2 activities, with genomic circles 70 evenly placed. In one navigation method, query patients are automatically selected to highlight the activity level of each biomarker on the query patient's circle relative to the remaining patients in the cohort. Lines (not shown) are optionally drawn from the three bars representing patients to regions in the center of the circles, where machine learning algorithms (such as principal component analysis) computed based on the cohort's ER / PR / HER2 data have been appropriately visualized. The lines are precisely drawn to the position of the patient of interest relative to the cohort. From this overall visualization, additional patients with similar ER / PR / HER2 expression levels are selected, for example, using bounding circles 80. Any additional patients selected in the machine learning space will have lines drawn around their respective ER / PR / HER2 activity levels.

[0080] You can also select a group (e.g., Figure 5 The groups {1, 2, 4, 8} in the dataset provide a statistical summary. For continuous variables such as age and gene expression values, the mean of the selected groups can be calculated. For discrete variables such as sex and ER status, enrichment analysis using tests such as the hypergeometric test can be performed, and the attributes are sorted in descending order of p-value. Table 1 shows a typical summary of selected patients for the breast cancer dataset.

[0081] Table 1 - Summary of selected patients in the breast cancer dataset

[0082] Average age 46 Average expression level of P53 2.3FPKM dominant sex Female (p-value 0.001) Dominant ER state Positive (p-value 0.003) Dominant PR status Negative (p-value 0.005) Dominant HER2 status Positive (p-value 0.007) Dominant breast cancer subtypes Baseline (p-value 0.009)

[0083] In this table, FPKM (fragment per kilobase per million exons mapped) represents the expression value of the p53 gene based on RNA sequencing data. Many of these variables are specific to the exemplary task of breast cancer diagnosis, and the statistical summary elements are appropriately described in advance in a summary format for each disease or clinical task.

[0084] Figure 4 and Figure 5The graphical visualization and navigation tools are exemplary examples. Besides the exemplary circular geometry, other geometries can be employed. The advantage of circular geometry is its ease of updating to a reasonable number of modal features relative to which it is drawn (i.e., it can comfortably fit any number of features around the circle); however, square geometry, for example, is only suitable for drawing for two modal features.

[0085] It is also envisioned that the operation of selecting patient clusters in a graphical modal representation could be performed by an entity / institution other than the clinician operating one or more user input devices 16, 18 (e.g., to perform enclosing selection 80, as in...). Figure 5 (in Chinese). For example, in Figure 4 In the exemplary example, another program executed on computers 10, 12, such as clustering process 30, selects a patient cluster as the query patient and a group of one or more sample patients (in the exemplary example). Figure 4 Among them, sample patients Bob Brown and Mickey Red, and query patient John Smith).

[0086] Return to reference Figure 1 and Figure 2 Another exemplary implementation of patient group identification using patient-level correlation feedback is described in text below and includes the following steps:

[0087] Step 1. Unsupervised learning is performed using hierarchical clustering of all patients and selected patient features on a large dataset (more than one million samples in some embodiments).

[0088] Step 2. Determine the number of clusters and calculate the cluster centroids.

[0089] Step 3. Select patients P that include the query patient based on all features. Q Clustering, and selecting additional seeds from the same cluster.

[0090] Step 4. For each seed, find the most similar patient based on the distance between that patient and all different cluster centroids as measured using the patient comparison metric.

[0091] Step 5. Select samples based on a priority list of similar patients (e.g., patients belonging to a single cluster) and samples similar to the current sample.

[0092] Step 6. Determine which features are important for patients’ similarity by removing one feature at a time.

[0093] Step 7. Use the patient comparison index to find the distance between the current patient and all selected patients.

[0094] Step 8. Find the column with a median close to 0. Discard columns with high values.

[0095] Step 9. Perform unsupervised clustering on the entire dataset based on the selected features—using only the selected clusters.

[0096] Step 10. Finally, present the original query patient P. Q The patients in the cluster, or the location where most of the selected patients appear.

[0097] Finally, repeat steps 1-10 iteratively until all samples in the group are relevant to clinicians.

[0098] The invention has been described with reference to preferred embodiments. Modifications and variations may occur to others upon reading and understanding the foregoing detailed description. The invention is intended to be construed as including all such modifications and variations, provided they fall within the scope of the appended claims or their equivalents.

Claims

1. A patient cohort identification apparatus comprising: a computer (10, 12) having a display component (14) and at least one user input device (16, 18), the computer being in communication with a patient database (20) storing patient data comprising values for features of sample patients in the patient database, the computer being programmed to perform a patient cohort identification method, the method comprising: performing (50) an automatic feature selection process (22) on the patient data to select a set of features (24), and automatically clustering (52) sample patients in the patient database using a patient comparison metric that depends on the set of features, wherein the automatic feature selection process is an unsupervised feature selection process; performing at least one iteration of: displaying (54) on the display component information about one or more sample patients that are similar or dissimilar to a query patient according to the automatic clustering, and receiving (56) via the at least one user input device a user-entered comparison value comparing the one or more sample patients to the query patient; adjusting (60) the patient comparison metric to increase agreement between comparison values comparing the one or more sample patients to the query patient as computed by the patient comparison metric and the user-entered comparison value, wherein the adjusting comprises adjusting at least one of: the set of features and feature weights of the patient comparison metric; and repeating (62) the automatic clustering using the adjusted patient comparison metric; and identifying a patient cohort for the query patient using the adjusted patient comparison metric resulting from the last iteration, wherein the displaying (54) and the receiving (56) comprise at least one of: (I) displaying a request to rank at least one sample patient on a quantitative ranking scale for similarity to the query patient, and receiving the user-entered comparison value for the sample patient as a received similarity ranking of the sample patient on the quantitative ranking scale; and (II) displaying a request to select which of two sample patients is most similar to the query patient, and receiving a user-entered comparison value as a received selection of which of the two sample patients is most similar to the query patient.

2. The patient cohort identification device of claim 1, wherein, the identifying comprises: identifying the patient cohort as at least part of a cluster (34) containing the query patient generated by the last repetition (62) of the automatic clustering.

3. The patient cohort identification device of claim 1 or 2, wherein, the displaying (54) and the receiving (56) comprise: (I) displaying on the display component (14) information about one or more similar sample patients belonging to a cluster (34) also containing the query patient generated by a most recently performed automatic clustering; or (II) displaying on the display component (14) information about one or more dissimilar sample patients not belonging to a cluster (34) containing the query patient generated by a most recently performed automatic clustering.

4. The patient cohort identification device of claim 1 or 2, wherein, the displaying (54) comprises: simultaneously displaying two or more graphical modality representations (70, 72, 74), wherein each graphical modality representation plots the one or more sample patients and the query patient for two or more features of the modality.

5. The patient cohort identification device of claim 4, wherein, the two or more graphical modality representations (70, 72, 74) include graphical modality representations for modalities selected from the group consisting of: a clinicality modality, a radiology modality, a genomics modality, a demographics modality, and a physiology modality.

6. The patient cohort identification device of claim 1 or 2, wherein, the adjusting (60) includes: (I) performing a plurality of feature set adjustment iterations, each feature set adjustment iteration including: (1) adjusting the set of features by adding or removing features to produce a candidate adjusted set of features; (2) computing a comparison value using the patient comparison metric with the candidate adjusted set of features, the comparison value comparing the one or more sample patients to the query patient; (3) accepting or rejecting the candidate adjusted set of features based on whether the comparison value computed in operation (2) is in agreement with the user-input comparison value is increased or decreased, respectively; or (II) performing dimensionality reduction to reduce the number of features in the feature set.

7. The patient cohort identification device of claim 1 or 2, wherein, the adjusting (60) includes adjusting a feature weight of the patient comparison metric.

8. The patient cohort identification device of claim 7, wherein, the adjusting (60) includes performing a plurality of feature weight adjustment iterations, each feature weight adjustment iteration including: (1) adjusting the patient comparison metric by increasing or decreasing a value of at least one feature weight of the patient comparison metric to produce a candidate adjusted patient comparison metric; (2) computing a comparison value using the candidate adjusted patient comparison metric, the comparison value comparing the one or more sample patients to the query patient; and (3) accepting or rejecting the candidate adjusted patient comparison metric based on whether the comparison value computed in operation (2) is in agreement with the user-input comparison value is increased or decreased, respectively.

9. The patient cohort identification device of claim 1 or 2, wherein, the automatic feature selection process (22) is one of: principal component analysis (PCA), information gain (IG), and pairwise feature correlation.

10. A patient cohort identification device, comprising: a computer (10, 12) having a display component (14) and at least one user input device (16, 18), the computer being in communication with a patient database (20) storing patient data including values for features for sample patients in the patient database, the computer being programmed to perform a patient cohort identification method, the method including: performing an automatic feature selection process on the patient data to select a set of features, wherein the automatic feature selection process is an unsupervised feature selection process; simultaneously displaying two or more graphical modality representations (70, 72, 74) on the display component, wherein each graphical modality representation plots patients in the database for two or more coordinate features of the modality; receiving a selection (80) of a cluster of patients in one graphical modality representation (70); and receiving a selection (82) of a cluster of patients in another graphical modality representation (72). In response to receiving the selection, highlighting the patients in the selected cluster of patients in other concurrently displayed one or more graphical modal representations (72, 74).

11. The patient cohort identification device of claim 10, wherein, The receiving includes: (I) receiving the selection (80) of the cluster of patients via the at least one user input device (16, 18) operating on the one graphical modal representation (70); or (II) receiving the selection of the cluster of patients from another computer program running on the computer (10, 12).

12. The patient cohort identification device of claim 11, wherein, The selection includes receiving a lasso (80) of the cluster of patients via the at least one user input device comprising one of: a mouse (18), trackball, trackpad, touchscreen, or other pointing device.

13. The patient cohort identification device of claim 10 or 11, wherein, The patient cohort identification method further includes: receiving an updated selection of the two or more coordinate features for one of the concurrently displayed graphical modal representations (70, 72, 74), wherein the graphical modal representation is updated to plot the patients in the database for the updated two or more coordinate features of the modality.

14. The patient cohort identification device of claim 10 or 11, wherein, The two or more graphical modal representations (70, 72, 74) include graphical modal representations for modalities selected from the group consisting of: a clinical science modality, a radiology modality, a genomics modality, a demographics modality, and a physiology modality.

15. A patient cohort identification method performed in coordination with a computer (10, 12) having a display component (14) and at least one user input device (16, 18) and in communication with a patient database (20) storing patient data including values for features of sample patients in the patient database, the patient cohort identification method comprising: performing (52) an automatic clustering of sample patients in the patient database using a patient comparison metric dependent on a set of features (24), wherein the automatic clustering includes an unsupervised feature selection process; performing at least one iteration of: displaying (54) on the display component information about one or more sample patients similar or dissimilar to a query patient according to the automatic clustering and receiving (56) via the at least one user input device a user-entered comparison value comparing the one or more sample patients to the query patient; adjusting (60) at least one of: the set of features of the patient comparison metric and a feature weight to generate an adjusted patient comparison metric having improved agreement with the user-entered comparison value compared to the patient comparison metric without the adjustment; and repeating (62) the automatic clustering using the adjusted patient comparison metric; and identifying a patient cohort for the query patient as at least part of a cluster (34) containing the query patient resulting from the repetition of the automatic clustering by the last iteration, wherein the displaying (54) and the receiving (56) include at least one of: (I) displaying a request to rank the sample patients on a quantitative ranking scale for similarity to the query patient, and receiving a ranking of the sample patients on the quantitative ranking scale for similarity; and (II) displaying a request to select which one of two sample patients is most similar to the query patient, and receiving a selection of which one of the two sample patients is most similar to the query patient.

16. The patient cohort identification method of claim 15, wherein, The displaying (54) includes: simultaneously displaying two or more graphical modal representations (70, 72, 74), wherein each graphical modal representation plots the one or more sample patients and the query patient for two or more features of the modality. The displaying (54) includes: simultaneously displaying two or more graphical modal representations (70, 72, 74), wherein each graphical modal representation plots the one or more sample patients and the query patient for two or more features of the modality.

Citation Information

Patent Citations

  • Retrieval of similar patient cases based on disease probability vectors

    CN101911078A

  • Deep learning-based clustering method

    CN103530689A

  • System and method for patient identification for clinical trials using content-based retrieval and learning

    US20050210015A1

  • Systems and methods for holistic analysis and visualization of pharmacological data

    US20120078522A1

  • Iterative Refinement of Cohorts Using Visual Exploration and Data Analytics

    US20140108379A1