Systems and methods for transforming and storing data from multiple studies
By standardizing and combining data from multiple studies, the method addresses the challenge of diverse data formats, enabling effective training of a machine learning algorithm for predicting treatment outcomes.
Patent Information
- Application Number
- JP2025507095
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-05
- Filing Date
- 2023-08-04
- Publication Date
- 2025-08-26
AI Technical Summary
Medical studies for the same medical condition often collect different data formats, making it difficult to combine data from multiple studies to train a machine learning algorithm for predicting treatment outcomes.
Convert questions from different studies into a standardized format, categorize them, group compatible questions, rescale answers, and combine them to create a unified dataset for training a machine learning algorithm.
Enables the training of a machine learning algorithm using data from multiple studies, improving the prediction of treatment outcomes by standardizing and combining diverse data formats.
Smart Images

Figure 2025528114000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 395,381, filed August 5, 2022, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Medical studies for the same medical condition may collect different data from the patients in the study. Studies may have patients complete different clinical questionnaires. Data from different medical studies may be used to compare different treatment options. It may be desirable to train a machine learning algorithm to predict the outcomes of the different treatment options used in the studies. Because the studies use different methods to collect data on patients, it may not be possible to combine data from multiple studies to train a machine learning algorithm. Summary of the Invention [Means for solving the problem]
[0003] Multiple studies may relate to the same or similar medical conditions. Different studies may study different treatments and / or treatment protocols for the medical condition. Each study collects different information about the patients who participate in the study. Each study may collect similar and / or overlapping data, but this data may be collected in different formats.
[0004] To gather information about patients participating in a study, patients may be given a clinical questionnaire. Different studies may use different questionnaires and / or different formats of questionnaires. For example, patients enrolled in one study may be given a short version of the questionnaire, while patients enrolled in a second study may be given a long version of the questionnaire. Even within a single study, different questionnaires and / or different formats of questionnaires may be used.
[0005] A machine learning algorithm (MLA) may be trained using data from different studies. To prepare the data for use in training the MLA, data from different studies may be converted into a common format. Identical or nearly identical questions may be merged into a standardized format. Each question may be categorized, and similar questions may be combined.
[0006] Data from multiple studies can be obtained and then used to generate a dataset that can be used to train the MLA. A set of data points can be extracted from each study, with each data point containing data corresponding to a single patient. Each data point can include various patient attributes extracted from the study data. The attributes can include data collected using a clinical questionnaire. Individual questions from the questionnaire can be extracted from the study.
[0007] Standardized questions can be created based on questions from a questionnaire used in a study. Each question asked in a study can be converted into one of the standardized questions. A taxonomy with various categories can be defined. Each standardized question can be assigned one or more categories from the taxonomy. Standardized questions can also be assigned flags.
[0008] Compatible questions may be identified because they are assigned the same or similar categories. Compatible questions may be grouped and combined. Responses to the grouped questions may be rescaled to a common scale. The scale may be categorical or binary. Whether the scale of a question is categorical or binary may be determined based on the format of the questions being combined. After the questions in the group are rescaled to a common scale, the grouped questions may be combined.
[0009] Study data may have some missing data. For example, data points in a study may have missing attribute values. This missing data may be imputed. If the amount of data missing from a study exceeds a predefined threshold, individual attributes with a high amount of missing data may be excluded from the study, or the entire study may be excluded.
[0010] After combining and rescaling the questions, a final dataset may be generated. The final dataset may be divided into a training set, a validation set, and / or a test set. The training set may be used to select features for the MLA. The features may be one or more categories from a classification, attributes from research data, and / or responses to a questionnaire.
[0011] The MLA can be trained using features selected from the training set. After training, the MLA can be trained using a validation set. The MLA can then be tested using a test set.
[0012] According to a first broad aspect of the present technology, there is provided a method that includes obtaining a plurality of studies, each study including a plurality of data points, each data point including a plurality of attributes; extracting a plurality of questions from the plurality of studies; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of the plurality of questions based on the one or more categories assigned to each question; determining an answer scale to use for each group of the plurality of groups; rescaling answers to the groups of questions based on the group's respective answer scale; combining questions for each group of the plurality of groups; imputing missing attribute values in the plurality of studies; generating a final dataset by combining the plurality of studies;
[0013] In some implementations of this method, the plurality of questions includes questions from a clinical questionnaire administered to patients enrolled in a clinical trial.
[0014] In some implementations of this method, each data point corresponds to a patient enrolled in a clinical trial, and the plurality of attributes includes an indication of a treatment given to the patient.
[0015] In some implementations of this method, the plurality of attributes includes an indication of responses to a clinical questionnaire administered to the patient.
[0016] In some implementations of this method, the plurality of attributes includes an indication of whether a treatment is effective for the patient.
[0017] In some implementations of this method, the plurality of attributes includes an indication of side effects experienced by the patient.
[0018] In some implementations of this method, the plurality of attributes includes physiological data of the patient.
[0019] In some implementations of this method, the plurality of attributes includes historical patient data.
[0020] In some implementations of this method, the plurality of attributes includes socio-demographic data of the patient.
[0021] In some implementations of this method, the plurality of attributes includes psychological data of the patient.
[0022] In some implementations of this method, rescaling the answers to the group of questions includes using equipercentile scaling to rescale the answers.
[0023] According to another broad aspect of the present technology, there is provided a method including obtaining a plurality of studies, each study including a plurality of data points, each data point including a plurality of attributes; extracting a plurality of questions from the plurality of studies; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of questions from the plurality of questions that are assigned the same set of categories; combining the groups of questions; generating a final dataset by combining the plurality of studies; selecting a plurality of features in the final dataset; and training a machine learning algorithm (MLA) using the plurality of features.
[0024] In some implementations of this method, combining the groups of questions includes determining an answer scale for the groups of questions and rescaling the answers corresponding to the groups of questions based on the answer scale.
[0025] According to a further broad aspect of the present technology, there is provided a method for predicting treatment efficacy for a patient, the method including receiving questionnaire responses from the patient; inputting the questionnaire responses into an MLA trained according to any of the methods described herein; outputting, by the MLA, a predicted efficacy of each of a plurality of treatments; generating an interface based on the predicted efficacy of each of the plurality of treatments; and outputting the interface for display.
[0026] According to another broad aspect of the present technology, there is provided a trained MLA, trained according to any of the methods described and claimed herein.
[0027] According to another broad aspect of the present technology, there is provided a method for a patient (e.g., for predicting treatment efficacy) that includes obtaining a plurality of study results corresponding to a plurality of treatments, each study result including a plurality of data points, each data point including a plurality of attributes; extracting a plurality of questions from the plurality of study results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of questions from the plurality of questions that are assigned the same set of categories; combining the groups of questions; generating a final dataset by combining the plurality of study results; and training the MLA using the final dataset; receiving questionnaire responses from the patient; inputting the questionnaire responses into the MLA; outputting, by the MLA, a predicted efficacy of each of the plurality of treatments; generating an interface based on the predicted efficacy of each of the plurality of treatments; and outputting the interface for display.
[0028] Various implementations of the present technology provide computer-based systems, such as, for example, but not limited to, electronic devices, including at least one processor and memory storing program instructions for performing one or more methods described herein, the program instructions being executable by the at least one processor of the electronic device.
[0029] It is expressly understood that not all technical advantages described herein need be enjoyed in every embodiment of the present technology.
[0030] As used herein, the term "and / or" is intended to represent an inclusive or, for example, "X and / or Y" is intended to mean either X or Y, or both. As a further example, "X, Y, and / or Z" is intended to mean X, Y, Z, or any combination thereof.
[0031] In the context of this specification, unless expressly specified otherwise, a computer system or computing environment may refer to, but is not limited to, an "electronic device," "computing device," "operation system," "system," "computer-based system," "computer system," "network system," "network device," "controller unit," "monitoring device," "control device," "server," and / or any combination thereof appropriate for the relevant task at hand.
[0032] In the context of this specification, unless expressly specified otherwise, any method and / or system described herein may be implemented in a cloud-based environment, such as, but not limited to, a Microsoft Azure environment, an Amazon EC2 environment, and / or a Google Cloud environment.
[0033] In the context of this specification, unless expressly stated otherwise, the terms "computer-readable medium" and "memory" are intended to include media of all natures and types, non-limiting examples of which include RAM, ROM, disks (e.g., CD-ROMs, DVDs, floppy disks, hard disk drives, etc.), USB keys, flash memory cards, solid-state drives, and tape drives. Furthermore, in the context of this specification, "a" computer-readable medium and "the" computer-readable medium should not be construed as the same computer-readable medium. Conversely, whenever appropriate, "a" computer-readable medium and "the" computer-readable medium may be construed as a first computer-readable medium and a second computer-readable medium.
[0034] In the context of this specification, unless expressly stated otherwise, the words "first," "second," "third," etc. are used as adjectives only to distinguish between the nouns they modify, and not to describe any particular relationship between those nouns.
[0035] Additional and / or alternative features, aspects, and advantages of implementations of the present technology will become apparent from the following description, the accompanying drawings, and the appended claims.
[0036] For a better understanding of the present technology, as well as other aspects and further features thereof, reference is made to the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0037] [Figure 1] FIG. 1 is a block diagram of an exemplary computing environment in accordance with various embodiments of the present technology. [Figure 2] FIG. 1 is a diagram of a process for combining research results to train a machine learning algorithm (MLA), according to various embodiments of the present technology. [Figure 3]FIG. 1 is a diagram of classification according to various embodiments of the present technology. [Figure 4] FIG. 1 is a flow diagram of a method for generating a combined dataset, in accordance with various embodiments of the present technology. [Figure 5] FIG. 1 is a flow diagram of a method for generating a combined dataset, in accordance with various embodiments of the present technology. DETAILED DESCRIPTION OF THE INVENTION
[0038] The examples and conditional language described herein are primarily intended to aid the reader in understanding the principles of the technology, and are not intended to limit its scope to such specifically described examples and conditions. Those skilled in the art will appreciate that various configurations may be devised that, although not explicitly described or illustrated herein, nevertheless embody the principles of the technology and are included within its spirit and scope.
[0039] Furthermore, to aid in understanding, the following description may describe a relatively simplified implementation of the technology. As those skilled in the art will appreciate, various implementations of the technology may be more complex.
[0040] In some cases, what are believed to be useful modifications to the technology may also be described. This is done merely to aid in understanding and, again, does not define the scope of the technology or indicate limitations. These modifications are not an exhaustive list, and one of ordinary skill in the art may make other modifications while still remaining within the scope of the technology. Furthermore, if an example of a modification is not described, it should not be construed as meaning that the modification is not possible and / or that what is described is the only way to practice that element of the technology.
[0041] Furthermore, all statements herein describing principles, aspects, and implementations of the present technology, as well as specific examples thereof, are intended to encompass both structural and functional equivalents, whether now known or developed in the future. Thus, for example, those skilled in the art will appreciate that any block diagrams herein represent conceptual diagrams of illustrative circuitry embodying the principles of the present technology. Similarly, it will be appreciated that any flowcharts, flow diagrams, state transition diagrams, pseudocode, and the like may be substantially represented on a computer-readable medium and represent various processes that may be so executed by a computer or processor, whether or not such a computer or processor is explicitly shown.
[0042] The functions of the various elements shown in the figures, including any functional block labeled as a "processor," may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors, some of which may be shared. In some embodiments of the present technology, the processor may be a general-purpose processor such as a central processing unit (CPU), or a processor dedicated to a specific purpose, such as a digital signal processor (DSP). Furthermore, the explicit use of the term "processor" should not be construed to refer solely to hardware capable of executing software, but may implicitly include, but is not limited to, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage. Other hardware, conventional and / or custom, may also be included.
[0043] Software modules, or modules that simply imply software, may be represented herein as any combination of flowchart elements or other elements illustrating the execution of process steps, and / or textual descriptions. Such modules may be performed by explicitly or implicitly shown hardware. Furthermore, it should be understood that one or more modules may include, for example, but not limited to, computer program logic, computer program instructions, software, stacks, firmware, hardware circuits, or combinations thereof.
[0044] Computing Environment FIG. 1 illustrates a computing environment 100 that may be used to implement and / or execute any of the methods described herein. In some embodiments, computing environment 100 may be implemented by any of a conventional personal computer, a computer dedicated to managing network resources, a network device and / or electronic device (e.g., without limitation, a mobile device, a tablet device, a server, a controller unit, a control device, etc.), and / or any combination thereof suitable for the relevant task at hand. In some embodiments, computing environment 100 includes various hardware components, including one or more single-core or multi-core processors collectively represented by processor(s) 110, a solid-state drive 120, random access memory 130, and input / output interfaces 150. Computing environment 100 may be a computer specially designed for running machine learning algorithms (MLAs). Computing environment 100 may also be a general-purpose computer system.
[0045] In some embodiments, computing environment 100 may be a subsystem of one of the systems listed above. In other embodiments, computing environment 100 may be a "commercially available" general-purpose computer system. In some embodiments, computing environment 100 may be distributed among multiple systems. Additionally, computing environment 100 may be dedicated to implementing the present technology. As will be appreciated by those skilled in the art, numerous variations on how computing environment 100 is implemented may be envisioned without departing from the scope of the present technology.
[0046] Those skilled in the art will appreciate that processor 110 generally refers to processing power. In some embodiments, one or more specialized processing cores may be provided instead of, or in addition to, one or more conventional central processing units (CPUs). For example, one or more graphics processing units (GPUs), tensor processing units (TPUs), and / or other so-called accelerated processors (or processing accelerators) may be provided in addition to, or instead of, one or more CPUs.
[0047] System memory typically includes random access memory 130, but is more generally intended to encompass any type of non-transitory system memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or a combination thereof. While solid-state drive 120 is shown as an example of a mass storage device, more generally, such mass storage may include any type of non-transitory storage device configured to store data, programs, and other information and to make the data, programs, and other information accessible via system bus 160. For example, mass storage may include one or more of a solid-state drive, a hard disk drive, a magnetic disk drive, and / or an optical disk drive.
[0048] Communication between the various components of the computing environment 100 may be enabled by a system bus 160, which may include one or more internal and / or external buses (e.g., PCI bus, Universal Serial Bus, IEEE 1394 "Firewire" bus, SCSI bus, Serial ATA bus, ARINC bus, etc.), to which the various hardware components are electronically coupled.
[0049] The input / output interface 150 may enable network functionality such as wired or wireless access. By way of example, the input / output interface 150 may comprise a networking interface such as, but not limited to, a network port, a network socket, or a network interface controller. Examples of how a networking interface may be implemented will be apparent to those skilled in the art. For example, the networking interface may implement a particular physical layer and data link layer standard, such as Ethernet, Fibre Channel, Wi-Fi, Token Ring, or a serial communications protocol. The particular physical layer and data link layer may provide the basis for a complete network protocol stack, enabling communication between small groups of computers on the same local area network (LAN) and large-scale network communication via routable protocols such as the Internet Protocol (IP).
[0050] Input / output interface 150 may be coupled to touchscreen 190 and / or one or more internal and / or external buses 160. Touchscreen 190 may be part of a display. In some embodiments, touchscreen 190 is a display. Touchscreen 190 may also be referred to as screen 190. In the embodiment shown in FIG. 1 , touchscreen 190 comprises touch hardware 194 (e.g., pressure-sensing cells embedded in a layer of the display that enable detection of physical interaction between a user and the display) and a touch input / output controller 192 that enables communication with display interface 140 and / or one or more internal and / or external buses 160. In some embodiments, input / output interface 150 may be connected to a keyboard (not shown), a mouse (not shown), or a trackpad (not shown) in addition to or instead of touchscreen 190 to enable a user to interact with computing device 100.
[0051] According to some implementations of the present technology, the solid-state drive 120 stores program instructions suitable for being loaded into the random access memory 130 and executed by the processor 110 to perform acts of one or more methods described herein. For example, at least some of the program instructions may be part of a library or an application.
[0052] Combining research findings 2 is a diagram of a process for combining study results to train a machine learning algorithm (MLA) according to various embodiments of the present technology. Study results 201-204 may be obtained. Each study may be testing the effects of a different treatment and / or treatment protocol for the same medical condition. For example, study result 201 may be a study result of a compound used to treat depression, and study result 202 may be a study result of a different compound used to treat depression.
[0053] Although described as research results 201-04, it should be understood that the methods described herein can be used to combine any type of dataset having common data, which may be in different formats. The datasets to be combined may be clinical trial results and / or any other type of dataset. The data to be combined may include any type of patient data, such as electronic medical records (EMRs). For example, patient demographics, treatment information about the patient, and outcome data corresponding to the treatment may be extracted from a patient's EMR and then combined with data extracted from other patients' EMRs.
[0054] Each set of study results 201-04 may include a set of data points for individual patients involved in each study. Each data point corresponds to an individual patient enrolled in the study. The data points may include various attributes. The attributes of a data point corresponding to a patient may include the patient's responses to a clinical questionnaire, information about the patient, information about the treatment administered to the patient, outcome data indicative of the patient's response to the treatment, the patient's medical history, the patient's family history, physiological information about the patient, and / or any other data collected about the patient.
[0055] The outcome data for the data points may include whether the treatment resulted in remission, time to remission, whether the treatment caused harm and / or had adverse side effects, whether the treatment resolved a particular symptom, whether the treatment resulted in a return to baseline physiological measurements, and / or any other measure of the effectiveness of the treatment.
[0056] The different studies that produced the findings may have used the same and / or different questionnaires. Even if different questionnaires or different formats of questionnaires were used, the different questionnaires may have similar questions. For example, the questionnaire used in one study and the questionnaire used in a second study may ask different questions about the same topic.
[0057] The clinical questionnaire may include questions about the patient's mental health status, medical history, family medical history, current medications, sociodemographic information, and / or any other type of question. Patients may be asked to update and / or complete new clinical questionnaires periodically to ensure that the information collected about them is current. Thus, individual data points in the study results 201-04 may include multiple responses to the same questionnaire at different times. Patients may be asked to update their clinical questionnaires after a predetermined period of time has elapsed. Clinical questionnaires may be completed by the patient, the patient's caregiver, and / or a physician.
[0058] Standardized versions of questions from study outcomes 201-04 may be generated. Similar, but not identical, questions may be asked in different studies and / or different questionnaires. Similar questions may be grouped and replaced with a standardized version of that question in the transformed outcomes 211-14.
[0059] The taxonomy of categories may be generated and / or predefined. A category may be a category of a questionnaire question. The taxonomy is organized in a tree structure, with each category having a parent node and / or a child node. FIG. 3 shows an example of a taxonomy. In the taxonomy shown in FIG. 3, the "Dysfunction" category is a branch of the root category of "Quality of Life." The "Dysfunction" category further has a leaf and another sub-branch, "Relationships," which itself has leaves, named "Family," "Social," and "Romantic."
[0060] Each question in the research findings 201-04 may be assigned one or more categories in the taxonomy. Questions may be assigned categories based on keywords in the question. Categories may be assigned to questions using an MLA, based on manual review, and / or any other suitable method. The MLA may be trained to assign categories to questions based on the question text. A question is input into the MLA, and the MLA may output one or more categories that it predicts correspond to the question. A human operator can then verify that the predicted categories correspond to the question and / or edit the categories assigned to the question. Other methods, such as extracting individual words from the question, may also be used to suggest categories to apply to a question.
[0061] After standardizing and categorizing the questions, questions that are compatible with each other may be grouped and combined. Compatible questions may be semantically similar to each other. Groups of questions assigned the same and / or similar categories may be formed. For each group of questions, a common answer format may be selected. Answers to questions within the group may be rescaled to the selected format. After rescaling the answers, the grouped questions may be combined.
[0062] The study results 201-04 may be converted into transformed results 211-14. In the transformed results 211-14, each question may be in a standardized format. Similar questions may be converted into a combined version of the question.
[0063] The transformed results 211-14 from each study may then be combined into a final dataset 215. Because the study results 201-04 have been converted into a common format in the transformed results 211-14, the data may potentially be combined into the final dataset 215. The final dataset 215 may include data points from each study result 201-04. Each data point in the final dataset 215 may be associated with an individual patient. Each data point may include attributes indicative of the standardized and transformed question, the treatment the patient received, and an indication of the effectiveness of the treatment for the patient.
[0064] The final dataset 215 may be divided into a training set 220, a validation set 221, and a test set 222. The training set 220, the validation set 221, and the test set 222 may each include a subset of the data points in the final dataset 215. Each data point in the final dataset 215 may be randomly assigned to the training set 220, the validation set 221, or the test set 222. The training set 220, the validation set 221, and the test set 222 may have predetermined sizes. For example, the training set 220 may be configured to be larger than the validation set 221.
[0065] After generating training set 220, training set 220 can be used to train MLA 225. Each data point in training set 220 can be input to MLA 225. MLA 225 can predict the likelihood of remission for a patient corresponding to the data point if the patient receives various treatments. For each potential treatment, MLA 225 can output a predicted likelihood of remission.
[0066] The MLA 225 may include one or more neural networks and / or any other type of machine learning, deep learning, and / or artificial intelligence (AI) model. The MLA 225 may include an encoder whose input x does not include the patient's assigned treatment, but is responsible for encoding features corresponding to the patient into some latent space e(x). The decoder can decode the encoded features back to the original input d(e(x)). The decoded features may not be identical to the original features that were encoded. Any suitable type of encoder and / or decoder may be used, such as a symmetric autoencoder.
[0067] The loss function can compare the predicted effectiveness of a treatment to the actual outcomes recorded during the study, which are stored in the training set 220. The actual outcomes may be referred to as the "labels" of the data points, and the MLA 225 can be trained to predict the labels. The MLA 225 can then adjust itself based on the amount of loss output by the loss function. The MLA 225 can be adjusted to reduce the amount of loss between the predictions and the labels. This training process can be repeated for all data points in the training set 220.
[0068] After the MLA 225 is trained using the training set 220, the MLA 225 can be tuned using the validation set 221. Each data point in the validation set 221 can be input to the MLA 225. The MLA 225 can then output a prediction result that can be compared to the corresponding label using a loss function. The parameters of the MLA 225 can be tuned based on the results of the loss function.
[0069] After training the MLA 225 using the validation set 221, a test set 222 can be input into the MLA 225 to determine how accurate the predictions of the MLA 225 are. Various methods can be used to determine whether the MLA 225 is accurate enough for further use. The predictions output by the MLA 225 can be compared to the results of the test set 222. An average loss amount for the MLA 225 on the test set 222 can be determined. If the average loss amount is less than a predetermined threshold loss amount, the MLA 225 can be determined to be ready for use.
[0070] To use the MLA 225 after training, a clinical questionnaire may be administered to the patient. The answers to the questions may be encoded into a vector of numbers using an encoder function e(x). The vector may then be input to the MLA 225.
[0071] prototype Various prototypes can be defined that correspond to clusters of patients. Each prototype may correspond to a group of patients with similar characteristics, exhibiting similar symptoms, and / or responding similarly to one or more treatments. Prototypes can be defined such that each prototype responds differently to available treatments. The prototypes can assist clinicians and / or patients in understanding the results output by the MLA 225. In other words, the prototypes can be used to increase the interpretability of the results for clinicians and / or patients. Each prototype can be used to generate an example patient corresponding to the prototype in order to compare actual patients with the prototype.
[0072] Training the MLA 225 may involve layers of a neural network that form the MLA 225 to extract these prototypes. Each prototype may indicate the importance of features in predicting patient remission and / or the differential effects of different treatments on a given prototype. Each prototype may be associated with a patient cluster, representing a group of patients that are relatively similar to the trained prototype. Extracting prototypes may improve the accuracy of the MLA 225 and / or improve the interpretability of the MLA 225. Extracting prototypes may help clinicians understand the output of the MLA 225 by showing how different feature clusters, representing different patient prototypes, respond to different treatments.
[0073] The number of prototypes to define can be determined empirically (through human / non-human initialization and experimental progression) and / or dynamically (by algorithmic determinism to optimize downstream objectives). The number of prototypes can be selected based on various considerations, such as increasing the interpretability and / or accuracy of the prototypes. For example, setting the number of prototypes to three can strike a balance between providing sufficient nuance between prototypes and providing a sufficiently accurate MLA225.
[0074] In some cases, prototypes may be defined in the original feature space without the use of an autoencoder, but then encoded into a latent space by an autoencoder for compatibility in comparison with already encoded features. Prototypes may be defined manually by an operator and / or automatically using various functions such as clustering algorithms. For example, an operator may input various parameters of the prototype.
[0075] Considering the symmetry of neural networks, both the encoder and decoder may include the same number of fully connected layers. The output of the encoding layer e(x) may be fed to a prototype layer p, which may be composed of k nodes to represent each prototype individually. The variable k may represent the number of patient archetypes that the prototype may learn to represent individually. Each node may be the size of the input data sample. The prototypes may be defined in a latent (encoded) space.
[0076] To compare a patient's data with a set of prototypes, both can be mapped to the encoded space. The prototypes may have learned parameters that can be configured to shift around the encoded feature space to achieve optimal downstream predictive performance of the MLA. The prototypes can be assigned "frozen" weights that can ensure they remain static over the duration of MLA training.
[0077] To enable an operator, such as a clinician, to interpret the prototype, the prototype can be decoded by a decoder d(p). The decoder can extract the original feature values corresponding to the prototype. A subject matter expert, such as a clinician, may view the original feature values of the prototype to better understand its relationship to the prototype and the probability of predicted treatment effectiveness.
[0078] Once the patient data is input into the MLA 225, an autoencoder can be used to calculate the distance between the patient and each prototype in the latent space. These distances can then be passed on to downstream prediction purposes.
[0079] Prototype configuration When defining prototypes, various hyperparameters can be configured, including (1) the number of prototypes that MLA225 supports and (2) the tunable parameters of each prototype.
[0080] Any number of prototypes may be defined. For purposes of improving interpretability, it may be desirable to have a relatively small number of prototypes, such as two, three, or four prototypes, since too many prototypes may make it difficult for a clinician to understand and / or explain why one treatment may benefit a patient over another. From a performance perspective, the number of prototypes may also be configured to optimize downstream objectives, such as predicting treatment remission rates. The operator may select the number of prototypes to define to balance interpretability with the overall performance of the MLA225.
[0081] The parameters of the prototype can be defined in a variety of ways, such as based on input from an operator and / or automatically using a function. The operator can define the parameters of the prototype. A prototype can then be generated based on the parameters defined by the operator. The parameters of the prototype can be generated using a function, such as a clustering algorithm.
[0082] Previously identified stereotypical patient clusters can be used as the basis for prototypes. An operator can define parameters corresponding to the previously identified clusters, and the previously identified clusters can be converted into prototypes by passing those parameter values to an encoder e(x) to initialize the prototype.
[0083] Algorithm initialization can be used to generate a prototype, which can be initialized using a Xavier-Glorot uniform / normal distribution, a He (i.e., Kaiming) uniform / normal distribution, or a normal or uniform distribution, or any other existing or custom distribution that can sample a set of parameters from a set of continuous or discrete values.
[0084] Prototype Output The output of the prototype layer may represent the distance between the patient's coded data e(x) and each prototype p. In other words, the output may indicate the distance between the patient and each prototype. This distance may be defined by the Frobenius norm between the coded sample and each prototype individually. Any other suitable distance measure may also be used, such as variance-based distance (assuming each prototype represents a statistically distinct sample distribution), Mahalanobis distance, or modeling each prototype cluster to a normal distribution to identify which patient samples are most likely to fall within a certain standard deviation from the cluster center. For example, if a sample is more likely to be within one standard deviation of a first cluster than within three standard deviations of a second cluster, the distance may reflect the degree of overlap.
[0085] The latent distances from the patient to the prototypes can be fed into a fully connected neural layer, which concatenates them with the omitted assigned treatments before passing the patient information to the encoder. The treatments can be one-hot encoded before being concatenated with the remainder of the distance vector. This concatenation can be fed into a final classification layer, whose goal is to extract the probability of remission for each assigned treatment and test a hypothetical case for each patient. These predicted remission rates for each treatment for each patient can then be aggregated and used to calculate a margin of error.
[0086] Methods for generating a combined data set 4 and 5 show a flow diagram of a method 400 for generating a combined dataset in accordance with various embodiments of the present technology. In one or more aspects, the method 400, or one or more steps thereof, may be performed by a computing system such as the computing environment 100. The method 400, or one or more steps thereof, may be embodied in computer-executable instructions stored in a computer-readable medium, such as a non-transitory mass storage device, loaded into a memory, and executed by a CPU. Some steps or portions of steps within the flow diagram may be omitted or reordered.
[0087] In step 405, data may be acquired from multiple studies, such as study results 201-04. As described above, the studies may be clinical studies testing various treatments for medical conditions. The acquired data may be de-identified patient-level data from clinical trials for the treatment of disorders such as depression. The data may include associated study protocols corresponding to the studies.
[0088] Patients enrolled in the studies may have completed clinical questionnaires. These studies may use different questionnaires and / or different formats of questionnaires. Different patients may have been given different questionnaires and / or different formats of questionnaires. The data obtained in step 405 may include responses to the questionnaires given by the patients. The data obtained in step 405 may include results of the study for each patient. The results of the study for each patient may include an indication of whether the treatment was effective for that patient and / or a measure of how effective the treatment was for that patient. The data obtained in step 405 may include an indirect measure of the effectiveness of the treatment, such as a quality of life assessment, scaled functional disability, and / or any other indirect measure.
[0089] In step 410, questionnaires may be extracted from each of the study outcomes 201-04. In some cases, a single questionnaire may be extracted from one of the study outcomes 201-04. In other cases, multiple questionnaires may be extracted from one of the study outcomes 201-04. The extracted questionnaires may include individual questions asked to study participants. Each question may be associated with information about available answers. The information about available answers may indicate whether participants could provide a binary or categorical answer to the question. A binary answer is one in which the patient could select "yes" or "no," "true" or "false," or any other binary answer. A categorical answer is one in which the patient can select from multiple answers and / or from answers on a scale, such as a Likert scale. An example of a question with a categorical answer is one in which the patient is asked to provide a rating from 1 to 5. Information regarding available responses to categorical questions may include the scale to which the response corresponds, information regarding the options available to the patient to respond, and / or information indicating the meaning of the highest and lowest values on the scale.
[0090] In step 415, standardized questions may be created. Each question extracted in step 410 may be converted to a standardized format. A standard question statement may be created for each question. Different questionnaires may include the same question in slightly different formats and / or the questions may be in a different order. A single standard version of a question may be selected, and all other iterations of that question may conform to the standard version.
[0091] Different studies may use slightly different versions of the same questionnaire. Different versions of a questionnaire may contain similar sets of questions, but may have slightly different wording and / or a different order of questions. A single version of the questionnaire may be selected as the standard version of the questionnaire. Responses to all other iterations of the questionnaire used may be modified to match the standard version of the questionnaire.
[0092] At step 420, each of the standardized questions may be assigned a category and / or a flag. The category may be a category within a taxonomy, such as the taxonomy of categories shown in Figure 3. Note that the taxonomy shown in Figure 3 is exemplary, and the actual taxonomy used may include any number of categories and / or levels. The taxonomy may include categories for sociodemographic, physiological, cognitive, and quality of life characteristics, and / or any other type of category.
[0093] Each question may be assigned a root category, and then the question may be assigned a child category of the root category that corresponds to the question.
[0094] For each question, one or more flags can be selected. The flags can represent various characteristics of the question, such as whether the question is answered by the patient or the clinician, whether the question refers to the current time or a past time, and / or other aspects of the question. The available flags can be predefined flags.
[0095] Categories can be applied manually and / or automatically. A semantic system can be used to apply and / or suggest categories that correspond to a question. An administrator can select from the suggested categories or apply other categories. If the selected category has parent categories, those parent categories can be automatically assigned to the question. For example, using the taxonomy shown in Figure 3, if the "Family" category is assigned to a question, the parent categories "Relationships," "Dysfunction," and "Quality of Life" may also be assigned to the question.
[0096] In step 425, compatible questions may be identified and grouped. Questions with the same or similar categories may be grouped. An administrator may review the grouped questions and determine whether they are compatible. An administrator may modify the grouped questions, such as adding or removing questions from the group.
[0097] Questions that have the same type of answer, categorical or binary, may be grouped together. Questions that do not have the same type of answer may not be grouped together. For example, a question with a categorical answer may be determined to be incompatible with a question with a binary answer, even though they may be considered compatible if they have the same type of answer.
[0098] In step 430, the answers to the grouped compatible questions may be rescaled. A single scale may be determined and applied to all questions in the group. For example, if two questions in the group have answer ranges of 1 to 10 and another question in the group has answer ranges of 1 to 5, the answers in the 1 to 5 range may be rescaled to the 1 to 10 range. Any suitable method of rescaling may be used, such as equipercentile equating. Equipercentile equating is an equating method that uses percentiles of the underlying values between two variables to convert one scale to the same scale as the other variable.
[0099] If a group has both binary and categorical questions, the questions may all be rescaled to either binary or categorical format. For example, in a group with binary and categorical questions, the categorical responses may be converted to binary responses.
[0100] In step 435, compatible questions can be combined. Each question in the research data can be associated with an identifier. For each group to be combined, an identifier for the combined question can be generated. The identifier for each question in the group can be replaced with the identifier for the combined question.
[0101] In step 440, any missing data may be imputed. The study results obtained in step 405 may be missing some data. For example, some attribute values may not have been collected for a patient. If too much data is missing from a study, the results of that study may be excluded from further use. Each set of study data may be compared to a threshold amount of missing data. The threshold may be selected by testing several different thresholds, such as 50%, 55%, and 60%. If the amount of data missing from a set of study data exceeds the threshold, the study data may not be included in the final dataset. In some cases, individual attributes or patients within a study may be excluded rather than the entire study if there is a lot of missing data for those attributes. For example, if more than 50% of the values for a patient, study, or attribute are missing, that patient, study, or attribute may be excluded from the study data. Missing data can be imputed using various methods, such as multiple imputation by chained equations (MICE).
[0102] In step 445, a final dataset 215 may be generated. The transformed results 211-14 from each study result 201-04 may be combined to form a single final dataset 215. All of the questions in the final dataset 215 may be standardized, and some of the questions in the final dataset 215 may be combined. Each data point in the final dataset 215 may include various attributes. The attributes may include the answer to the question, any other data about the patient, data about the treatment the patient received, and data about the effectiveness of the treatment for the patient. The data about the effectiveness of the treatment for the patient may be referred to as the "label" of the data point. The MLA 225 may be trained to predict the label of a data point.
[0103] In step 450, the final dataset 215 may be divided into a training set 220, a validation set 221, and / or a test set 222. Each data point in the final dataset 215 may be assigned to either the training set 220, the validation set 221, and / or the test set 222. Any technique may be used to divide the final dataset 215 into the training set 220, the validation set 221, and / or the test set 222, such as randomly selecting data points in the final dataset 215 for each set. A predetermined amount or proportion of data points may be assigned to the training set 220, the validation set 221, and / or the test set 222. For example, the training set 220 may include 80% of the data points in the final dataset 215, the test set 222 may include 10% of the data points in the final dataset 215, and the validation set 221 may include 10% of the data points in the final dataset 215.
[0104] In step 455, the training set 220 may be used to select features for the MLA 225. A subset of the attributes in the final dataset 215 may be selected to be used as features by the MLA 225. Features that are most likely to be useful in predicting the effectiveness of a treatment may be selected. A feature selection process may be applied to the training set 220 to determine which features to use. Any feature selection algorithm may be used. The feature selection algorithm may output one or more attributes of the final dataset 215 to be used as features.
[0105] Features can be selected based on their determined impact on the ultimate outcome of interest. For example, to predict treatment efficacy, MLA 225 can be trained to determine the likelihood that a particular prescribed treatment will lead to remission for any given patient. Features that are likely to influence whether the prescribed treatment will lead to remission can be selected.
[0106] Features may be selected based on unique patterns present in the training set 220. The ability of the MLA 225 to associate treatment effects with patients may be affected by which features are selected. If the selected features are not rich enough, the performance of the MLA 225 may be reduced. Features may be selected that align the patterns discovered in the training set 220 with the ability to determine whether a treatment will lead a patient to remission. In step 455, the most influential features may be selected.
[0107] Examples of features that may be included in and / or selected from the dataset are included below in Table 1. It should be understood that the features listed in Table 1 are exemplary and that other features may be included in and / or generated using the dataset.
[0108] [Table 1A]
[0109] [Table 1B]
[0110] [Table 1C]
[0111] [Table 1D]
[0112] In step 460, the MLA 225 may be trained using the training set 220 and the features selected in step 455. The MLA 225 may be trained to accurately predict the remission rate of each treatment in the final dataset 215. The MLA 225 may receive the training set 220 of data points. For each data point in the training set 220, the MLA 225 may predict the likelihood that each treatment will lead the patient to remission based on the features. The predicted likelihood of the treatment actually administered to the patient during the study may be compared to the label of the data point. The label indicates whether the treatment led to remission during the study. A loss function may be used to compare the label to the prediction output by the MLA 225. The MLA 225 may be adjusted based on the difference between the predicted likelihood and the label. In this manner, the MLA 225 may be trained to receive data points including the features selected in step 455 and output the predicted likelihood that each treatment will lead the patient to remission.
[0113] The loss function may be composed of various subsections that act as regularizers and control terms for the intended behavior of the MLA 225. The global loss that may be used to train the overall MLA 225 may be a weighted sum of some or all of the following components: (1) Remission classification, which refers to whether the probability of remission for a given treatment matches the true incidence (target) of that patient. This can be characterized as a cross-entropy loss function. (2) The loss of an autoencoder may be defined as the Euclidean distance between the original sample x and the decoded sample d(e(x)). Other distance metrics may be used, such as, but not limited to, the change in entropy between distributions and the Wasserstein distance. (3) Controlling the Variance of Prototype-Sample Distances. The variance of distances between prototypes and samples can consist of both (I) the (within-group) variance of the distance between the closest sample of a given prototype and the prototype itself, and / or (II) the (between-group) variance of pairwise distances between all prototypes. These two components can be linearly combined with a coefficient that can adjust their impact on the global objective. This allows for control over prototypes with the goal of ensuring that they are sufficiently spread across the latent sample space to potentially capture topically useful and mutually independent properties of the original patient population. In scenarios where prototypes are learned during the training process, this component may spread prototypes to avoid generating redundant prototypes that do not resemble and / or correctly capture the nuances and characteristics of actual patients. (4) Controlling for differential treatment of remission predictions by prototype treatment. The prototype remission differential variance loss can be composed of both (I) the (between-group) variance in remission predictions across all treatment types among different prototypes, and (II) the (within-group) variance within prototypes and between different treatment predictions. These can be linearly combined by a weighted sum, which allows customization of the composition between these loss components. Because the objective function may encourage greater variance across these two regions, this component of the loss function may be disabled to induce that behavior during the training cycle.
[0114] A weighting coefficient can be assigned to each loss component defined above. For example, the weighting coefficients can be as follows: (1) 1, (2) 0.01, (3.I) 0.001, (3.II) 0.01, (4) 0.01 [The internal module coefficient configuration can be (4.I) 0.05, (4.II) 0.95]. The performance of the classification problem, which is loss component (1) above, can be prioritized over all other loss components, such as by assigning the highest weight to that component. This loss component corresponds to predicting the remission rate for each assigned drug, so the classification problem can be assigned the highest weight. Increasing the weight of this component can improve the accuracy of remission prediction by MLA225.
[0115] The weighting of components (2) through (4) may affect how patient samples are distributed across prototypes using variance. These component weightings may be configured using trainable parameters. The values of these weightings may be continuously updated during the training process of MLA225 to optimize downstream objectives (e.g., supervised / unsupervised / reinforcement learning objectives applied to mental health outcomes).
[0116] An Adam optimizer may be used to direct the training of MLA 225, and / or any other suitable optimizer may be used to train MLA 225. The optimizer may train MLA 225 using training set 220 and the features selected in step 455. The optimizer may configure all trainable parameters of MLA 225, such as the autoencoder, prototype, and / or predictive downstream layers. The optimizer may run data points from training set 220 through MLA 225, calculate individual loss components for each data point, determine changes to make to the parameters to minimize each loss component, and repeat this process to minimize the global loss.
[0117] The aforementioned loss components form a series of sub-optimization problems that the global optimizer uses to determine whether the existing parameters are optimally configured to perform well on each of these sub-problems. The optimizer tracks each operation between each data point and parameter, allowing it to determine the proportion of change to make to each independent parameter in each training cycle to minimize the downstream loss component. The proportion of change made in each training cycle (also known as the learning rate) is a hyperparameter set for the optimizer that affects how quickly it can explore a plausible solution space to output an optimal MLA 225. The learning rate may be predetermined. For example, the learning rate may be set to 0.0001. This may optimize the results to ensure that the MLA 225 learns differences in treatment effects.
[0118] The trained MLA 225 may be tuned using the validation set 221 in step 465. The trained MLA 225 may be tested using the test set 222 in step 470. A loss may be calculated for each data point in the test set 222. The loss may indicate the difference between the label of the data point and the prediction output by the MLA 225 for that data point. An average loss may be determined for all data points in the test set 222. If the average loss is below a predetermined threshold, the MLA 225 is considered ready for use.
[0119] After the MLA 225 has been trained, it can be used to predict a patient's treatment outcome. A questionnaire response corresponding to the patient can be received. The questionnaire response can be completed by the patient and / or a clinician. The questionnaire completed by the patient can be a standardized clinical questionnaire. The questionnaire can be created based on the features used to train the MLA 225.
[0120] The questionnaire responses can be input into the MLA 225. Other data corresponding to the patient, such as physiological data, historical data, sociodemographic data, psychological data, and / or other relevant data about the patient, can also be input into the MLA 225. The MLA 225 can output a predicted efficacy of each treatment that the MLA 225 was trained to predict. An interface can be generated using the predicted efficacy. The interface can display all available treatment options and the predicted efficacy of each treatment option for the patient. The interface can be output for display on a screen and / or monitor. [Explanation of symbols]
[0121] 100 Computing Environments 110 processors 120 Solid State Drive 130 Random Access Memory 150 Input / Output Interfaces 160 System Bus 190 Touchscreen 192 Touch Input / Output Controller 194 Touch Hardware 201 Research results 202 Research results 203 Research results 204 Research results 211 Result after conversion 212 Result after conversion 213 Result after conversion 214 Result after conversion 215 Final Dataset 220 Training Set 221 validation set 222 test set 225 MLA 400 ways
Claims
1. obtaining a plurality of study results, each study result comprising a plurality of data points, each data point comprising a plurality of attributes; extracting a plurality of questions from the plurality of research results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of a plurality of questions based on the one or more categories assigned to each question; determining a response scale to use for each of said plurality of groups; rescaling answers to the group of questions based on the respective answer scales of the group; combining questions for each group of the plurality of groups; imputing missing attribute values in the plurality of studies; generating a final data set by combining the results of the multiple studies; Splitting the final dataset into a training set, a validation set, and a test set; selecting, based on the training set, a plurality of features in the training set; training a machine learning algorithm (MLA) using the features of the training set; A method comprising:
2. 10. The method of claim 1, wherein the plurality of questions comprises questions from a clinical questionnaire administered to patients enrolled in a clinical trial.
3. 10. The method of claim 1, wherein each data point corresponds to a patient enrolled in a clinical trial, and the plurality of attributes includes an indication of a treatment given to the patient.
4. The method of claim 3 , wherein the plurality of attributes comprises an indication of responses to a clinical questionnaire administered to the patient.
5. 5. The method of claim 3 or 4, wherein the plurality of attributes comprises an indication of whether the treatment is effective for the patient.
6. 6. The method of claim 3, wherein the plurality of attributes comprises an indication of side effects experienced by the patient.
7. The method of claim 3 , wherein the plurality of attributes comprises physiological data of the patient.
8. The method of claim 3 , wherein the plurality of attributes includes historical data of the patient.
9. 9. The method of claim 3, wherein the plurality of attributes comprises socio-demographic data of the patient.
10. 10. The method of claim 3, wherein the plurality of attributes comprises psychological data of the patient.
11. 11. The method of claim 1, wherein rescaling the answers to the group of questions comprises using equipercentile scaling to rescale the answers.
12. obtaining a plurality of study results, each study result comprising a plurality of data points, each data point comprising a plurality of attributes; extracting a plurality of questions from the plurality of research results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of questions from the plurality of questions that are assigned the same set of categories; combining said groups of questions; generating a final data set by combining the results of the multiple studies; selecting a plurality of features in the final data set; training a machine learning algorithm (MLA) using the plurality of features; A method comprising:
13. combining said groups of questions determining an answer scale for said group of questions; rescaling answers corresponding to the group of questions based on the answer scale; 13. The method of claim 12, comprising:
14. 1. A method for predicting a patient's therapeutic efficacy, comprising: obtaining a plurality of studies corresponding to a plurality of treatments, each study comprising a plurality of data points, each data point comprising a plurality of attributes; extracting a plurality of questions from the plurality of research results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of questions from the plurality of questions that are assigned the same set of categories; combining said groups of questions; generating a final data set by combining the results of the multiple studies; training the MLA using the final dataset; training a machine learning algorithm (MLA) by receiving questionnaire responses from the patient; entering the questionnaire responses into the MLA; outputting, by the MLA, a predicted efficacy of each of the plurality of treatments; generating an interface based on the predicted efficacy of each of the plurality of treatments; outputting the interface for display; A method comprising:
15. 15. A system comprising at least one processor and a memory storing a plurality of executable instructions that, when executed by the at least one processor, cause the system to perform the method of any one of claims 1 to 14.