System and method for converting and storing data for multiple studies

By converting data from different medical studies into standardized formats and training machine learning algorithms, the combinational difficulties caused by data format differences are solved, and the accuracy and interpretability of treatment choice prediction are improved.

CN120239887APending Publication Date: 2025-07-01AVERED HEALTH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069450.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-05
Filing Date
2023-08-04
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Different questionnaires and data formats are used in different medical studies, making it difficult to combine data to train machine learning algorithms, affecting the prediction accuracy of treatment choices.

Method used

The data from different studies are converted into standardized formats, and the compatibility problems are combined through classification and labeling problems, missing data are filled, and the final data set is generated, which is divided into training, validation and test sets. The features are trained using machine learning algorithms to predict treatment effects.

Benefits of technology

Standardized processing of cross-research data is realized, the training efficiency of machine learning algorithms and the accuracy of treatment choice prediction are improved, and an interpretable therapeutic effect interface is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239887A_ABST
    Figure CN120239887A_ABST
Patent Text Reader

Abstract

A method and system for combining data sets is disclosed. And a research result can be obtained. Each research result may include data points. Each data point may include an attribute. A research problem can be extracted from the research results. The problem may be converted into a standardized format. A category may be specified for the question. The questions may be grouped together to form groups. An answer scale may be determined for each group. The answer may be readjusted using a corresponding answer scale. A final data set may be generated by combining the research results. Features may be selected using the final data set. Features of the final data set may be used to train a machine learning algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 395,381, filed on August 5, 2022, the entire content of which is incorporated herein by reference. Background Art

[0003] Medical research for the same medical condition may collect different data from the patients in the study. These studies may have patients fill out different clinical questionnaires. Data from different medical studies can be used to compare different treatment options. It may be necessary to train machine - learning algorithms to predict the outcomes of the different treatment options used in the studies. Since these studies use different methodologies to collect data on patients, it may not be possible to combine data from multiple studies to train machine - learning algorithms. Summary of the Invention

[0004] Multiple studies may involve the same or similar medical conditions. Different studies may investigate different treatments and / or treatment regimens for the medical condition. Each study collects a variety of information about the patients participating in the study. Each study may collect similar and / or overlapping data, but the data may be collected in different formats.

[0005] To collect information about the patients participating in the study, patients may be given clinical questionnaires. Different studies may use different questionnaires and / or different forms of questionnaires. For example, patients participating in a first study may be given a short - form version of the questionnaire, while patients participating in a second study may be given a long - form version of the questionnaire. Even within a single study, different questionnaires and / or different formats of questionnaires may be used.

[0006] Machine - learning algorithms (MLAs) can be trained using data from different studies. To prepare the data intended for training the MLA, data from different studies can be converted into a common format. The same or nearly the same questions can be combined into a standardized format. Each question can be classified, and similar questions can be grouped.

[0007] Data can be obtained from multiple studies and then used to generate a data set that can be used to train the MLA. A set of data points can be extracted from each study, where each data point includes data corresponding to an individual patient. Each data point can include multiple attributes of the patient extracted from the study data. These attributes can include data collected using a clinical questionnaire. The individual questions in the questionnaire can be extracted from the study.

[0008] Standardized questions can be created based on the questions of the questionnaires used in the study. Each question asked in the study can be converted into one of the standardized questions. A taxonomy with multiple categories can be defined. Each standardized question can be assigned one or more categories of the taxonomy. Standardized questions can also be assigned labels.

[0009] Compatible questions can be identified as they are assigned the same or similar categories. Compatible questions can be grouped together and combined. The responses of the grouped questions can be rescaled to a common scale. The scale can be a categorical scale or a binary scale. The scale of the questions can be determined to be a categorical scale or a binary scale based on the format of the questions being combined. After rescaling the questions in the group to a common scale, the grouped questions can be combined.

[0010] There may be some missing data in the research data. For example, the data points in the study may be missing attribute values. The missing data can be filled. If the amount of missing data in the study exceeds a predefined threshold, individual attributes with a large amount of missing data can be removed from the study or the entire study can be deleted.

[0011] After combining and rescaling the questions, a final dataset can be generated. The final dataset can be divided into a training set, a validation set, and / or a test set. The training set can be used to select features for the MLA. The features can be one or more categories of the taxonomy, attributes of the research data, and / or responses of the questionnaires.

[0012] The MLA can be trained using the features selected from the training set. After training, the MLA can be adjusted using the validation set. Then, the MLA can be tested using the test set.

[0013] According to a first broad aspect of the present technology, there is provided a method, which includes: obtaining a plurality of research results, where each research result includes a plurality of data points, and where each data point includes a plurality of attributes; extracting a plurality of questions from the plurality of research results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining a plurality of question groups based on the one or more categories assigned to each question; determining a response scale for each of the plurality of groups; rescaling the responses of the question groups based on the respective response scales of the groups; combining the questions of each of the plurality of groups; filling in the missing attribute values in the plurality of research results; generating a final dataset by combining the plurality of research results; dividing the final dataset into a training set, a validation set, and a test set; selecting a plurality of features in the training set based on the training set; and training a machine learning algorithm (MLA) using the plurality of features of the training set.

[0014] In some embodiments of the method, the plurality of questions includes questions from clinical questionnaires given to patients participating in a clinical trial.

[0015] In some embodiments of the method, each data point corresponds to a patient participating in a clinical trial, and wherein the plurality of attributes includes indications of the treatment given to the patient.

[0016] In some embodiments of the method, the plurality of attributes includes indications of responses to clinical questionnaires given to the patient.

[0017] In some embodiments of the method, the plurality of attributes includes indications of whether the treatment is effective for the patient.

[0018] In some embodiments of the method, the plurality of attributes includes indications of side effects experienced by the patient.

[0019] In some embodiments of the method, the plurality of attributes includes physiological data of the patient.

[0020] In some embodiments of the method, the plurality of attributes includes historical data of the patient.

[0021] In some embodiments of the method, the plurality of attributes includes sociodemographic data of the patient.

[0022] In some embodiments of the method, the plurality of attributes includes psychological data of the patient.

[0023] In some embodiments of the method, re-scaling the responses of the question groups includes re-scaling the responses using equi-percentile scaling.

[0024] According to another broad aspect of the technology, there is provided a method comprising: obtaining a plurality of research results, wherein each research result includes a plurality of data points, and wherein each data point includes a plurality of attributes; extracting a plurality of questions from the plurality of research results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of questions having the same set of assigned categories; combining the groups of questions; generating a final data set by combining the plurality of research results; selecting a plurality of features in the final data set; and training a machine learning algorithm (MLA) using the plurality of features.

[0025] In some embodiments of the method, combining the groups of questions includes: determining a response scale for the group of questions; and re-scaling the responses corresponding to the group of questions based on the response scale.

[0026] According to a broader aspect of the present technology, a method for predicting a patient's treatment outcome is provided, the method comprising: receiving questionnaire responses from the patient; inputting the questionnaire responses into an MLA trained according to any of the methods described herein; outputting, by the MLA, a predicted outcome for each of the plurality of treatments; generating an interface based on the predicted outcome for each of the plurality of treatments; and outputting and displaying the interface.

[0027] According to another broader aspect of the present technology, a trained MLA trained according to any of the methods described and claimed herein is provided.

[0028] According to another broader aspect of the present technology, a method for a patient (e.g., for predicting treatment outcome) is provided, the method comprising: training a machine learning algorithm (MLA) by: obtaining a plurality of study results corresponding to a plurality of treatments, wherein each study result comprises a plurality of data points, and wherein each data point comprises a plurality of attributes; extracting a plurality of questions from the plurality of study results; converting the plurality of questions into questions in a standardized format; assigning one or more categories to each of the plurality of questions; determining groups of questions having the same set of assigned categories among the plurality of questions; combining the groups of questions; generating a final data set by combining the plurality of study results; and training the MLA using the final data set; receiving questionnaire responses from the patient; inputting the questionnaire responses into the MLA; outputting, by the MLA, a predicted outcome for each of the plurality of treatments; generating an interface based on the predicted outcome for each of the plurality of treatments; and outputting and displaying the interface.

[0029] Multiple embodiments of the present technology provide a computer-based system, e.g., an electronic device including at least one processor and a memory storing program instructions for performing one or more of the methods described herein, the program instructions being executed by the at least one processor of the electronic device.

[0030] It should be clearly understood that not all of the technical effects mentioned herein need to be embodied in each embodiment of the present technology.

[0031] As used herein, the phrase "and / or" is intended to mean inclusive or; e.g., "X and / or Y" is intended to mean X or Y or both. As a further example, "X, Y and / or Z" is intended to mean X or Y or Z or any combination thereof.

[0032] In the context of this specification, unless otherwise expressly specified, a computer system or computing environment may refer to, but is not limited to, an "electronic device", "computing device", "operating system", "system", "computer-based system", "computer system", "network system", "network device", "controller unit", "monitoring device", "control device", "server", and / or any combination thereof suitable for the relevant task at hand.

[0033] In the context of this specification, unless otherwise expressly specified, any method and / or system described herein may be implemented in a cloud-based environment (such as, but not limited to, the Microsoft Azure environment, the Amazon EC2 environment, and / or the Google Cloud environment).

[0034] In the context of this specification, unless otherwise expressly specified, the expressions "computer-readable medium" and "memory" are intended to include any medium of any nature and type, non-limiting examples of which include RAM, ROM, disks (such as CD-ROMs, DVDs, floppy disks, hard disk drives, etc.), USB keys, flash memory cards, solid state drives, and tape drives. Still in the context of this specification, "a" computer-readable medium and "the" computer-readable medium should not be construed as being the same computer-readable medium. On the contrary and where appropriate, "a" computer-readable medium and "the" computer-readable medium may also be construed as a first computer-readable medium and a second computer-readable medium.

[0035] In the context of this specification, unless otherwise expressly specified, words such as "first", "second", "third", etc. used as adjectives are only for distinguishing the nouns they modify and not for describing any specific relationship between these nouns.

[0036] Additional and / or alternative features, aspects, and advantages of embodiments of the present technology will become apparent from the following specification, drawings, and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] For a better understanding of the present technology and its other aspects and further features, please refer to the following description, which will be used in conjunction with the Figure 1 accompanying

[0038] Figure 1 is a block diagram of an exemplary computing environment according to various embodiments of the present technology;

[0039] Figure 2 is a schematic diagram of a process for combining research results to train a machine learning algorithm (MLA) according to various embodiments of the present technology;

[0040] Figure 3Shows a schematic diagram of a taxonomy according to multiple embodiments of the present technology; and

[0041] Figure 4 and Figure 5 Shows a schematic flowchart of a method for generating a combined dataset according to multiple embodiments of the present technology. Detailed Description

[0042] The examples and conditional language described herein are mainly intended to assist the reader in understanding the principles of the present technology and not to limit its scope to such specifically described examples and conditions. It will be understood that those skilled in the art can design various arrangements, which, although not explicitly described or shown herein, still embody the principles of the present technology and are included within its spirit and scope.

[0043] In addition, for the sake of understanding, the following description may describe relatively simplified embodiments of the present technology. As those skilled in the art will understand, multiple embodiments of the present technology may be more complex.

[0044] In some cases, useful examples that are considered to be modifications of the present technology may also be set forth. This is only for the sake of understanding and is not intended to limit the scope or set forth the boundaries of the present technology. These modifications are not an exhaustive list, and those skilled in the art can make other modifications while still remaining within the scope of the present technology. In addition, in the case where examples of modifications are not set forth, it should not be construed that modifications are not possible and / or that the described is the only way to implement the element of the present technology.

[0045] In addition, all descriptions herein that set forth the principles, aspects, and embodiments of the present technology and their specific examples are intended to include their structural equivalents and functional equivalents, whether currently known or developed in the future. Thus, for example, those skilled in the art will understand that any block diagram herein represents a conceptual view of an illustrative circuit embodying the principles of the present technology. Similarly, it will be understood that any flowchart, flow schematic diagram, state transition diagram, pseudocode, etc. represents multiple processes that can be substantially represented in a computer-readable medium and thus executed by a computer or processor, whether or not such a computer or processor is explicitly shown.

[0046] The functionality of the various elements shown in the figures, including any functional blocks labeled as "processor", can be provided by using dedicated hardware as well as hardware capable of executing software in association with appropriate software. When the functionality is provided by a processor, the functionality can be provided by a single dedicated processor, a single shared processor, or multiple individual processors, some of which may be shared processors. In some embodiments of the present technology, the processor can be a general-purpose processor, such as a central processing unit (CPU), or a processor dedicated to a specific purpose, such as a digital signal processor (DSP). Additionally, the explicit use of the term "processor" should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, but is not limited to, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), read only memories (ROMs) for storing software, random access memories (RAMs), and non-volatile memories. Other conventional hardware and / or custom hardware may also be included.

[0047] Software modules or simple modules implied to be software can herein be represented as any combination of flowchart elements or other elements representing the performance of process steps and / or text descriptions. Such modules can be executed by hardware explicitly or implicitly shown. Additionally, it should be understood that one or more modules can include, for example but not limited to, computer program logic, computer program instructions, software, stacks, firmware, hardware circuits, or combinations thereof.

[0048] Computing environment

[0049] Figure 1 A computing environment 100 is shown, which can be used to implement and / or execute any method described herein. In some embodiments, the computing environment 100 can be implemented by any conventional personal computer, a computer dedicated to managing network resources, network devices, and / or electronic devices (such as but not limited to mobile devices, tablet devices, servers, controller units, control devices, etc.) and / or any combination thereof suitable for the relevant task at hand. In some embodiments, the computing environment 100 includes a plurality of hardware components, including one or more single-core or multi-core processors collectively represented by a processor 110, a solid state drive 120, a random access memory 130, and an input / output interface 150. The computing environment 100 can be a computer specifically designed to run machine learning algorithms (MLAs). The computing environment 100 can be a general-purpose computer system.

[0050] In some embodiments, the computing environment 100 can also be a subsystem of one of the above systems. In some other embodiments, the computing environment 100 can be an "off-the-shelf" general-purpose computer system. In some embodiments, the computing environment 100 can also be distributed across multiple systems. The computing environment 100 can also be dedicated to implementing the present technology. As those skilled in the art can understand, various variations on how to implement the computing environment 100 can be envisioned without departing from the scope of the present technology.

[0051] Those skilled in the art will understand that the processor 110 generally represents processing capabilities. In some embodiments, one or more dedicated processing cores can be provided as a replacement or supplement to one or more conventional central processing units (CPUs). For example, one or more graphics processing units (GPUs), tensor processing units (TPUs), and / or other so-called acceleration processors (or processing accelerators) can be provided as a supplement or replacement to one or more CPUs.

[0052] The system memory will generally include random access memory 130, but more generally is intended to include any type of non-transitory system memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or combinations thereof. The solid-state drive 120 is shown as an example of a mass storage device, but such mass storage can more generally include any type of non-transitory storage device configured to store data, programs, and other information and to make the data, programs, and other information accessible via the system bus 160. For example, mass storage can include one or more of solid-state drives, hard disk drives, disk drives, and / or optical disc drives.

[0053] Communication between the multiple components of the computing environment 100 can be achieved via a system bus 160 that includes one or more internal buses and / or external buses (e.g., PCI bus, Universal Serial Bus, IEEE 1394 "FireWire" bus, SCSI bus, Serial ATA bus, ARINC bus, etc.), with multiple hardware components electronically coupled to the system bus 160.

[0054] The input / output interface 150 may allow for the implementation of networking functions such as wired access or wireless access. As an example, the input / output interface 150 may include a networking interface, such as but not limited to a network port, a network socket, a network interface controller, etc. To those skilled in the art of the present technology, multiple instances of how the networking interface may be implemented will become apparent. For example, the networking interface may implement specific physical layer and data link layer standards, such as Ethernet, Fibre Channel, Wi-Fi, Token Ring, or serial communication protocols. The specific physical layer and data link layer may provide a basis for a complete network protocol stack, enabling communication between small groups of computers on the same local area network (LAN) and large-scale network communication via routing protocols such as the Internet Protocol (IP).

[0055] The input / output interface 150 may be coupled to the touch screen 190 and / or one or more internal and / or external buses 160. The touch screen 190 may be part of a display. In some embodiments, the touch screen 190 is the display. The touch screen 190 may also be referred to as the screen 190. In Figure 1 the illustrated embodiment, the touch screen 190 includes touch hardware 194 (e.g., a pressure-sensitive unit embedded in a display layer capable of detecting physical interaction between a user and the display) and a touch input / output controller 192 capable of communicating with the display interface 140 and / or one or more internal and / or external buses 160. In some embodiments, the input / output interface 150 may be connected to a keyboard (not shown), a mouse (not shown), or a touchpad (not shown) such that a user may interact with the computing device 100 in addition to or instead of the touch screen 190.

[0056] According to some embodiments of the present technology, the program instructions stored in the solid-state drive 120 are suitable for being loaded into the random access memory 130 and executed by the processor 110 for performing the acts of one or more methods described herein. For example, at least some of the program instructions may be part of a library or an application.

[0057] Combined research results

[0058] Figure 2 is a schematic diagram of a process of combining research results to train a machine learning algorithm (MLA) according to multiple embodiments of the present technology. Research results 201-04 may be obtained. Each research may test the effects of different treatments and / or treatment regimens on the same medical condition. For example, the research result 201 may be the research result of a compound for treating depression, and the research result 202 may be the research result of a different compound for treating depression.

[0059] Although described as Study Result 201-04, it should be understood that the methods described herein can be used for any type of data set that combines general data that may be in different formats. The data sets to be combined can be the results of clinical trials and / or any other type of data set. The data to be combined can include any type of patient data, such as electronic medical records (EMRs). For example, patient attributes, treatment information related to the patient, and outcome data corresponding to the treatment can be extracted from the patient's EMR and then combined with data extracted from the EMRs of other patients.

[0060] Each set of Study Results 201-04 can include a set of data points for individual patients involved in each study. Each data point corresponds to an individual patient participating in the study. The data points can include multiple attributes. Attributes of a data point corresponding to a patient can include the patient's responses to a clinical questionnaire, information related to the patient, information related to the treatment applied to the patient, outcome data indicating the patient's response to the treatment, the patient's medical history, the patient's family history, physiological information related to the patient, and / or any other data collected about the patient.

[0061] The outcome data of a data point can include whether the treatment resulted in remission, the time to remission, whether the treatment caused harm and / or had adverse side effects, whether the treatment resolved certain symptoms, whether the treatment led to a return to baseline physiological measurements, and / or any other measurement of treatment effectiveness.

[0062] The different studies that produced Study Results 201-04 may have used the same and / or different questionnaires. Even if different questionnaires were used or questionnaires in different formats were used, the different questionnaires may have similar questions. For example, the questionnaire used in a first study and the questionnaire in a second study may ask different questions about the same topic.

[0063] Clinical questionnaires can include questions related to the patient's mental health, medical history, family medical history, current medications, sociodemographic information, and / or any other type of question. Patients can be periodically asked to update the clinical questionnaire and / or complete a new clinical questionnaire so that the information collected about the patient is up-to-date. Thus, a single data point in Study Results 201-04 may include multiple responses to the same questionnaire given at different times. Patients can be asked to update the clinical questionnaire after a predetermined time has passed. The clinical questionnaire can be completed by the patient, the patient's caregiver, and / or the doctor.

[0064] A standardized version of the questions in Research Outcome 201-04 can be generated. Similar but not identical questions can be asked in different studies and / or different questionnaires. Similar questions can be grouped together and replaced with the standardized version of that question in Transformation Outcome 211-14.

[0065] A taxonomy of categories can be generated and / or predefined. The categories can be categories of questions in a questionnaire. The taxonomy can be organized in a tree structure where each category can have parent nodes and / or child nodes. Figure 3 An example of the taxonomy is shown. In Figure 3 the shown taxonomy, the "Functional Impairment" category is a branch of the "Quality of Life" root category. The "Functional Impairment" category also has leaves and another sub-branch, Interpersonal Relationships, and Interpersonal Relationships itself has leaves: "Family", "Social", and "Romance".

[0066] The questions in Research Outcome 201-04 can each be assigned one or more categories in the taxonomy. Categories can be assigned to questions based on keywords in the questions. The MLA can be used to assign categories to questions based on manual review and / or based on any other suitable method. The MLA can be trained to assign categories to questions based on the text of the questions. The questions can be input into the MLA and the MLA can output one or more predicted categories corresponding to the questions. Then, a human operator can confirm that the predicted categories correspond to the questions and / or edit the categories assigned to the questions. Other methods can be used to suggest categories applicable to the questions, such as by extracting individual words from the questions.

[0067] After standardizing and categorizing the questions, questions that are compatible with each other can be grouped together and then combined. Compatible questions may be semantically similar to each other. Groups of questions with the same and / or similar assigned categories can be formed. For each group of questions, a common response format can be selected. The responses to the questions in the group can be re-formatted to the selected format. After re-formatting the responses, the grouped questions can be combined.

[0068] Research Outcome 201-04 can be transformed into Transformation Outcome 211-14. In Transformation Outcome 211-14, each question can be in a standardized format. Similar questions can be transformed into a combined version of that question.

[0069] Then the conversion results 211-14 of each study can be combined into the final dataset 215. Since the study results 201-04 have been converted to a common format in the conversion results 211-14, the data can be combined into the final dataset 215. The final dataset 215 can include data points from each of the study results 201-04. Each data point in the final dataset 215 can be associated with an individual patient. Each data point can include multiple attributes that represent the standardized and transformed questions, the treatments received by the patient, and the indications of the effects of the treatments on the patient.

[0070] The final dataset 215 can be divided into a training set 220, a validation set 221, and a test set 222. The training set 220, the validation set 221, and the test set 222 can each include a subset of the data points in the final dataset 215. Each data point in the final dataset 215 can be randomly assigned to the training set 220, the validation set 221, or the test set 222. The training set 220, the validation set 221, and the test set 222 can have predetermined sizes. For example, the training set 220 can be configured to be larger than the validation set 221.

[0071] After generating the training set 220, the training set 220 can be used to train the MLA 225. Each data point in the training set 220 can be input into the MLA 225. If a patient is given multiple treatments, the MLA 225 can predict the likelihood of remission of the patient corresponding to the data point. For each possible treatment, the MLA 225 can output the predicted likelihood of remission.

[0072] The MLA 225 can include one or more neural networks and / or any other type of machine learning model, deep learning model, and / or artificial intelligence (AI) model. The MLA 225 can include an encoder that takes as input x the treatment not specified for the patient and is responsible for encoding the features corresponding to the patient into a certain latent space e(x). The decoder can decode the encoded features back to the original input d(e(x)). The decoded features may not be the same as the original encoded features. Any suitable type of encoder and / or decoder can be used, such as a symmetric autoencoder.

[0073] The loss function can compare the predicted treatment effect with the actual result recorded during the study, which is stored in the training set 220. The actual result can be referred to as the "label" of the data point, and the MLA 225 can be trained to predict the label. Then the MLA 225 can adjust itself based on the amount of loss output by the loss function. The MLA 225 can be adjusted to reduce the amount of loss between the prediction and the label. This training process can be repeated for each data point in the training set 220.

[0074] After training the MLA 225 using the training set 220, the MLA 225 can be adjusted using the validation set 221. Each data point in the validation set 221 can be input into the MLA 225. Then the MLA 225 can output a prediction result, and the loss function can be used to compare the prediction result with the corresponding label. The parameters of the MLA 225 can be adjusted based on the result of the loss function.

[0075] After adjusting the MLA 225 using the validation set 221, the test set 222 can be input into the MLA 225 to determine the prediction accuracy of the MLA 225. Multiple methods can be used to determine whether the accuracy of the MLA 225 is sufficient for further use. The predictions output by the MLA 225 can be compared with the results in the test set 222. The average loss amount of the MLA 225 for the test set 222 can be determined. If the average loss amount is lower than a pre-determined threshold loss amount, it can be determined that the MLA 225 is ready for use.

[0076] After training the MLA 225, in order to use the MLA 225, a clinical questionnaire can be conducted on patients. The encoder function e(x) can be used to encode the answers to the questions into digital vectors. Then the vectors can be input into the MLA 225.

[0077] Prototype

[0078] Multiple prototypes can be defined for patient groups. Each prototype can correspond to a group of patients with similar characteristics, presenting similar symptoms, and / or responding similarly to one or more treatments. The prototypes can be defined such that each prototype responds differently to the available treatments. The prototypes can help clinicians and / or patients understand the results output by the MLA 225. In other words, the prototypes can be used to improve the interpretability of the results by clinicians and / or patients. Each prototype can be used to generate an exemplary patient corresponding to the prototype to compare the real patient with the prototype.

[0079] The training of the MLA 225 can involve forming a neural network layer of the MLA 225 that extracts these prototypes. Each prototype can indicate the importance of features in predicting patient remission and / or the differential effects of different treatments on a given prototype. Each prototype can be associated with a patient group, that is, a group of patients relatively similar to the learned prototype. Prototype extraction can improve the accuracy of the MLA 225 and / or improve the interpretability of the MLA 225. Prototype extraction can help clinicians understand the output of the MLA 225 by showing how different groups of features representing different patient prototypes may respond to different treatments.

[0080] The number of prototypes to be defined can be determined empirically (through human / non - human initialization and experimental processes) and / or dynamically (through algorithmic determinism to optimize downstream objectives). The number of prototypes can be selected based on various considerations, such as increasing the interpretability and / or accuracy of the prototypes. For example, the number of prototypes can be set to three, which can provide a balance between providing sufficient nuances between prototypes and also providing sufficiently accurate MLA225.

[0081] In some cases, prototypes can be defined in the original feature space without using an auto - encoder, but then encoded into the latent space by the auto - encoder for compatibility when comparing with the encoded features. Prototypes can be defined manually by an operator and / or automatically using various functions such as clustering algorithms. For example, the operator can input multiple parameters for the prototypes.

[0082] Given the symmetric nature of the neural network, both the encoder and the decoder can include the same number of fully - connected layers. The output of the encoding layer e(x) can be input into the prototype layer p, which can be configured with k nodes to represent each prototype respectively. The variable k can represent the number of patient archetypes that the prototypes can learn to represent respectively. Each node can be the size of the input data sample. Prototypes can be defined in the latent (encoded) space.

[0083] To compare the patient's data with the set of prototypes, both can be mapped into the encoded space. Prototypes can have learning parameters that can be configured to move around the encoded feature space to achieve optimal downstream prediction performance for the MLA. Prototypes can be assigned "frozen" weights that can ensure the prototypes remain static throughout the MLA training.

[0084] To make the prototypes interpretable by an operator such as a clinician, the prototypes can be decoded by the decoder d(p). The decoder can extract the original feature values corresponding to the prototypes. Content experts such as clinicians can consult the original feature values of the prototypes to better understand the prototypes and their relationship with the predicted probability of treatment effectiveness.

[0085] When inputting the patient's data into the MLA225, an auto - encoder can be used to calculate the distance between the patient and each prototype in the latent space. Then, these distances can be passed down for downstream prediction objectives.

[0086] Prototype Configuration

[0087] When defining prototypes, multiple hyperparameters can be configured, including (1) the number of prototypes that the MLA225 will support and (2) the adjustable parameters for each prototype.

[0088] Any number of prototypes can be defined. For the purpose of improving interpretability, it may be preferred to have a relatively small number of prototypes, such as two, three, or four prototypes, because having too many prototypes may make it difficult for clinicians to understand and / or interpret why a patient may benefit more from one treatment than another. From a performance perspective, the number of prototypes can also be configured to optimize downstream objectives, such as predicting the remission rate of a treatment. The operator can select the number of prototypes to define to balance the interpretability and overall performance of the MLA 225.

[0089] The parameters of the prototypes can be defined in various ways, such as based on operator input and / or automatically defined using functions. The operator can define the parameters for the prototypes. The prototypes can then be generated based on the parameters defined by the operator. The parameters of the prototypes can be generated using functions such as clustering algorithms.

[0090] Previously determined patterned patient groups can be used as the basis for the prototypes. The operator can define the parameters corresponding to the previously determined groups, and by passing these parameter values to the encoder e(x), the previously determined groups can be converted into prototypes to initialize the prototypes.

[0091] Prototypes can be initialized using algorithms. The following can be used to initialize the prototypes: Xavier-Glorot uniform distribution / normal distribution, He (i.e., Kaiming) uniform distribution / normal distribution, or normal distribution or uniform distribution, or other pre-existing distributions or custom distributions that allow sampling of a set of parameters from a continuous set or a discrete set of values.

[0092] Prototype output

[0093] The output of the prototype layer can represent the distance between the patient-encoded data e(x) and each prototype p. In other words, the output can represent the distance between the patient and each prototype. This distance can be defined by the separate Frobenius norms between the encoded samples and each prototype. Any other suitable distance metric can be used, such as variance-based distance (assuming each prototype represents a statistically distinct distribution of samples), Mahalanobis distance, or modeling each prototype group as a normal distribution to identify which patient samples are most likely to be within a certain standard deviation of the group center. For example, if a sample is more likely to be within one standard deviation of the first group than within three standard deviations of the second group, the distance can reflect the degree of overlap.

[0094] The potential distance from a patient to a prototype can be input into a fully connected neural layer that is connected to a specified treatment that was omitted before passing patient information to an encoder. The treatment can be encoded in a one-hot manner before being concatenated with the rest of the distance vector. This concatenation can be input into a final classification layer that aims to extract the likelihood of remission for each specified treatment to test the hypothetical scenarios for each patient. Then, these predicted remission rates for each patient for each treatment can be aggregated and used to calculate differential benefits.

[0095] Method for generating a combined dataset

[0096] Figure 4 and Figure 5 FIG. 400 is a flow diagram showing a method 400 for generating a combined dataset according to various embodiments of the present technology. In one or more aspects, method 400 or one or more of its steps may be performed by a computing system such as computing environment 100. Method 400 or one or more of its steps may be embodied in computer-executable instructions stored in a computer-readable medium (such as a non-transitory mass storage device), loaded into memory, and executed by a CPU. Some steps or portions of steps in the flow diagram may be omitted or the order may vary.

[0097] At step 405, data may be obtained from multiple studies, such as study results 201-04. As described above, the studies may be clinical studies testing multiple treatments for a medical condition. The data obtained may be de-identified patient-level data from clinical trials for treating a condition such as depression. The data may include the relevant study protocols corresponding to the studies.

[0098] Patients participating in the studies may have completed clinical questionnaires. Different questionnaires and / or different formats of questionnaires may be used in the studies. Different patients may be given different questionnaires and / or different formats of questionnaires. The data obtained at step 405 may include the answers to the questionnaires given by the patients. The data obtained at step 405 may include the study results for each patient. The study results for each patient may include an indication of whether the treatment was effective for the patient, and / or a measure of the degree to which the treatment was effective for the patient. The data obtained at step 405 may include indirect measures of treatment effectiveness, such as quality of life assessments, functional impairment ratings, and / or any other indirect measures.

[0099] In step 410, questionnaires can be extracted from each of the research results 201-04. In some cases, a single questionnaire can be extracted from one of the research results 201-04. In other cases, multiple questionnaires can be extracted from one of the research results 201-04. The extracted questionnaires can include each individual question asked of the research participants. Each question can be associated with information on the available responses. The information related to the available responses can indicate whether the participant can provide a binary response or a categorical response to the question. A binary response is a response in which the patient can choose "yes" or "no", "true" or "false", or any other binary response. A categorical response is a response in which the patient can choose from multiple responses and / or choose from responses on a scale such as a Likert scale. An example of a question with a categorical response is a question that asks the patient to provide a rating from 1 to 5. The information related to the available responses to a categorical question can include the scale corresponding to the response, information on the options available to the patient for the response, and / or information indicating the meaning of the highest and lowest values on the scale.

[0100] In step 415, standardized questions can be created. Each question extracted in step 410 can be converted to a standardized format. Standard question text can be created for each question. Different questionnaires can include the same question with slightly different formats and / or the order of the questions may be different. A single standard version of the question can be selected, and all other iterations of the question can be matched to the standard version.

[0101] Different studies can use the same questionnaire with slightly different versions. Different versions of the questionnaire can include a similar set of questions, but the text of the questions may be slightly different and / or the order of the questions may be different. A single version of the questionnaire can be selected as the standard version of the questionnaire. Responses to all other iterations of the questionnaire used can be modified to match the standard version of the questionnaire.

[0102] In step 420, each standardized question can be assigned a category and / or flagged. The category can be a category in a taxonomy, such as the category taxonomy shown in Figure 3 It should be noted that the taxonomy shown in Figure 3 is exemplary, and the actual taxonomy used can include any number of categories and / or levels. The taxonomy can include categories related to sociodemographic characteristics, physiological characteristics, cognitive characteristics, and quality of life characteristics, and / or any other type of category.

[0103] Each question can be assigned a root category. Then, subcategories corresponding to the root category of the question can be assigned to the question.

[0104] One or more tags can be selected for each question. These tags can represent multiple characteristics of the question, such as whether the question was answered by a patient or a clinician, whether the question refers to the current time or a past time, and / or any other aspect of the question. The available tags can be predefined tags.

[0105] Categories can be applied manually and / or automatically. A semantic system can be used to apply and / or suggest categories corresponding to the questions. An administrator can select from the suggested categories and / or apply other categories. If the selected category has a parent category, then these parent categories can be automatically assigned to the question. For example, using the Figure 3 taxonomy shown, if the "Family" category is assigned to a question, then the parent categories "Interpersonal Relationships", "Functional Impairment", and "Quality of Life" can also be assigned to that question.

[0106] In step 425, compatible questions can be identified and grouped together. Questions with the same or similar categories can be grouped together. An administrator can review the questions that have been grouped together to determine if they are compatible. The administrator can modify the grouped questions, for example, by adding or deleting questions from the group.

[0107] Questions with the same type (binary or categorical) of answers can be grouped together. Questions with different types of answers may not be grouped together. For example, questions with categorical answers may not be determined to be compatible with questions with binary answers, even if they would be considered compatible if they had the same type of answers.

[0108] In step 430, the answers of the compatible questions grouped together can be re-scaled. A single scale can be determined and applied to all the questions in the group. For example, if the answers of two questions in the group range from 1 to 10 and the answer of another question in the group ranges from 1 to 5, then the answers ranging from 1 to 5 can be re-scaled to range from 1 to 10. Any suitable re-scaling method can be used, such as percentile equating. Percentile equating is an equating method that uses the percentile of the latent values between two variables to transform one scale to the same scale as another variable.

[0109] If a group has both binary questions and categorical questions, the questions can all be re-scaled to a binary format or a categorical format. For example, for a group with binary questions and categorical questions, the categorical answers can be converted to binary answers.

[0110] In step 435, compatible questions can be combined. Each question in the research data can be associated with an identifier. For each group that is being combined, an identifier for the combined question can be generated. The identifier of each question in the group can be replaced with the identifier of the combined question.

[0111] In step 440, any missing data can be filled. The research results obtained in step 405 may have some missing data. For example, some attribute values of patients may not have been collected. If too much data is missing in a study, the results of that study may be deleted and not further used. Each research data set can be compared with a threshold amount of missing data. The threshold can be selected by testing multiple different thresholds, such as 50%, 55%, 60%. If the amount of missing data in a research data set exceeds the threshold, that research data may not be included in the final data set. In some cases, if there is a large amount of missing data for individual attributes or patients in a study, those attributes or patients can be deleted rather than deleting the entire study. For example, if the missing values for a patient, study, or attribute exceed 50%, that patient, study, or attribute can be deleted from the research data. Multiple methods can be used to fill in the missing data, such as multiple imputation by chained equations (MICE).

[0112] In step 445, the final data set 215 can be generated. The transformation results 211-14 of each of the research results 201-04 can be combined to form a single final data set 215. All the questions in the final data set 215 can be standardized, and some questions in the final data set 215 can be combined. Each data point in the final data set 215 can contain multiple attributes. These attributes can include answers to questions, any other data about the patient, data about the treatment given to the patient, and data about the effect of the treatment on the patient. The data about the effect of the treatment on the patient can be referred to as the "label" of the data point. The MLA 225 can be trained to predict the label of the data point.

[0113] In step 450, the final data set 215 can be divided into a training set 220, a validation set 221, and / or a test set 222. Each data point in the final data set 215 can be assigned to the training set 220, the validation set 221, and / or the test set 222. Any technique can be used to divide the final data set 215 into the training set 220, the validation set 221, and / or the test set 222, such as randomly selecting data points for each set in the final data set 215. The training set 220, the validation set 221, and / or the test set 222 can be assigned a predetermined number or proportion of data points. For example, the training set 220 can include 80% of the data points in the final data set 215, the test set 222 can include 10% of the data points in the final data set 215, and the validation set 221 can include 10% of the data points in the final data set 215.

[0114] In step 455, the training set 220 can be used to select features for the MLA 225. A subset of the attributes in the final data set 215 can be selected to be used as features for the MLA 225. Features that are likely to be most useful for predicting the treatment effect can be selected. A feature selection method can be applied to the training set 220 to determine which features will be used. Any feature selection algorithm can be used. The feature selection algorithm can output one or more attributes of the final data set 215 to be used as features.

[0115] Features can be selected based on their determining influence on the outcome of the final prediction target. For example, to predict the treatment effect, the MLA 225 can be trained to determine the likelihood that a certain prescription treatment will result in remission for any given patient. Features that seem to affect whether the prescription treatment will result in remission can be selected.

[0116] Features can be selected based on the inherent patterns present in the training set 220. The ability of the MLA 225 to associate the treatment effect with the patient can be affected by the features selected. If the selected features are not rich enough in information, the performance of the MLA 225 may decline. Features that align the patterns found in the training set 220 with its ability to determine whether the treatment will result in patient remission can be selected. Features with the maximum influence amount can be selected in step 455.

[0117] Examples of features that can be included in the data set and / or features that can be selected are shown in Table 1 below. It should be understood that the features listed in Table 1 are exemplary, and other features can be included in the data set and / or generated using the data set.

[0118] Table 1 - Examples of FeaturesThis table shows a list of features in a list form.

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128] In step 460, the training set 220 and the features selected in step 455 can be used to train the MLA 225. The MLA 225 can be trained with the goal of accurately predicting the remission rate of each treatment in the final dataset 215. The MLA 225 can receive the training set 220 of data points. For each data point in the training set 220, the MLA 225 can predict the likelihood that each treatment will result in patient remission based on the features. The predicted likelihood of the treatment actually given to the patient during the study can be compared with the label of the data point. The label indicates whether the treatment resulted in remission during the study. A loss function can be used to compare the label with the prediction output by the MLA 225. The MLA 225 can be adjusted based on the difference between the predicted likelihood and the label. In this way, the MLA 225 can be trained to receive data points including the features selected in step 455 and output the predicted likelihood that each treatment will result in patient remission.

[0129] The loss function can include multiple sub-parts that serve as regularization terms and control terms for the expected behavior of the MLA 225. The global loss that can be used to train the overall MLA 225 can be a weighted sum of some or all of the following components:

[0130] (1) Remission classification regarding whether the likelihood of remission for a given treatment matches the true occurrence (target) for that patient. This can be characterized as a cross-entropy loss function.

[0131] (2) The autoencoder loss can be defined using the Euclidean distance between the original sample x and the decoded sample d(e(x)). Other distance metrics can be used, such as but not limited to the change in entropy between distributions and the Wasserstein distance.

[0132] (3) Controlling the variance of the prototype sample distances. The variance of the distances between prototypes and samples can consist of two parts: (I) the (intra-group) variance of the distances between the nearest sample to a given prototype and the prototype itself and / or (II) the (inter-group) variance of the pairwise distances between all prototypes. These two components can be linearly combined with coefficients that can adjust their impact on the global objective. This can control the prototypes, with the goal of keeping the prototypes sufficiently dispersed in the latent sample space so as to potentially capture locally useful and mutually independent properties of the original patient population. In the case where prototypes are learned during the training process, this component may cause the prototypes to disperse, thus avoiding the generation of redundant prototypes that may not be similar and / or may not correctly capture the nuances and features of the true patients.

[0133] (4) Controlling differential treatment remission prediction for prototypes. The differential prototype remission variance loss can consist of two parts: (I) the (between-group) variance of remission prediction among different prototypes across all treatment types and (II) the (within-group) variance within a prototype and between different treatment predictions. These can be linearly combined through a weighted sum with customizable configuration between these loss components. Since the objective function may promote greater variance in these two domains, this component of the loss function can be negated to induce this behavior during the training cycle.

[0134] Weighting coefficients can be specified for each of the loss components defined above. For example, the weighting coefficients can be as follows: (1) 1, (2) 0.01, (3.I) 0.001, (3.II) 0.01, (4) 0.01 [whose internal module coefficient composition can be (4.I) 0.05, (4.II) 0.95]. The performance of the classification problem is the above loss component (1), which can be prioritized over all other above loss components, for example, by specifying the maximum weight for this component. Since this loss component corresponds to predicting the remission rate for each specified drug, the classification problem can be assigned the maximum weight. By increasing the weight of this component, the accuracy of MLA 225 remission prediction can be improved.

[0135] Weighting of components (2) to (4) can affect how patient samples are distributed among prototypes using variance. The weighting of these components can be configured using trainable parameters. These weighting values can be continuously updated during the MLA225 training process to optimize downstream objectives (e.g., supervised / unsupervised / reinforcement objectives applied to mental health outcomes).

[0136] The training of MLA225 can be specified using the Adam optimizer and / or MLA225 can be trained using any other suitable optimizer. The optimizer can use the training set 220 and the features selected in step 455 to train MLA225. The optimizer can configure all trainable parameters of MLA225, such as the autoencoder, prototypes, and / or prediction downstream layers. The optimizer can pass data points from the training set 220 through MLA225, calculate the respective loss components for each data point, determine the changes to be made to the parameters to minimize each loss component, and repeat this process to minimize the global loss.

[0137] The loss components described above form a series of sub-optimization problems, which the global optimizer uses to determine whether the existing parameters are optimally set to perform well on each sub-problem. The optimizer tracks every operation that occurs between each data and parameter, such that the proportional amount of change to each individual parameter can be determined in each training cycle to minimize the downstream loss component. The proportional change per learning cycle (also known as the learning rate) is a hyperparameter set for the optimizer, which affects the speed at which the optimizer can explore the credible solution space to output the optimal MLA 225. The learning rate can be predetermined. For example, the learning rate can be set to 0.0001. This can optimize the results to ensure that the MLA225 can learn differential treatment benefits.

[0138] The trained MLA225 can be adjusted using the validation set 221 in step 465. The trained MLA225 can be tested using the test set 222 in step 470. The amount of loss for each data point in the test set 222 can be calculated. The amount of loss can indicate the difference between the label of the data point and the prediction output by the MLA 225 for that data point. The average amount of loss for all data points in the test set 222 can be determined. If the average amount of loss is below a predetermined threshold, the MLA 225 can be considered ready for use.

[0139] After the MLA225 is trained, the MLA225 can be used to predict the treatment effect of a patient. Survey responses corresponding to the patient can be received. The survey responses can be completed by the patient and / or the clinician. The survey completed by the patient can be a standardized clinical survey. The survey can be generated based on the features used to train the MLA225.

[0140] The survey responses can be input into the MLA225. Other data corresponding to the patient, such as physiological data, historical data, sociodemographic data, psychological data, and / or any other relevant data about the patient, can also be input into the MLA225. The MLA225 can output the predicted effects of each treatment that the MLA225 is trained to predict. An interface can be generated using the predicted effects. The interface can display all available treatment options and the predicted effects of each treatment option on the patient. The interface can be output and displayed on a screen and / or monitor.

Claims

1. A method, comprising: Obtaining a plurality of research results, wherein each research result includes a plurality of data points, and wherein each data point includes a plurality of attributes; Extracting a plurality of questions from the plurality of research results; Converting the plurality of questions into questions in a standardized format; Assigning one or more categories to each of the plurality of questions; Determining a plurality of question groups based on the one or more categories assigned to each question; Determining a response scale for each of the plurality of groups; Readjusting the responses of the question groups based on the respective response scales of the groups; Combining the questions of each of the plurality of groups; Filling in missing attribute values in the plurality of research results; Generating a final data set by combining the plurality of research results; Dividing the final data set into a training set, a validation set, and a test set; Selecting a plurality of features in the training set based on the training set; and Training a machine learning algorithm (MLA) using the plurality of features of the training set.

2. The method of claim 1, wherein the plurality of questions include questions from a clinical questionnaire given to patients participating in a clinical trial.

3. The method of claim 1, wherein each data point corresponds to a patient participating in a clinical trial, and wherein the plurality of attributes include indications of the treatment given to the patient.

4. The method of claim 3, wherein the plurality of attributes include indications of responses to a clinical questionnaire given to the patient.

5. The method of claim 3 or 4, wherein the plurality of attributes include indications of whether the treatment is effective for the patient.

6. The method of any one of claims 3 to 5, wherein the plurality of attributes include indications of side effects experienced by the patient.

7. The method of any one of claims 3 to 6, wherein the plurality of attributes include physiological data of the patient.

8. The method of any one of claims 3 to 7, wherein the plurality of attributes include historical data of the patient.

9. The method of any one of claims 3 to 8, wherein the plurality of attributes include sociodemographic data of the patient.

10. The method of any one of claims 3 to 9, wherein the plurality of attributes include psychological data of the patient.

11. The method of any one of claims 1 to 10, wherein readjusting the responses of the question groups includes readjusting the responses using equal percentile scaling.

12. A method, comprising: Obtaining a plurality of research results, wherein each research result includes a plurality of data points, and wherein each data point includes a plurality of attributes; Extracting a plurality of questions from the plurality of research results; Converting the plurality of questions into questions in a standardized format; Assigning one or more categories to each of the plurality of questions; Determining question groups having the same set of assigned categories among the plurality of questions; Combining the question groups; Generating a final data set by combining the plurality of research results; Selecting a plurality of features in the final data set; and Training a machine learning algorithm (MLA) using the plurality of features.

13. The method of claim 12, wherein combining the question groups includes: Determine a response scale for the set of questions; and Based on the response scale, re-adjust the responses corresponding to the set of questions.

14. A method for predicting a patient's treatment outcome, the method comprising: Training a machine learning algorithm (MLA) by the following steps: Obtaining a plurality of study results corresponding to a plurality of treatments, wherein each study result includes a plurality of data points, and wherein each data point includes a plurality of attributes, Extracting a plurality of questions from the plurality of study results, Converting the plurality of questions into questions in a standardized format, Assigning one or more categories to each of the plurality of questions, Determining groups of questions having the same set of assigned categories among the plurality of questions, Combining the groups of questions, Generating a final data set by combining the plurality of study results, and Training the MLA using the final data set; Receiving a questionnaire response from the patient; Inputting the questionnaire response into the MLA; Outputting, by the MLA, a predicted outcome for each of the plurality of treatments; Generating an interface based on the predicted outcome for each of the plurality of treatments; and Outputting and displaying the interface.

15. A system comprising at least one processor and a memory storing a plurality of executable instructions, which when executed by the at least one processor cause the system to perform the method according to any one of claims 1 to 14.