Techniques for generating predictive outcomes regarding spinal muscular atrophy using artificial intelligence

By employing AI systems to analyze subject records and extract relevant features, the challenges of variability in SMA disease progression and symptom severity are addressed, resulting in improved prediction, clinical study candidate identification, and personalized treatment selection for SMA patients.

JP7682270B2Active Publication Date: 2025-05-23F HOFFMANN LA ROCHE & CO AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023531619
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-26
Filing Date
2021-11-22
Publication Date
2025-05-23
Estimated Expiration
2041-11-22

AI Technical Summary

Technical Problem

Current treatments for spinal muscular atrophy (SMA) are challenging due to the variability in disease progression and symptom severity across patients, making it difficult to define effective treatment workflows and schedules.

Method used

The use of artificial intelligence (AI) systems to predict disease progression, identify candidate subjects for clinical studies, and select personalized therapeutic treatments for SMA patients by analyzing subject records and extracting relevant features.

Benefits of technology

AI-driven approaches enable more accurate prediction of disease progression, improved identification of suitable candidates for clinical studies, and personalized treatment selection, potentially leading to enhanced treatment efficacy and outcomes for SMA patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682270000003
    Figure 0007682270000003
  • Figure 0007682270000004
    Figure 0007682270000004
  • Figure 0007682270000005
    Figure 0007682270000005
Patent Text Reader

Abstract

Techniques are disclosed for using artificial intelligence (AI) to facilitate treatment of subjects diagnosed with spinal muscular atrophy (SMA). The methods and systems disclosed herein relate to techniques for using AI to predict disease progression in subjects diagnosed with SMA, detect potential commonalities across subjects with SMA to identify candidate subjects for new or existing clinical studies, and intelligently select subject-specific therapeutic treatments for treating SMA.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to European Patent Application No. 20211555.6, entitled "Techniques for Generating Predictive Outcomes Relating to Spinal Muscular Atrophy using Artificial Intelligence," filed November 24, 2020, which is incorporated by reference in its entirety for all purposes.

[0002] Field The methods and systems disclosed herein generally relate to techniques for using artificial intelligence (AI) to facilitate treatment of subjects diagnosed with spinal muscular atrophy (SMA). More specifically, the methods and systems disclosed herein relate to techniques for using AI to predict disease progression in subjects diagnosed with SMA, detect hidden commonalities across subjects with SMA to identify candidate subjects for new or existing clinical studies, and intelligently select subject-specific therapeutic treatments for treating SMA. [Background technology]

[0003] background The brain contains specialized cells called motor neurons that control the voluntary movement of over 500 muscles throughout the body. Motor neurons contain axons, which are long fibers that carry signals from the brain along the spinal cord to the target muscle. However, the health of motor neurons depends heavily on the presence of a protein called survival motor neuron (SMN) protein. SMN1, a gene located on chromosome 5, produces sufficient amounts of SMN protein to maintain healthy motor neurons.

[0004] People with a neuromuscular disease called spinal muscular atrophy (SMA) produce insufficient amounts of SMN protein due to a mutation in the SMN1 gene. The lack of SMN protein gradually degenerates motor neurons. However, the degenerated motor neurons prevent brain signals to control voluntary movements from reaching the target muscles. Although SMN1 may not produce enough SMN protein, most people have at least one functional copy of SMN1, called the SMN2 gene. SMN2 can produce about 10-20% of the normal levels of SMN protein, allowing at least some motor neurons to survive. People with SMA generally experience progressive muscle atrophy, primarily of proximal muscles, leading to muscle weakness and muscle breakdown.

[0005] SMA presents a variety of unique challenges.For example, symptoms and severity of symptoms vary widely across subjects with SMA.Therefore, defining treatment workflows for treating subjects is particularly difficult for subjects diagnosed with SMA.Since SMA-related treatments can be highly situational to the disease progression that subjects experience, defining treatment workflows with specific treatment schedules is a difficult and complex task.

[0006] In many cases, defining a schedule for treating a subject is responsive to symptoms rather than predictions. For example, there is a large variability across subjects as to which muscle groups weaken first and to what extent over the course of the disease. Subjects generally experience weakening of the muscle groups that support the spine, putting strain on the respiratory system. However, for some subjects, the atrophy of this muscle group progresses quickly, while for others, the progression is slow. Furthermore, certain subjects experience weakening in the muscle groups that support swallowing, putting strain on daily eating activities. For some subjects, the muscle groups that support swallowing weaken before the muscle groups that support the spine, while for others, the order of degeneration of the muscle groups is reversed. Treating a subject with weakened muscles that support swallowing is very different from treating a subject with weakened muscles that support the spine. Typically, defining treatment for an individual subject involves closely monitoring the subject's symptoms and responding with treatment accordingly.

[0007] In another example illustrating the challenges inherent to SMA, one treatment involves increasing the expression of SMN protein using gene replacement therapy. However, increasing SMN protein expression only leads to improvement in the subject's motor function when performed within the therapeutic window. For example, in animal models, performing SMN restoration therapy is effective in improving motor function only if the therapy is performed within the first three days after birth. The same therapy may not be effective at all if performed more than 10 days after birth. There is a narrow time window for performing a specific SMN therapy to improve motor function, and that time window is contextual for each subject. For a new subject (e.g., a patient), identifying a therapeutic window for SMN protein expression is a technically challenging and complex task. Often, identifying a treatment and treatment schedule for a new subject involves manually comparing many different and complex attributes of the new subject with those of previously treated subjects.

[0008] Symptom severity across subjects with SMA is also highly variable. Symptom severity can be based on a variety of factors, including, for example, the time between symptom onset and diagnosis or treatment, the type of SMA, the subject's daily activities, etc. It is difficult to gain insight into the potential severity and / or timing of future SMA-related events for a given subject with a diagnosed SMA type. This can lead to treatment being administered too late. Studies have found that, on average, SMA type I patients are diagnosed and then treated for 4 months after symptom onset, and SMA type III patients are diagnosed and then treated for 10 months after symptom onset.

[0009] In addition, lack of data availability is another unique challenge in the SMA context. SMA is characterized as a rare disease, affecting roughly 1 in 10,000 births. Experienced physicians may never have the opportunity to treat a subject developing SMA over their career. Even at the local level, the number of previously treated subjects with SMA may be limited. Physicians treating subjects newly diagnosed with SMA may not have access to a sufficient amount of data to inform new treatment schedules for new subjects. Furthermore, using clinical studies to test new treatments on SMA subjects is a challenge given the potentially sparse availability of subjects at the hospital or local level.

[0010] Bai Tian et al. ("EHR phenotyping via jointly embedding medical concepts and words into a unified vector space", BMC Medical Informatics and Decision Making, vol. 18, no. S4, December 1, 2018 (2018-12-01), page 13, XP055804407, DOI:10.1186 / s12911-018-0672-0) disclose using predictive modeling to address the heterogeneity of electronic health record (EHR) data and gain insight into patient phenotyping by incorporating both (1) diagnostic medical codes and (2) words from clinical notes into the same continuous vector space and building connections between them. To evaluate the quality of their vector representations, Tian et al. disclose two types of experiments: (1) phenotype and treatment discovery by evaluating the association between codes and words in the vector space, and (2) prediction of the codes assigned to patients during the second visit by evaluating the association between codes and words in the vector space from the first visit. Tian et al. evaluated six diseases (acute liver failure, female breast cancer, schizophrenia disorder, brain symptoms, depressive disorder, and HIV) for their baseline method, none of which are as rare or difficult to treat as SMA.

[0011] Thus, there is a need for improved personalized selection of SMA treatments, personalized treatment schedules, and formation of subject groups for new clinical studies to improve treatment efficacy for individual subjects diagnosed with SMA. Summary of the Invention

[0012] overview In some embodiments, a computer-implemented method is provided. The computer-implemented method can include searching a subject record associated with the subject and extracting a subset of a set of features included in the subject record. For example, the subject record can include a set of features characterizing the subject. The subject may have previously been diagnosed with spinal muscular atrophy (SMA). Furthermore, each feature of the subset of the set of features can be associated with an SMA trait. The computer-implemented method can also include generating a partial word sequence by combining the subset of the set of features to result in a sequence of one or more words. Each word of the one or more words represents a characteristic of the subset of features. The computer-implemented method can include converting the partial word sequence into a numerical representation using a trained word-to-vector model. The computer-implemented method can also include inputting the numerical representation of the partial word sequence into a natural language processing (NLP) model trained to predict a completed word or phrase to complete the partial word sequence. The computer-implemented method can further include generating a disease progression representing a predicted progression of one or more SMA phenotypes specific to the subject over a period of time based on the completed word or phrase output by the NLP model. The computer-implemented method can also include outputting an indication that the subject is predicted to exhibit one or more SMA phenotypes involved in disease progression.

[0013] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.

[0014] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more processors to perform some or all of one or more of the methods disclosed herein.

[0015] Some embodiments of the present disclosure include a system including one or more processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more processors, cause the one or more processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause the one or more processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0016] The terms and expressions employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described or portions thereof, recognizing that various modifications are possible within the scope of the invention as claimed. Thus, although the invention as claimed has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be made by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief description of the drawings]

[0017] The present disclosure is described in conjunction with the accompanying drawings.

[0018] [Figure 1]A diagram showing a network environment in which a cloud-based application is hosted, according to some aspects of the present disclosure. [Diagram 2] A flowchart showing an example of a process executed by a cloud-based application to deliver a reduced subject record to a user device in connection with a consultation broadcast requesting assistance regarding the treatment of a subject, according to some aspects of the present disclosure. [Diagram 3] A flowchart showing an example of a process for monitoring user integration of a treatment plan definition (e.g., a decision tree or treatment workflow) and automatically updating the treatment plan definition based on the results of the monitoring, according to some aspects of the present disclosure. [Figure 4] A flowchart showing an example of a process for recommending a treatment for a subject, according to some aspects of the present disclosure. [Diagram 5] A flowchart showing an example of a process for obscuring query results to comply with data privacy rules, according to some aspects of the present disclosure. [Figure 6] A flowchart showing an example of a process for communicating with a user using a bot script such as a chatbot, according to some aspects of the present disclosure. [Figure 7] A block diagram showing an example of a network environment for deploying a trained artificial intelligence model to facilitate subject-specific identification of treatments and treatment schedules, according to some aspects of the present disclosure. [Figure 8] A block diagram showing an example of a network environment for deploying a trained artificial intelligence model to predict disease progression for a subject diagnosed with SMA, according to some aspects of the present disclosure. [Figure 9] A block diagram showing an example of a network environment for intelligently identifying candidate subjects for a new or existing clinical study, according to some aspects of the present disclosure. [Figure 10]FIG. 1 is a block diagram illustrating an example of a network environment for deploying a trained artificial intelligence model to intelligently select a treatment, in accordance with some aspects of the present disclosure. [Figure 11] 1 is a flow chart illustrating an example of a process for predicting disease progression in a subject diagnosed with SMA, according to some embodiments of the present disclosure. [Figure 12] 1 is a flow chart illustrating an example of a process for intelligently identifying candidate subjects for a new or existing clinical study, according to some aspects of the present disclosure. [Figure 13] 1 is a flowchart illustrating an example of a process for deploying an artificial intelligence model to facilitate the selection of a treatment to perform on a subject diagnosed with SMA, according to some embodiments of the present disclosure.

[0019] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes the similar components. If only a first reference label is used in this specification, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] Detailed Description I. Overview In Europe, a rare disease is defined as a disease that affects less than 1 in 2,000 people. Although SMA is one of the leading genetic causes of infant mortality in Europe, it remains a rare disease, given that approximately 10,000 people in Europe are affected by SMA. The population of subjects diagnosed with SMA presents several unique challenges. First, experienced physicians may not have had the opportunity to treat subjects with SMA over their careers. Even at the hospital or regional level, the number of previously treated subjects with SMA may be limited. Without experience in diagnosing and treating subjects with SMA, it may be difficult to treat subjects correctly. The small number of subjects affected with SMA limits the ability to gain insight into the pathophysiological mechanisms of SMA and test new treatments.

[0021] Second, SMA is unique in that disease progression and phenotypic severity vary substantially within each SMA type. Although SMA generally degenerates proximal muscles, there are over 500 skeletal muscles that can be affected by SMA. Thus, SMA phenotypes and phenotypic severity range widely across subjects. For example, certain subjects first experience degeneration of the pharyngeal muscles that assist in swallowing activities, while others first experience degeneration of the muscles around the thigh that assist in knee extension during walking activities. The initial treatment of these two groups of subjects is very different. Subjects experiencing difficulty swallowing may be treated with a semi-solid diet by a dietitian, whereas subjects experiencing difficulty walking may be provided with a wheelchair or a cane as a treatment to reduce fatigue of the thigh muscles. Thus, identifying treatments and treatment schedules is often informed by the onset of symptoms rather than predictively, before the onset of symptoms or before the severity of symptoms increases.

[0022] Certain aspects of the present disclosure provide a cloud-based application comprised of an AI system to solve the unique challenges of SMA. AI-based techniques have recently been used to transform the diagnosis and treatment of rare diseases. AI techniques can be used to learn patterns and correlations across various types of datasets (e.g., structured datasets, unstructured datasets, streaming data, etc.) from different sources. For example, AI techniques can be implemented to facilitate the improvement of care lines and the development of new treatments for SMA, even when rare diseases are characterized by a limited number of subjects that are geographically distributed.

[0023] Certain aspects of the present disclosure relate to AI systems configured to perform specific predictive functions, such as predicting disease progression for a particular subject developing SMA, predicting candidate subject groups for evaluation or enrollment in a new or existing clinical study, or predicting a situation-specific treatment schedule for a particular subject.

[0024] As described in more detail with respect to FIG. 8 and FIG. 11, certain aspects of the present disclosure relate to techniques for predicting disease progression of a particular subject diagnosed with SMA. The AI ​​system can train an AI model, such as a natural language processing (NLP) model, or an AI model on word sequences (e.g., sentences) that represent disease progression of SMA patients. Training an NLP model on word sequences that represent disease progression of previously treated SMA patients allows the AI ​​model to learn patterns in various combinations of words within those word sequences. The trained AI model can then receive as input the current health state of the particular subject. In some implementations, the trained AI model treats the current health state of the particular subject as a partial word sequence and then generates a prediction of the next word that is likely to complete the partial word sequence. The predicted next word represents the predicted future disease progression for the particular subject. For example, the predicted disease progression can indicate changes in SMA-specific phenotypes, symptoms, or other disease-related events that the particular subject is predicted to exhibit over the course of the disease.

[0025] As described in more detail with respect to FIG. 9 and FIG. 12, certain aspects of the present disclosure relate to techniques for intelligently identifying groups of subjects predicted to be suitable candidates for enrollment in a new or existing clinical study. For example, a subject is a suitable candidate for enrollment in a clinical study when a treatment being investigated in the clinical study is predicted to be effective for the subject. In some implementations, intelligently identifying subject groups based on high-dimensional subject records includes selectively reducing the dimensionality of the subject records to improve the computational efficiency of subspace clustering of the subject records (e.g., clustering along many dimensions, not just one or two, as in k-means clustering). The reduced dimensional subject records can be used to automatically predict new groups of subjects that may be suitable candidates for a new or existing clinical study. Illustratively, according to certain implementations, if 40 subjects undergoing treatment for SMA at a hospital in Italy experience improvements in motor function after a particular physical therapy, and 17 subjects undergoing treatment for SMA at a research facility in Bogota also experience similar improvements in motor function after the same physical therapy, an AI system can process the data records corresponding to the subjects to detect common latent features across these two groups of subjects. Furthermore, after the AI ​​system detects shared latent features, such as specific biomarkers shared across subjects, the two groups of subjects can be enrolled in existing clinical studies investigating the specific biomarkers, or new clinical studies to investigate the specific biomarkers can be suggested if no existing clinical studies exist.

[0026] As described in more detail with respect to Figures 10 and 13, certain aspects of the present disclosure relate to techniques for intelligently selecting a treatment from a group of available treatments using a treatment selection system trained to maximize a contextually predetermined reward function based on a subject-specific dataset (e.g., a subject record for a particular subject) when selecting a treatment. The output of the trained AI model can predict, for a particular subject specifically suffering from SMA, which treatment to select to achieve the highest likelihood of treatment efficacy, delay in disease progression, extended survival, etc.

[0027] The applications (e.g., operating locally on the device and / or using at least in part the results of computations performed on one or more remote and / or cloud servers) can be used (for example) by a subject with SMA and / or a caregiver caring for a subject with SMA. The applications can perform one or more operations disclosed herein. In some examples, the one or more applications can facilitate communication between a subject with SMA and a caregiver. Such communication can facilitate (for example) alerting a caregiver to abnormal weakness of muscles supporting the spine and / or can facilitate telemedicine (which can be particularly beneficial, for example, when a subject or part of the community is experiencing an epidemic, when a subject has mobility issues, and / or when a subject is physically distant from the caregiver's office).

[0028] II. Overview of Spinal Muscular Atrophy (SMA) Subtypes, Diagnostic Protocols, Related Medical Tests, Progression Assessment, and Available Treatments II. A. Genetic Causes of SMA SMA is a neuromuscular disease characterized by atrophy of skeletal muscles used in voluntary movement. Subjects with SMA experience progressive degeneration of certain nerve cells located in the anterior horn of the spinal cord. These nerve cells, called spinal motor neurons, control muscle movement. The degeneration of motor neurons weakens skeletal muscles, causing generalized weakness in the subject.

[0029] The genetic cause of SMA is a mutation in the survival motor neuron 1 (SMN1) gene located on chromosome 5. In healthy individuals, the SMN1 gene produces the survival motor neuron (SMN) protein, a protein necessary for the survival of motor neurons. The SMN1 gene produces the entire amount of SMN protein required for motor neurons to survive. However, in individuals affected by SMA, the SMN1 gene is mutated due to a deletion or other point mutation occurring in exon 7. The deletion in exon 7 of chromosome 5 in the SMN1 gene causes a decrease in the amount of SMN protein produced by the SMN1 gene or completely prevents the production of SMN protein.

[0030] SMN1 has at least one functional copy, called the survival motor neuron 2 (SMN2) gene, which inefficiently produces the SMN protein that supports healthy motor neurons. For example, the SMN2 gene can only produce about 10-20% of the normal levels of SMN protein necessary for motor neuron survival. Because SMN1 and SMN2 are nearly identical except for a single nucleotide in exon 7, SMN1 and SMN2 produce the same SMN protein in different amounts. However, ultimately, without enough SMN protein, motor neurons cannot function properly and will eventually shrink and die, resulting in weakened and sometimes fatal muscle weakness.

[0031] In some cases, SMA may not be the result of a mutation in the SMN1 gene on chromosome 5, but rather a mutation in another gene on another chromosome.For example, spinal muscular atrophy with respiratory distress (SMARD), sometimes called autosomal recessive distal spinal muscular atrophy (DSMA1), is not caused by a mutation in the SMN1 gene.Instead, SMARD is caused by a mutation in the IGHMBP2 gene, which is located on the long arm of chromosome 11.Subjects with SMARD have severe respiratory distress and muscle weakness.

[0032] Most forms of SMA affect proximal muscles, as do forms associated with mutations on chromosome 5, while other forms of SMA affect distal muscles. Genetic causes of distal muscle atrophy may include mutations in the UBA1 gene located on chromosome X, the DYNC1H1 gene located on chromosome 14, the TRPV4 gene located on chromosome 12, the PLEKHG5 gene located on chromosome 1, the GARS gene located on chromosome 7, and the FBXO38 gene located on chromosome 5. The UBA1 gene listed above may cause X-linked SMA (e.g., XL-SMA or SMAX2). X-linked SMA is similar to SMA type I, but in X-linked SMA, the joints may also be affected. Other symptoms of X-linked SMA may include hypotonia, lack of response to stimuli, and congenital contractures.

[0033] II.B. Types of SMA SMA generally appears early in a subject's lifespan and is the leading genetic cause of death in infants, affecting approximately 1 in every 10,000 births. Approximately 1 in 40-60 people are carriers of a mutation in the SMN1 gene that causes SMA. SMA is inherited in an autosomal recessive pattern, with no significant difference in incidence between ethnicities. If both parents are carriers of a mutation in the SMN1 gene, the newborn has approximately a 25% chance of developing SMA.

[0034] There are four main types of SMA: Types I, II, III, and IV, plus the extremely rare and severe Type 0. SMA types differ based on the age at which symptoms begin and the highest milestones reached in motor development.

[0035] II.B.1.SMA type 0 SMA type 0 is a very rare prenatal form of SMA disease. SMA type 0 is detectable in utero as the subject's fetus will exhibit severe SMA symptoms before birth. For example, the subject's fetus diagnosed with type 0 exhibited generalized osteopenia in the lower extremities.

[0036] SMA type 0 generally has a fatal prognosis with symptom onset in utero, resulting in hypotension, facial paralysis, and potentially death within the first few weeks to months of the subject's infantile life. Homozygous mutations in the SMN1 gene may be the cause of SMA type 0. Available diagnostic tests can show the absence of SMN1 exon 7, demonstrating a homozygous deletion of the SMN1 gene.

[0037] Furthermore, SMA Type 0 subjects exhibit reduced muscle movement in utero, severe asphyxiation, significant hypotonia, respiratory insufficiency at birth, and the need for resuscitation and ventilator support. Additionally, subjects are uniformly observed to have an alert facial appearance.

[0038] II.B.2. SMA type I SMA type I, also known as Werdnig-Hoffmann disease, usually develops within the first few months of life. The most severe form of SMA type I has a rapid and unexpected onset. As the disease progresses, rapid motor neuron death causes inefficiency in major body organs, especially the respiratory system. Pneumonia-induced respiratory failure is the most frequent cause of death. If untreated and without respiratory support, infants diagnosed with SMA type I usually do not survive past the age of 2 years. With appropriate respiratory support, those developing milder SMA type I phenotypes can survive into adolescence and adulthood.

[0039] II.B.3. SMA type II SMA type II, also known as Dubowitz disease, affects individuals who have been able to maintain a sitting position at some point in their lifespan but have never learned to walk unassisted. Onset of SMA type II usually occurs between 6 and 18 months of age, and progression varies widely as some children gradually become weaker while others remain relatively stable. In these children, scoliosis is usually present, and spinal correction can improve breathing. Although life expectancy is reduced, most people with SMA type II survive well into adulthood.

[0040] II.B.4. SMA type III SMA Type III, also known as Kugelberg-Welander disease, is a juvenile form of the disease that usually appears after 12 months of age. Patients with SMA Type III are characterized as having the ability to walk unassisted for at least some time in their lifespan, even if this ability is lost later. With this form of the disease, respiratory problems are infrequent and life expectancy is normal or near normal.

[0041] II.B.5.SMA type IV SMA4 is an adult-onset disease that usually begins after age 30, causing progressive weakening of the leg muscles and requiring subjects to frequently use movement aids. Other complications are rare, and life expectancy is normal.

[0042] II.B.6. Phenotypic Severity Across SMA Subtypes All subjects with SMA have at least one SMN2 copy of the SMN1 gene.For a given subject, the number of SMN2 copies that a subject has is correlated with the severity of SMA phenotype, so the number of SMN2 copies affects the subject's prognosis.For example, the more copies of SMN2 gene a subject has, the milder the symptoms are, and the later the onset of symptoms is.The more copies of SMN2 gene present, the more functional SMN protein is available, and therefore the later the onset of disease symptoms is due to increased survival of motor neurons.

[0043] Thus, the severity of SMA across SMA subtypes is influenced by the number of SMN2 copies a subject has. For example, approximately 70% of SMA-I subjects carry two SMN2 copies, and 82% of SMA-II subjects carry three SMN2 copies. However, subjects developing SMA-III have by far the smallest number of SMN2 copies, between three and four. The SMN1 gene produces approximately 100% full-length mRNA for the SMN protein. However, the SMN2 gene produces a transcript of the SMN protein that lacks exon 7. As a result, approximately 10% of the SMN protein encoded by the SMN2 gene is correctly spliced ​​and encodes a protein identical to SMN1. Thus, the more SMN2 copies there are, the less SMN protein is missing.

[0044] II.C. Diagnosis of SMA Subtypes Diagnosing SMA involves a series of steps. First, a physician may perform an office physical examination and a review of the subject's family history. Certain non-invasive tests may be performed to determine whether genetic testing should be performed. Non-invasive tests help the physician distinguish SMA from other neuromuscular conditions (e.g., muscular dystrophies). For example, if the subject is ambulatory, the physician may perform motor function tests such as the Hammersmith Functional Motor Scale-Extended (HFMSE) test and the 6-minute walk test (6MWT). The HFMSE and 6MWT motor function tests are highly correlated with predicting the severity of the SMA phenotype. In addition, the physician may evaluate muscle weakness and hypotension, which are early signs of the presence of motor function problems associated with SMA. Other evaluations may include evaluating the subject for a history of motor dysfunction, loss of motor skills, proximal muscle weakness, lack of reflexes, tongue fasciculations, and other indicators of degeneration of motor neurons. Additionally, the most common symptoms prompting diagnostic genetic testing for SMA include progressive bilateral muscle weakness (usually in the upper arms and legs), bell-shaped chest, and hypotension associated with absent reflexes. These symptoms are more common and often more severe in subjects with SMA Type 0 and SMA Type I.

[0045] Creatine kinase is an enzyme excreted from degenerated muscle, so a blood test for creatine kinase may indicate the possibility of SMA. Although creatine kinase enzyme levels are above the threshold level for several neuromuscular diseases, the results of such a blood test are nevertheless useful to the physician diagnosing the subject. Creatine kinase enzyme levels may be normal for certain subjects with SMA type I, but creatine kinase levels may be useful for diagnosing SMA types II and III.

[0046] If an early evaluation of the subject's symptoms indicates motor function problems associated with SMA, genetic testing may be performed on the subject. A diagnosis of SMA can only be confirmed by genetic testing, for example, by detecting a biallelic deletion of exon 7 or other point mutations in the SMN1 gene. Although other methods are available for genetic testing, multiplex ligation-dependent probe amplification (MLPA) is often used because it also allows for detection of the number of SMN2 gene copies in a subject. Several MLPA genetic testing kits are commercially available, for example, Asuragen's Amplidex PCT / CE SMN1 / 2 Kit, and Prevention Genetics' Spinal Muscular Atrophy via MLPA of SMN1 and SMN2 test.

[0047] In addition to or instead of genetic testing, electromyography (EMG) testing may be performed. EMG testing measures the electrical activity of a muscle or muscle group, and muscle biopsy and / or creatine kinase (CPK) testing may also be used to diagnose SMA, as well as to differentiate the diagnosis from other types of neuromuscular diseases if necessary.

[0048] In addition to diagnostic testing of symptomatic individuals, prenatal genetic testing and newborn screening can be performed to diagnose early stages of severe forms of SMA, such as SMA Type 0 and SMA Type I.

[0049] II.D.Newborn Screening for SMA Neonatal screening for SMA can be part of routine screening of newborns during the first few days of an infant's life. Neonatal screening for SMA is a genetic test on the newborn's blood. The genetic test involves evaluating a blood sample of the newborn subject for abnormalities associated with the SMN1 gene. While genetic testing on blood is invasive, neonatal screening for SMA uses the same blood sample already taken for screening of other disorders. When the results of the newborn blood sample analysis show that the newborn is missing a portion of the SMN1 gene located on chromosome 5, the newborn is likely to have or is at high risk of having SMA. Further testing can be performed to determine whether the infant subject has SMA and, if so, to identify the infant subject's target treatment.

[0050] According to certain studies, for example, by screening all newborns in the United States for SMA, approximately 364 newborns with the disorder would likely be detected each year. Furthermore, widespread newborn screening could prevent approximately 50 newborns from needing mechanical ventilation and approximately 30 deaths due to SMA type I. Furthermore, newborn screening is important because treatment early relative to the onset of symptoms is more effective than treatment later relative to the onset of symptoms.

[0051] Newborn screening programs can also be used to identify presymptomatic newborns. In many cases, if therapeutic treatment is initiated before the onset of symptoms, treatment can prevent irreversible motor neuron damage. Homozygous mutations in the SMN1 gene have been shown to be accurately detectable in blood samples from newborns, proving that screening newborns for SMA using blood samples taken on the day of birth is a useful screening approach.

[0052] Neonatal screening for SMA has limitations. For example, point mutations in the SMN1 gene of certain subjects are difficult to detect. Prenatal screening and treatment may be suitable for certain subjects. For example, in mouse cell models, the SMN protein supports neural differentiation and the formation of neuromuscular junctions in utero. The SMN protein is also involved in neurodevelopment and synaptogenesis. Thus, prenatal screening for SMA for certain subjects, and potentially prenatal or neonatal treatment for subjects diagnosed with SMA type 0, may be feasible and useful for early detection and treatment. For certain subjects, prenatal screening for SMA may be feasible, and fetal gene replacement therapy may be given, such as administering an adeno-associated virus (AAV) that can infect and deliver the SMN1 gene to the cells of the subject. Furthermore, for certain subjects, chorionic villus sampling or amniocentesis may be performed at 10-14 weeks or 15-20 weeks of gestation. This sampling is indicated to identify the possibility or risk that the fetus has developed SMA. However, prenatal screening also has its own challenges and limitations. Prenatal screening is invasive and can pose risks to the mother and fetus. Non-invasive prenatal screening for SMA is possible. In certain studies, fetal trophoblast cells or cell-free fetal DNA were isolated from maternal blood samples and evaluated to detect SMA.

[0053] II. Clinical Symptoms of E.SMA Symptoms vary depending on the type of SMA, the stage of the disease, and individual factors, but signs and symptoms of SMA include delayed gross motor skills, difficulty standing, sitting, or walking, frog-legged posture when sitting, areflexia (especially in the limbs), generalized muscle weakness, hypotonia, relaxation, tendency to limp, loss of respiratory muscle strength, gastrointestinal problems, coughing, buildup of secretions in the lungs or throat, respiratory distress, bell-shaped body, scoliosis, twitching of the tongue (fasciculations), difficulty sucking or swallowing, and poor feeding.

[0054] II.F.SMA Treatment Treatment of SMA varies based on severity and type. In the most severe forms (SMA0 and SMA1), individuals experience maximal muscle wasting and require prompt intervention. In contrast, individuals developing SMA4 or adult-onset SMA may not require treatment until much later in life. Treatment of severe SMA is often difficult, as the timeline for diagnosis and treatment can be very short due to the patient's age or current health status. SMA can become life-threatening very quickly, as it is a rapidly progressive disease that affects the muscles involved in swallowing, breathing, and feeding. Thus, early diagnosis and aggressive treatment of individuals developing SMA0 and SMA1 is important.

[0055] Currently, nusinersen (Spinraza®), an antisense oligonucleotide that modifies alternative splicing of the SMN2 gene, is used to treat SMA. SMN2 splicing modulation causes the SMN2 gene to produce increased amounts of full-length SMN protein. Nusinersen is administered directly to the central nervous system via intrathecal injection to prolong survival and improve motor function in infants with SMA. Other SMN2 gene splice modulators that increase the availability of SMN protein in motor neurons include orally administered small molecules such as Branaplam (LMI070, NVS-SM1) and Evrysdi (risdiplam, RG7916, R07034067) (F. Hoffman-La Roche AG). Evrysdi can be administered to treat types 1, 2, and 3 SMA in adults and children aged 2 months or older. Zolgensma® (onasemnogene abeparvovec) is a gene therapy that uses self-complementary adeno-associated virus type 9 (scAAV-9) as a vector to deliver the SMN1 transgene. The treatment has been approved in the United States as an intravenous formulation to treat patients under the age of 2.

[0056] Other treatments include the neuroprotective compound olesoxime (F. Hoffman-La Roche AG) and the SMN 2 gene activator albuterol.

[0057] Depending on the severity and type of SMA, respiratory support is often used to manage SMA. In some cases, respiratory problems are caused by the accumulation of airway secretions. Manual or mechanical chest physiotherapy with postural drainage can be used to remove secretions. In addition, manual or mechanical cough assist devices, or non-invasive ventilation (BiPAP) can be used. In more severe cases, a tracheotomy can be performed.

[0058] Nutritional support may also be essential, as ingestion, jaw opening, chewing, and swallowing may be impaired due to SMA. Other nutritional problems include food not passing through the stomach quickly enough, gastric reflux, constipation, vomiting, and bloating. Thus, SMA patients, especially SMA1 patients, may require a feeding tube or gastrostomy. Metabolic abnormalities caused by SMA may impair beta-oxidation of fatty acids in muscle, leading to organic acidemia and consequent muscle damage, especially during fasting. Individuals developing SMA, especially those with more severe forms of the disease, should choose softer foods to avoid aspiration, reduce fat intake, and avoid prolonged fasting.

[0059] Management of SMA may also include treatment of orthopedic problems resulting from disease progression. Skeletal problems associated with the weak muscles of SMA include stiff joints, hip dislocation, spinal deformity, osteopenia, increased risk of fractures, and pain. Weak muscles may result in the development of kyphosis, scoliosis, and / or joint contractures. Spinal fusion is sometimes performed in people with SMA I / II to relieve the pressure of the deformed spine on the lungs. In addition, mobility devices (e.g., wheelchairs, crutches, canes, walkers), range of motion exercises, and bone strengthening can help prevent orthopedic complications. Occupational and physical therapy are also helpful. Orthotic devices, e.g., ankle-foot braces, and thoracic-lumbar-sacral braces, can also be used to support the body and aid in walking.

[0060] In recent years, survival rates for SMA patients have increased with available drug treatments as well as aggressive respiratory, orthopedic, and nutritional support.

[0061] II.F.1 Treatment Time Window for Effective Restoration of SMN Protein Levels Early treatment of SMA is important. For example, studies have shown that preemptive treatment of SMA subjects before or around the onset of symptoms can improve motor function and quality of life. Restoring SMN protein levels early, for example, in some cases between 1 and 3 days of age, is more effective at improving motor function than restoring SMN protein levels after 5 days of age.

[0062] II.G.SMA Disease Progression The various types of SMA are degenerative. SMA may present differently across the various types of SMA.

[0063] II.G.1.SMA Type I Disease Progression For a given SMA type, the subject's proximal muscles degenerate first. Then, given the degeneration of the subject's proximal muscles, the distal muscles tighten. For example, the subject's thigh muscles may weaken first, which causes the subject's leg muscles to tighten. For most subjects with SMA, the hands maintain strength the longest, so that daily tasks (e.g., computer use) are feasible even as the disease progresses.

[0064] SMA can result in scoliosis (e.g., an "S" shaped curvature of the spine) as the subject's muscles that support the spine weaken over time. Subjects with scoliosis may present with uneven shoulders and hips, or the hip or shoulder on one side may be larger than the corresponding hip or shoulder on the other side of the subject. Given the weakness of the muscles that support the spine, subjects with SMA often experience respiratory problems that can be life-threatening.

[0065] For children who have SMA type I, the disease is also called Werdnig-Hoffmann disease, a severe form of SMA. Werdnig-Hoffmann disease can be diagnosed between birth and 6 months of age. SMA type I in certain children can result in significant muscle weakness, resulting in the child being unable to sit or stand on their own. Children may also experience difficulty in aspirating or swallowing, which can lead to malnutrition.

[0066] II.G.2.SMA Type II Disease Progression The disease progression in children with SMA type II varies significantly. Some children are able to sit up on their own early in life but not later, such as in their teenage years. Additionally, ambulatory type II subjects may experience difficulty walking a few feet without assistance. Fingers may begin to tremble. Tendon reflexes may also be diminished. By the mid-teens or later, SMA type II subjects are usually unable to sit up on their own. As with other types of SMA, subjects with type II often experience muscle weakness in the muscles near the spine, causing potentially life-threatening breathing problems.

[0067] II.G.3. SMA Type III Disease Progression SMA type III, also known as Kugelberg-Welander syndrome, can be diagnosed as early as 18 months of age. Symptoms may be detected earlier. For example, children with type III may be able to walk but experience difficulty climbing or climbing stairs. Children also experience difficulty sitting up from a supine position. Additionally, as with other forms of SMA, subjects with type III are likely to exhibit problems with breathing or other respiratory issues as the muscles that support the spine degenerate. For some subjects, SMA type III may be diagnosed between the ages of 20 and 30, and in these situations, disease progression may be slow. However, adults with SMA type III are usually ambulatory and may experience difficulty walking as they age.

[0068] III. Overview of Cloud-Based Network Architecture for Deploying Intelligent Functions The techniques relate to configuring a server to execute code that enables a user of an entity (e.g., a physician) to perform machine learning or artificial intelligence techniques using a subject record. The subject record includes a complex combination of data elements that characterize the subject. By way of illustration, the subject record may include a combination of thousands of data fields. Some data fields may include fixed non-numeric values ​​(e.g., the subject's ethnicity), other data fields may include unstructured text data (e.g., notes prepared by the physician), other data fields may include a time-varying series of collected measurements (e.g., glycosylated hemoglobin measurements taken 2-4 times per year), and other data fields may include images (e.g., an MRI of the subject's brain). Because machine learning and artificial intelligence models are often configured to process data in numerical or vector formats, the complexity and distribution of data types and formats in the subject record makes processing the subject record technically difficult, if not impossible. In light of the technical problems of this objective, certain aspects and features of the present disclosure relate to the conversion of the subject record into a converted representation, such as a vector representation, that characterizes the various data elements of the subject record.

[0069] The technique relates to converting non-numeric values ​​contained in the subject record into a numerical representation (e.g., feature vector) that can be input to a machine learning or artificial intelligence model to generate a predictive output. A server executing the code provides a technical effect of solving a desired technical problem by converting the subject record into a converted representation that is consumable by the machine learning or artificial intelligence model. "Consumable" may refer to data in a format or form that the machine learning or artificial intelligence model is configured to process to generate a predictive output. The machine learning or artificial intelligence model is not configured to process the subject records (as they exist stored in a data registry) due to the complex combination of data elements in multiple different data formats and data types contained in individual subject records. For example, for a given subject record, a data element may include a long sequence of events (e.g., an immunization record), another data element may include measurements taken from the subject (e.g., vitals), yet another data element may include text entered by a user (e.g., notes recorded by a physician), and yet another data element may be an image (e.g., an x-ray). Limited or overly simplistic analysis may be performed on the subject records (before any transformations), such as grouping subjects based on the values ​​of data elements (e.g., age groups). However, as the complexity and size of subject records reach big data scale, limited or overly simplistic analysis becomes problematic or infeasible. To process and extract analytical assessments from subject records at big data scale, machine learning or artificial intelligence techniques can be used for data mining of subject records. However, the machine learning or artificial intelligence models are configured to receive numerical or vector inputs. For example, a clustering operation such as k-means clustering is configured to receive vectors as input.Thus, to perform clustering operations on subject records, the present disclosure provides a technical effect of solving a technical problem of interest by transforming the subject records into a transformed representation, such as a numeric vector representation, that is consumable by a machine learning or artificial intelligence model. Intelligent analytics can be performed on the subject records in the transformed representation. Non-limiting examples of intelligent analytics (performed on the server executing the code) may include automatically detecting subject groups using clustering techniques, generating output that predicts a particular outcome based on values ​​of data elements in the subject records, and identifying existing subject records that are similar to a given or new subject record.

[0070] For example, and by way of non-limiting example only, a subject's subject record may include four data elements: a first data element includes a unique code representing a diagnosis of a condition; a second data element includes an MRI of the subject's brain; a third data element includes a time-varying series of measurements, such as blood pressure measurements over the course of a year; and a fourth data element includes unstructured notes, e.g., notes of symptoms detected by examining or performing one or more tests. According to a particular implementation, each of the first data element, the second data element, the third data element, and the fourth data element may be converted to a transformation representation (e.g., a vector). The technique used to convert the values ​​contained within the four data elements may depend on the type of data contained in the data elements. In the case of the first data element, for example, the unique code representing a diagnosis may be represented as a fixed-length vector, such that the size of the vector is determined by the size of the vocabulary of codes, and each code in the vocabulary is represented by a vector element of the fixed-length vector. The one or more unique codes contained within the first data element may be compared to the vocabulary of codes. If the unique code matches a code in the vocabulary, a vector element at a position in the vector corresponding to the unique code may be assigned a "1" and all remaining vector elements of the vector may be assigned a "0". In light of the above, a first vector may be generated to represent the value of the first data element. As another example, for the second data element, a latent space representation of the image may be generated using a trained autoencoder neural network. The latent space representation of the input image may be a low-dimensional version of the input image. The trained autoencoder neural network may include two models: an encoder model and a decoder model. The encoder model may be trained to extract a subset of salient features from a set of features detected in the image. The salient features (e.g., keypoints) may be regions of high intensity in the image (e.g., edges of objects). The output of the encoder model may be a latent space representation of the input image. The latent space representation may be output by a hidden layer of the trained autoencoder model, and thus the latent space representation may be only interpretable by the server.The decoder model may be trained to recover the original input image from a subset of the extracted salient features. The output of the encoder model may be used as a feature vector representing pixel values ​​of the image contained in the second data element. In light of the above, a second vector (e.g., a latent space representation) may be generated to represent the image contained in the second data element. As another example, for the third data element, the time-varying sequence of measurements may be represented numerically. In some implementations, the time-varying sequence may be represented by a sum of instances where measurements were taken from the subject. In other implementations, the time-varying sequence may be represented numerically using an average, mean, or median value of the measurements taken over instances of measurements that occurred during a period of time (e.g., a year). In other implementations, the frequency of the measurements may be calculated and used to numerically represent the time-varying sequence of measurements. In light of the above, a third vector may be generated to represent the time-varying sequence of values ​​contained within the third data element. As yet another example, for the fourth data element, the notes entered by the user may be processed and vectorized using any number of natural language processing (NLP) text vectorization techniques. In some implementations, a word-to-vector machine learning model, such as a Word2Vec model, may be performed to convert the notes included in the fourth data element into a single vector representation. In other implementations, a convolutional neural network may be trained to detect words or numbers in the text indicating symptoms, treatments, or diagnoses from the notes included in the fourth data element. In light of the above, a fourth vector may be generated to represent the text of the notes included in the fourth data element as a vector representation. In this manner, the final feature vector representing the entire subject record may be a vector of vectors, including a concatenation of the first vector, the second vector, the third vector, and the fourth vector. In other examples, an average of the first vector, the second vector, the third vector, and the fourth vector may be used to numerically represent the entire subject record.Other combinations of the first vector, second vector, third vector, and fourth vector may be used to generate a final feature vector that numerically represents the entire subject record.

[0071] In some implementations, instead of generating vectors to numerically represent each data element of the subject record, techniques may be performed to reduce the dimensionality of the subject record by identifying and selecting a subset of data elements from the set of data elements. The subset of data elements may represent "important" data elements, and the "importance" of the data elements is determined based on a prediction using a feature extraction technique such as Singular Value Decomposition (SVD). For example, converting the subject record into a converted representation consumable by machine learning and artificial intelligence models may include performing one or more feature extraction techniques on non-numeric values ​​contained in the data elements of the subject record to generate a feature vector that numerically represents the decomposed version of the non-numeric values. In some implementations, the feature extraction technique may include, for example, reducing the dimensionality of a set of data elements of the subject record (e.g., each data element representing a characteristic or dimension of the subject) to an optimal subset of features that can be used to predict, for example, an outcome or event. Reducing the dimensionality of the set of data elements may include reducing N data elements to a subset of M elements, where M is less than N. In these implementations, each element of the subset of M elements may be converted to a numerical value. In some implementations, a feature vector may be generated to represent the N data elements of the subject record. The feature vector may include a vector for each data element of the set of data elements. For example, the feature vector may be a numerical representation of a complex combination of the data elements of the subject record. Each non-numeric value in the data elements of the subject record may be vectorized to generate a representative vector. The vectors representing the set of data elements in the subject record may be concatenated or combined (e.g., as an average or weighted average) to generate a feature vector that numerically characterizes the entire set of data elements of the subject record. The feature vector is consumable by a trained machine learning or artificial intelligence model. Once a feature vector is generated for a subject record, the subject record can be evaluated individually or within a group of other subject records using machine learning and artificial intelligence techniques.After the feature vectors representing each subject record are generated and stored, the feature vectors of the subject records stored in the central data store can be input to a machine learning or artificial intelligence model or other extended analysis can be performed on the numerical representation of the subject records. For example, two different subject records can be compared with respect to one or more dimensions. A dimension can represent a feature or data element of the subject record along which a comparison between two or more subject records is made. For example, a data element of a first subject record includes text entered by a first user (e.g., a physician) that describes a symptom of the first subject. The text (e.g., the value of the data element of the first subject record) can be vectorized using the text vectorization techniques described above (e.g., Word2Vec) to generate a first vector that numerically represents the text associated with the data element. The text vectorization technique can generate an N-dimensional word vector for each word contained in the text. A matching data element of a second subject record (e.g., a data element of another subject record that also includes text entered by a physician that describes a symptom of another subject) may include text entered by a second user that describes a symptom of the second subject. The text (e.g., the value of the data element of the second subject record) may be vectorized using the text vectorization techniques described above to generate a second vector (e.g., an N-dimensional word vector) representing the text associated with the data element. The server may compare the first vector to the second vector in Euclidean or cosine space to quantify the similarity or dissimilarity between the first subject record and the second subject record with respect to at least the dimension of the patient's presentation of symptoms. If the first vector and the second vector are close to each other in Euclidean space (e.g., if the Euclidean distance between the first vector and the second vector is small), then the symptoms experienced by the first subject (as described in the text of the data element) are likely to be similar to the symptoms experienced by the second subject (as described in the text of the data element).However, if the Euclidean distance between the first vector and the second vector is large or exceeds a threshold distance (e.g., or if the Euclidean distance is above a threshold), then the symptoms experienced by the first subject can be predicted to be different from the symptoms experienced by the second subject.

[0072] In some implementations, the server may be configured to execute an application that allows users of an entity to build a data registry that serves to store subject records for subsequent processing. The data of the subject records may include unstructured data, such as electronic copies of physician notes and / or answers to open-ended questions. The unstructured data may be captured in the data registry by mapping portions of the unstructured data to fixed portions (e.g., data elements) of structured data records. The structure of the structured data records may be defined using specifications from modules that correspond (for example) to a particular use case (e.g., a particular disease, a particular test, etc.). For example, each word (e.g., text) of the unstructured note data may be converted to a numerical representation, and the various numerical representations associated with the unstructured note data may be decomposed (e.g., using SVD) to find words that describe a particular set of symptoms exhibited by the subject. The decomposition of the numerical representation of the unstructured note data may remove non-informative words such as "and," "the," "or." The remaining words represent a particular set of symptoms. Some portions of the note data may be irrelevant with respect to a data element in the structured data and / or may be more or less specific than the data contained in the data element. In some examples, various mappings (e.g., mapping "balance disorder" symptoms to "neurological" symptoms), natural language processing, or interface-based approaches (e.g., prompting a user for new information) may be used to obtain a structured data record. An interface may also be used to receive input identifying new information about a new or existing subject, and the interface may include input components and selection options that map to the structure of the data record.

[0073] Further, the technique relates to configuring a cloud-based application to convert non-numeric values ​​contained in data elements of the subject record to a numerical representation, such that the cloud-based application can use the numerical representation (e.g., the converted representation) of the subject record stored in the data registry to perform intelligent analytical functions. The conversion of non-numeric values ​​of data elements of the subject record to a numerical representation may depend on the type of data contained in the data element. For example, for a data element that includes text, such as notes taken by a user, the text may be converted to a numerical representation of the text using natural language processing techniques such as Word2Vec or other text vectorization techniques. As another example, for a data element that includes image frames of an image (e.g., an MRI) or video (e.g., an ultrasound video), each image or image frame may be converted to a numerical representation (e.g., a vector) using a trained autoencoder neural network that is trained to generate a latent space representation of the input image. The reduced representation (e.g., the latent space representation) of the input image can serve as a vector that numerically represents the input image. As yet another example, for a data element that includes a time-varying sequence of information (e.g., events occurring over a period of time), the time-varying information can be represented as a numerical representation using several exemplary transformations. In some examples, a count of an event may be used as a vector representing the time-varying information. In other examples, a frequency or rate of an event occurring (e.g., once a week, once a month, once a year, etc.) may be used as a vector representing the time-varying information. In still other examples, an average or combination of measurements associated with each event in the time-varying information may be used as a vector representing the time-varying information. The disclosure is not limited to these examples, and thus other numerical representations of the time-varying information may be used as vectors representing the numerical representations. The intelligent analysis function may be performed by running a machine learning or artificial intelligence model trained using the data records. The model output may be used to indicate a particular analysis extracted from the data records.

[0074] In some examples, transmission of data from subject records may be provided to develop treatment plans for individual subjects. For example, subject record information (e.g., conforming to data privacy restrictions by selecting omission and / or obfuscation of data) may be broadcast and / or transmitted to a selected group of user devices. For example, the broadcast may be transmitted to user devices associated with similar data records in response to input from a user corresponding to a request to initiate a consultation with a user associated with a similar subject. If the user receiving the broadcast accepts the consultation request (through providing a corresponding input), a secure data channel may be established between the users, and potentially more subject records may be shared (e.g., while conforming to data privacy restrictions applicable to the two users). Subject records similar to a given subject may be identified by performing a nearest neighbor technique using vector representations of two or more subject records. The nearest neighbor technique may be performed by comparing vectors of individual data elements across multiple subject records (e.g., nearest neighbors may be determined in relation to a dimension or feature of the subject record). Alternatively, the nearest neighbor technique may be performed by comparing an overall vector characterizing an entire subject record to an overall vector characterizing another entire subject record. The overall vector may be a concatenation of the individual vectors representing the values ​​of the data elements, or it may be an average or combination of the individual vectors representing the values ​​of the data elements.

[0075] As another example, the one or more processed data records may be returned in response to a query for subject records matching certain constraints. In some examples, a first user may submit a query identifying a first subject record. The query may correspond to a request to identify other subject records similar to the first subject record. The server may convert the first subject record into a converted representation using certain conversion techniques described above and herein. Alternatively, the converted representation of the first subject record may have been previously generated and stored in the database. Regardless of whether the converted representation of the first subject record is generated before or after receiving the query, converting the first subject record into the converted representation of the first subject record may include generating a vectorization of one or more non-numeric values ​​of data elements of the first subject record. Vectorizing the one or more non-numeric values ​​contained within the first subject record may include generating a numeric vector representation for each value (e.g., for each non-numeric text, such as a note) contained in each data element of the first subject record. The various vector representations may be concatenated or otherwise combined (e.g., averaged) to generate a feature vector that represents the entire first subject record. The vector representation that numerically represents the first subject record may be compared to vector representations of other subject records in a domain space (e.g., Euclidean or cosine space). For example, when the Euclidean distance between the two vector representations is within a threshold distance, the two subject records associated with the two vector representations may be interpreted (e.g., by a server) as similar with respect to at least one or more dimensions.

[0076] For each data element in the subject record, the technique used to generate a vector representation of the values ​​associated with the data element may depend on the type of data associated with the data element. In some examples, the data elements of the subject record may be associated with one or more images, such as an x-ray of the subject. Feature extraction techniques may be performed to generate a vector representation of each image associated with the data element. For example, the server may be configured to execute a trained autoencoder neural network to generate a reduced dimensional version of the image. The trained autoencoder neural network may include two models, an encoder model and a decoder model. The encoder model may be trained to extract a subset of salient features from a set of features detected in the image. The salient features (e.g., keypoints) may be regions of high intensity in the image (e.g., edges of objects). The output of the encoder model may be a latent space representation of the input image. The latent space representation may be output by a hidden layer of the trained autoencoder model, and thus the latent space representation may be interpretable only by the server. The subset of salient features of the latent space representation characterizing the subject record may be compared to the subset of salient features of the latent space representation characterizing another subject record to yield specific analytical insights. The decoder model may be trained to recover the original input image from the extracted subset of salient features. The output of the encoder model may be a vector representation of the data elements associated with the image contained in the subject record. In another example, a keypoint matching technique may be performed to match keypoints of an image contained in a data element of a first subject record with keypoints of another image contained in a data element of a second subject record. The vector representation (e.g., latent space representation) of the input image is consumable by a machine learning or artificial intelligence model, such that two different subject records (each containing an image) may be compared to each other to determine similarities or dissimilarities between the two different subject records.

[0077] For example, and by way of non-limiting example only, a magnetic resonance image (MRI) of the subject's brain is captured. The MRI is stored in a subject record associated with the subject. The server is configured to generate a transformed representation, such as a vector representation of the MRI included in the subject record, using feature extraction techniques such as keypoint detection, auto-encoding into a latent space representation, SVD, and other suitable computer vision techniques. The vector representation of the data element including the MRI may be concatenated or otherwise combined (e.g., averaged) with the vector representations of each remaining data element of the set of data elements to generate a feature vector that characterizes the entire subject record. A user may access the application to query a database of other subject records to search for sets of other subject records that include MRIs similar to the MRI of the subject's brain. Identifying other subject records that are similar to the subject record (at least in terms of similarity between the MRIs) may include computing k-nearest neighbors of the subject record. For example, the transformed representation may be plotted (visually or internally by the computing system) on a domain space, such as Euclidean space or cosine space. The transformed representation of each other subject record may also be plotted (visually or internally by the computing system). Nearest neighbor techniques may be performed to compare the vector representation of the subject record with vector representations of other subject records to identify k nearest neighbors to the subject vector. The identified k nearest neighbors may be predicted to have MRIs similar to the MRI of the subject's brain. Each other subject record identified as a nearest neighbor may be identified and retrieved for further evaluation or processing using an application.

[0078] In some implementations, the computing system can perform data processing techniques (e.g., nearest neighbor techniques) to identify similar subject records. In this search, various data elements may be weighted differently (e.g., according to predefined data element weightings, the importance of matching various data elements, and / or user input indicating the prevalence of particular data element values ​​across the subject record set). When searching the set of records for potential matches, some records may lack values ​​for various data elements. In these cases, it may be determined (for example) that the data element values ​​do not match and / or that the data elements may not be weighted when evaluating potential matches. Handling of missing values ​​may depend on the distribution of the values ​​of the data elements across the set of records and / or the values ​​of the data elements in the query.

[0079] Furthermore, some techniques relate to defining and using a set of rules to identify a potential treatment regimen for a subject, taking into account a set of symptoms identified within the subject record. For example, a subject record of interest can represent a subject who has recently experienced three symptoms: upper respiratory infection, fever, and sore throat. The three symptoms may be written as text within the data elements of the subject record of interest (e.g., the separation between words marked by tags such as semicolons). A server, such as server 135, can individually input the text "upper respiratory infection", "fever", and "sore throat" into a trained Word2Vec model or other vector model from text, such as a vocabulary mapping. The Word2Vec model may be trained to generate a vector representation for each word representing a symptom. The vector representations for the three symptoms may be averaged to generate a single vector representation for the "symptoms" data element of the subject record of interest. The single vector representation for the "symptoms" data element of the subject record of interest may be processed to identify other subject records that include synonyms in the "symptoms" data element. Each subject record stored in the database may be associated with an existing "symptoms" data element that has been converted to a numerical representation, such as a vector. The vectors for the "symptoms" data element may be plotted and compared to the vector for the "symptoms" data element of the subject record of interest. The server can identify the vector that is closest to the vector characterizing the "symptoms" data element. The vector of the "symptoms" data element that is closest to the vector of the subject record of interest may be predicted to be similar to the subject. The subject record associated with the vector that is closest to the vector of the subject record of interest is identified and further evaluated to determine the treatment regimen to be provided to that subject. The treatment provided to the subject associated with the vector that is closest to the vector for the subject record of interest may be used as a potential treatment regimen for treating the subject of interest. Furthermore, each potential treatment regimen may be weighted by the responsiveness experienced by other subjects. The potential treatment regimens may be classified according to the responsiveness experienced by other subjects.

[0080] The set of rules may be defined based on user interaction with a user interface, which may include specification of a particular criterion and an associated particular medical procedure, and / or selection of one or more previously defined rules (specifying the criterion and procedure). For example, one or more existing rules may be presented via the interface, and a user may select a rule for incorporation into a rule base associated with an account associated with the user. The one or more rules may be selected from among a set of rules defined by multiple users (e.g., associated with one or more institutions) and / or may be generated based on rules generated by multiple users. When a user selects a rule for incorporation into the rule base, the application may generate a feedback signal to the cloud server 135. The feedback signal may include metadata associated with the user's selection. The metadata may indicate whether the rule was incorporated into the rule base without modification or with modification. If the rule base was modified, the metadata should indicate what modifications were made to the rule. The metadata may also indicate whether the rule was rejected, deleted, or otherwise determined to be not useful to the user. For example, as a non-limiting example, the computing system may detect that rules associating one or more particular types of symptoms and / or test results with a given treatment are defined and / or selected relatively frequently by a user, and the computing system may then generate general rules for the particular types of symptoms and / or test results and treatments. General rules may be defined to have (for example) the most restrictive, the most inclusive, or median criteria. In some examples, the user's rule base may be processed to detect overlaps of any criteria between the rules. Upon identifying overlaps, a warning identifying the overlaps may be presented. The rules in the rule base may be used to evaluate and classify subject records to define populations associated with the subject records.Evaluating the subject record using the rules may be implemented as a decision tree, for example, in that a first criterion of the rule is compared to an attribute contained in the subject record. If the first criterion is met, the next criterion is compared to the attribute contained in the subject record. If the next criterion is met, the comparison continues for each criterion contained in the rule. The comparison may continue even if the next criterion is not met. In this case, the fact that the criterion (and any other criteria contained in the rule) is not met is stored and presented to the user device along with the criteria that were met.

[0081] Thus, embodiments of the present disclosure provide a cloud-based application configured to exchange subject information with external entities without violating data privacy regulations. The cloud-based application is configured to automatically evaluate data privacy regulations involved in sharing subject information across various jurisdictions. The cloud-based application is configured to execute protocols that obfuscate or otherwise modify the subject information, thereby algorithmically ensuring compliance with data privacy regulations.

[0082] IV. Network Environment for Hosting Cloud-Based Applications Configured with Intelligent Capabilities 1 illustrates a network environment 100 in which an embodiment of a cloud-based application is hosted. The network environment 100 may include a cloud network 130 that includes a cloud server 135, a data registry 140, and an AI system 145. The cloud server 135 may execute source code underlying the cloud-based application. The data registry 140 may store data records captured from or identified using one or more user devices, such as a computer 105, a laptop 110, and a mobile device 115.

[0083] The data records stored in the data registry 140 may be structured according to a skeleton structure of fixed parts (e.g., data elements). The computer 105, the laptop 110, and the mobile device 115 may each be operated by different users. For example, the computer 105 may be operated by a doctor, the laptop 110 may be operated by an administrator of an entity, and the mobile device 115 may be operated by a subject. The mobile device 115 may connect to the cloud network 130 using the gateway 120 and the network 125. In some examples, the computer 105, the laptop 110, and the mobile device 115 are each associated with the same entity (e.g., the same hospital). In other examples, the computer 105, the laptop 110, and the mobile device are associated with different entities (e.g., different hospitals). The user devices of the computer 105, the laptop 110, and the mobile device 115 are examples for illustrative purposes, and thus the present disclosure is not limited thereto. The network environment 100 may include any number or configuration of user devices of any device type.

[0084] In some embodiments, the cloud server 135 can obtain data (e.g., subject records) for storage in the data registry 140 by interacting with either the computer 105, the laptop 110, or the mobile device 115. For example, the computer 105 interacts with the cloud server 135 using an interface to select locally stored (e.g., stored on a network local to the computer 105) subject records or other data records for inclusion in the data registry 140. As another example, the computer 105 interacts with the interface to provide the cloud server 135 with an address (e.g., a network location) of a database that stores the subject records or other data records. The cloud server 135 then retrieves the data records from the database and populates the data registry 140.

[0085] In some embodiments, the computer 105, the laptop 110, and the mobile device 115 are associated with different entities (e.g., a medical center). The data records that the cloud server 135 retrieves from the computer 105, the laptop 110, and the mobile device 115 may be stored in different data registries. The data records from each of the computer 105, the laptop 110, and the mobile device 115 may be stored within the cloud network 130, but the data records are not intermingled. For example, the computer 105 cannot access the data records retrieved from the laptop 110 due to constraints imposed by data privacy regulations. However, the cloud server 135 may be configured to automatically obscure, obfuscate, or mask portions of the data records when those data records are queried by different entities. In this manner, data records captured from an entity may be released to the different entities in an obfuscated, obfuscated, or masked form to comply with data privacy regulations.

[0086] As data records are collected from computers 105, laptops 110, and mobile devices 115, the data records may be used as training data to train machine learning or artificial intelligence models to provide the intelligent analytic functions described herein. The data records may also be available for query by any entity, given that when a user device associated with an entity queries the data registry 140 and the query results include data records originating from different entities, those data records may be provided or exposed to the user device in an obfuscated format that complies with data privacy regulations.

[0087] The cloud server 135 may be configured in a special manner to execute code that, when executed, causes an intelligent function to be performed using the transformed representation of the subject record (e.g., a vector that numerically represents information stored in the subject record). For example, the intelligent function may be performed by executing code using the cloud server 135. The executed code may represent a trained neural network model. The neural network model may be trained to perform an intelligent function, such as predicting a subject's responsiveness to a treatment regimen, identifying similar patients, generating treatment regimen recommendations for the patient, and other intelligent functions. The neural network model may be trained using a training dataset that includes subject records of subjects who have previously been treated for a condition and experienced an outcome (e.g., overcoming the condition, increasing the severity of the condition, decreasing the severity of the condition, etc.). Additionally, the executed code may be configured to cause the cloud server 135 to convert non-numeric values ​​of existing subject records into a numerical representation (e.g., the transformed representation) that can be processed by the trained neural network model. For example, code executed by cloud server 135 can be configured to receive as input each subject record of a set of subject records, and for each subject record, the code, when executed, can cause cloud server 135 to perform the operations described herein to convert each data element of each subject record into a transformed representation, such as a vector representation. Performing the intelligent function may include inputting at least a portion of the data records stored in data registry 140 into a trained machine learning or artificial intelligence model to generate an output for further analysis. In some embodiments, the output can be used to extract patterns in the data records or to predict values ​​or outcomes associated with data fields of the data records. Various embodiments of the intelligent functions performed by cloud server 135 are described below.

[0088] In some embodiments, the cloud server 135 is configured to enable a user device (e.g., operated by a physician) to access a cloud-based application to send a consultation broadcast to a set of destination devices. The consultation broadcast may be a request for support or assistance regarding the treatment of a subject associated with the subject record. The destination device may be a user device operated by another user associated with another entity (e.g., a physician at another medical center). If the destination device accepts the request for assistance associated with the consultation broadcast, the cloud-based application may generate an abridged representation of the subject record that omits or obscures certain data fields of the subject record. The abridged representation may comply with data privacy regulations, and thus the abridged representation of the subject record cannot be used to uniquely identify the subject associated by the subject record. The cloud-based application may send the abridged representation of the subject record to the destination device that accepted the request for assistance. The user operating the destination device may evaluate the abridged representation and communicate with the user device using a communication channel to discuss options for treating the subject. For example, the communication channel may be configured as a secure chat room that allows a user device (e.g., operated by a physician requesting a consultation) to securely communicate with a destination device (e.g., operated by another physician providing the consultation).

[0089] In some embodiments, the cloud server 135 is configured to provide a treatment plan definition interface to the user device. The treatment plan definition interface allows the user device to define a treatment plan for the condition. For example, the treatment plan may be a workflow for treating a subject having the condition. The workflow may include one or more criteria for defining a population of subjects as having the condition. The workflow may also include a particular type of treatment for the condition. The cloud server 135 receives and stores a treatment plan definition for a particular condition from each user device of the set of user devices. A cloud-based application may distribute a treatment plan for a given condition to the set of user devices. Two or more user devices of the set of user devices may be associated with different entities. Each of the two or more user devices may be provided with an option to integrate any portion or the entirety of the treatment plan into a customer rule set. The cloud server 135 may monitor whether the user device integrates the entire shared treatment plan or a portion of the treatment plan. The interaction between the user device and the shared treatment plan may be used to determine whether to update the treatment plan or rules created based on the treatment plan.

[0090] In some embodiments, the cloud server 135 allows a user operating a user device to access a cloud-based application to determine a suggested treatment for a subject having a condition. The user device loads an interface associated with the cloud-based application. The interface allows the user operating the user device to select a subject record associated with a subject being treated by the user. The cloud-based application can evaluate other subject records to identify previously treated subjects similar to the subject being treated by the user. The similarity between subjects may be determined, for example, using an array representation of the subject record. The array representation (e.g., a transformed representation such as a vector, an N-dimensional matrix, or any numerical representation of a non-numeric value) may be any numerical and / or categorical representation of values ​​of data fields of the subject record. For example, the array representation of the subject record may be a vector representation of the subject record in a domain space, such as in Euclidean space. In some examples, the cloud server 135 may be configured to convert the entire subject record into a numerical representation, such as a vector. For a given subject record, the cloud server 135 may evaluate each data element to determine the type of data encompassed or contained in that data element. The type of data can inform the cloud server 135 as to what process or technique to perform to convert the numeric or non-numeric values ​​of that data element into a numeric representation. By way of illustration, the cloud server 135 can convert the non-numeric values ​​of the data elements of a subject record (e.g., the text of a doctor's note) into a numeric representation (e.g., a vector). The conversion may include using natural language processing techniques such as Word2Vec or other text vectorization techniques to generate numeric values ​​that represent each word of the text. The generated numeric values ​​can serve as vectors that can be input into a trained neural network to perform intelligent analytics.As another illustrative example, for a data element that includes image frames of an image (e.g., MRI data) or video (e.g., ultrasound video data), each image or image frame may be converted to a numerical representation (e.g., a vector) using a trained autoencoder neural network that is trained to generate a latent space representation of the input image. The reduced representation (e.g., latent space representation) of the input image may serve as a numerical representation of the input image. This numerical representation may be input to a neural network or other machine learning model to perform intelligent analysis of the associated subject records. As yet another example, for a data element that includes a time-varying sequence of information (e.g., events that occurred or measurements taken from a subject over a period of time), the time-varying information may be represented as a numerical representation using several illustrative transformations. In some examples, a count of events may be used as a vector representing the time-varying information. For example, if measurements were taken on a subject four times in a year, the numerical representation may be "4". In other examples, the frequency or rate of an event occurring (e.g., once a week, once a month, once a year, etc.) may be used as a vector representing the time-varying information. In yet other examples, an average or combination of measurements associated with each event in the time-varying information can be used as a vector representing the time-varying information. The disclosure is not limited to these examples, and thus other numerical representations of the time-varying information can be used as vectors representing the numerical representations.

[0091] The AI ​​system 145 can be configured to collect datasets on a big data scale, convert the collected datasets into curated training data, run learning algorithms using the curated training data, and store the detected patterns, correlations, and / or relationships of the training data in one or more trained AI models. In some implementations, the AI ​​system 145 can be configured to perform certain predictive functions, such as predicting disease progression for a particular subject developing SMA, predicting candidate subject groups for inclusion in a new or existing clinical study, or predicting a situational treatment schedule specific to a particular subject. In some implementations, as described in more detail with respect to FIG. 8 and FIG. 11, the output of the AI ​​system 145 can predict disease progression for a particular subject diagnosed with SMA. In other implementations, as described in more detail with respect to FIG. 9 and FIG. 12, the output of the AI ​​system 145 can predict new groupings of subjects that may be suitable candidates for new clinical studies. In other embodiments, as described in more detail with respect to FIG. 10 and FIG. 13, the output of the AI ​​system 145 can predict treatment selection for a particular subject diagnosed with SMA.

[0092] In some examples, the multiple values ​​in the array representation correspond to a single field. For example, the value of a data element may be represented by multiple binary values ​​generated via one-hot encoding. As another example, each value of the multiple values ​​in a single data element of a subject record may be individually converted to a numeric representation as described above. The numeric representations representing each value of the multiple values ​​may be combined into a single numeric representation corresponding to the data element. Combining the multiple numeric representations may be performed using any vector combining technique, such as averaging vector magnitudes, adding vectors, or concatenating multiple vectors into a single vector. In some examples, the cloud-based application may generate an array representation for each subject record of a group of subject records. The similarity between two subject records may be represented by comparing the two array representations to determine the distance between them. Instead of comparing the numeric representation of an entire subject record to another numeric representation of another subject record, the subject records may also be compared along a dimension (e.g., data element). For example, comparing two subject records along a dimension may include comparing a numeric representation of a data element of a subject record to another numeric representation of a matching data element of another subject record. Further, the cloud-based application may be configured to identify subjects that are nearest neighbors to a subject record selected by a user device using the interface. The nearest neighbors may be determined by comparing the numerical representations of the various subject records to the numerical representation of the subject record of interest. The cloud-based application may identify actions previously performed on the subjects that are nearest neighbors. The cloud-based application may utilize the actions previously performed on the nearest neighbors on the interface.

[0093] In some embodiments, the cloud server 135 is configured to create a query that searches a database of previously treated subjects. The cloud server 135 may execute the query and search for subject records that satisfy the query constraints. However, in presenting the query results, the cloud-based application may present subject records in full only for subjects that were or are being treated by the user creating the query. The cloud-based application masks or otherwise obscures portions of subject records for subjects not treated by the user creating the query. Masking or obscuring portions of subject records included in the query results allows users to comply with data privacy regulations. In some embodiments, the query results (whether or not the query results are obscured) may be automatically evaluated for patterns or common attributes within the subject records.

[0094] In some embodiments, the cloud server 135 embeds the chatbot in the cloud-based application. The chatbot is configured to automatically communicate with a user device. The chatbot can communicate with the user device in a communication session in which messages are exchanged between the user device and the chatbot. The chatbot may be configured to select an answer to a question received from the user device. The chatbot can select an answer from a knowledge base accessible to the cloud-based application. When a user device sends a question to the chatbot and the chatbot does not have an existing answer stored in the knowledge base, a different representation of the question for which there is an existing answer is stored in the knowledge base. A user communicating with the chatbot can be prompted as to whether the answer provided by the chatbot is accurate or useful.

[0095] It will be appreciated that any machine learning or artificial intelligence algorithm may be executed to generate any of the trained machine learning models described herein. A variety of different types and techniques of artificial intelligence-based and machine learning models can be trained and then executed to generate one or more outputs that predict user outcomes for performing a protocol or function. Non-limiting examples of models include naive Bayes models, random forest or gradient boosting models, logistic regression models, deep learning neural networks, ensemble models, supervised learning models, unsupervised learning models, collaborative filtering models, and any other suitable machine learning or artificial intelligence models.

[0096] It will be appreciated that the cloud-based application can be configured to perform intelligent functions with respect to consulting with external physicians, determining a diagnosis, and recommending treatment for any disease, condition, area of ​​study, or disorder, including, but not limited to, COVID-19, cancers of the lung, breast, colorectal, prostate, stomach, liver, cervical (cervix), esophageal, bladder, kidney, pancreatic, endometrial, oral, thyroid, brain, ovarian, skin, and gallbladder, solid tumors such as sarcomas and carcinomas, cancers of the immune system including lymphomas (such as Hodgkin's lymphoma or non-Hodgkin's lymphoma), and oncology including cancers of the blood (hematological cancers) and cancers of the bone marrow such as leukemias (such as acute lymphocytic leukemia (ALL) and acute myeloid leukemia (AML)), lymphomas, and myelomas. Further disorders include blood disorders such as anemia, hemophilia, bleeding disorders such as blood clots, ophthalmologic disorders including diabetic retinopathy, glaucoma, and macular degeneration, neurological disorders including multiple sclerosis, Parkinson's disease, spinal muscular atrophy, Huntington's disease, amyotrophic lateral sclerosis (ALS), and Alzheimer's disease, autoimmune disorders including multiple sclerosis, diabetes, systemic lupus erythematosus, myasthenia gravis, inflammatory bowel disease (IBD), psoriasis, Guillain-Barre syndrome, chronic inflammatory demyelinating polyneuropathy (CIDP), Graves' disease, Hashimoto's thyroiditis, eczema, vasculitis, allergies, and asthma.

[0097] Other diseases and disorders include, but are not limited to, kidney disease, liver disease, heart disease, stroke, gastrointestinal disorders such as celiac disease, Crohn's disease, diverticular disease, irritable bowel syndrome (IBS), gastroesophageal reflux disease (GERD) and peptic ulcers, arthritis, sexually transmitted diseases, high blood pressure, bacterial and viral infections, parasitic infections, connective tissue diseases, celiac disease, osteoporosis, diabetes, lupus, diseases of the central and peripheral nervous system such as attention deficit / hyperactivity disorder (ADHD), catalepsy, encephalitis, epilepsy and seizures, peripheral neuropathy, meningitis, migraine headaches, myelopathy, autism, bipolar disorder, and depression.

[0098] IV.A. The cloud-based application enables user devices to broadcast consultation requests to other user devices and automatically condense subject records to comply with data privacy regulations. 2 is a flow chart illustrating a process 200 performed by a cloud-based application to deliver abbreviated subject records to user devices in association with a consultation broadcast requesting assistance regarding a subject's treatment. Process 200 may be performed by the cloud server 135 to enable user devices associated with different entities (e.g., hospitals) to collaborate or consult regarding a treatment for a subject while complying with data privacy regulations.

[0099] Process 200 begins at block 210, where cloud server 135 receives a set of attributes from a user device. Each attribute in the set of attributes can represent any characteristic of a subject (e.g., a patient). The set of attributes may be identified by a user using an interface provided by cloud server 135. For example, the set of attributes identifies demographic information of the subject and recent symptoms experienced by the subject. Non-limiting examples of demographic information include age, sex, ethnicity, state or city of residence, income range, education level, or any other suitable information. Non-limiting examples of recent symptoms include a subject currently or recently (e.g., at last clinic visit, intake, within 24 hours, within one week) experiencing a particular symptom (e.g., difficulty breathing, fever above a threshold temperature, blood pressure above a threshold blood pressure, etc.).

[0100] In block 220, the cloud server 135 generates a record for the subject. The record may be a data element including one or more data fields. The record indicates each of a set of attributes associated with the subject. The record may be stored in a central data store, such as the data registry 140 or any other cloud-based database. In block 230, the cloud server 135 receives a request submitted by a user using the interface. The request may be to initiate a consultation broadcast. For example, a user associated with an entity is a physician at a medical center treating the subject. The user may access a cloud-based application and operate a user device to broadcast a request to assist in treating the subject. The broadcast may be sent to a set of other user devices associated with a different entity.

[0101] In block 240, the cloud server 135 queries the central data store using one or more recent symptoms included in the set of attributes associated with the subject. The query result includes a set of other records. Each record in the set of other records is associated with another subject. In some examples, the cloud server 135 can query the central data store to identify other subject records similar to the subject record. Similarity may be determined by comparing the transformed representation of the entire subject record to the transformed representation of each other subject record. The comparison of the transformed representations can result in a distance (e.g., Euclidean distance) that represents the similarity between the two subject records. In other examples, similarity may be determined based on values ​​included in the data elements. For example, the subject record of the subject may include a subject data element that includes text that represents symptoms experienced by the subject. Each other subject record stored in the central data store may also include a data element that includes text that represents the associated subject's symptoms. The cloud server 135 can convert the text included in the subject data element into a numerical representation using techniques described above (e.g., trained convolutional neural networks, text vectorization techniques such as Word2Vec, etc.). The numerical representation of the text contained in the subject's data element may be compared to the numerical representation of the text contained in the matching data element of each of the other subject records. The result of the comparison (e.g., in a domain space such as Euclidean space) between the two numerical representations may indicate the degree to which the text contained in the subject's data element is similar to the text contained in the data element of the other subject record. In block 250, the cloud server 135 identifies a set of destination addresses (e.g., other user devices associated with different entities). Each destination address in the set of destination addresses is associated with a care provider for another subject associated with one or more other records in the set of other records identified in block 240. In block 260, the cloud server 135 generates a condensed representation of the record for the subject. The condensed representation of the record omits, obscures, or obscures at least a portion of the record.The condensed representation of the record cannot be used to uniquely identify the subject associated with the record and therefore can be exchanged between external systems without violating data privacy regulations. The cloud server 135 can perform any masking or obfuscation techniques to generate the condensed representation of the record.

[0102] At block 270, the cloud server 135 utilizes the abbreviated representation of the recording with a connection input component (e.g., a selectable link such as a hyperlink that allows a communication channel to be established) to each destination address of the set of destination addresses. The connection input component may be a selectable element presented at each destination address. Non-limiting examples of a connection input component include a button, a link, an input element, and other suitable selectable elements. At block 280, the cloud server 135 receives a communication from a destination device associated with the destination addresses. The communication includes an indication that a user operating the destination device has selected a connection input component associated with the abbreviated representation of the recording. At block 290, the cloud server 135 establishes a communication channel between the user device and the destination device for which the connection input component was selected. The communication channel enables a user operating the user device (e.g., a physician treating the subject) to exchange messages or other data (e.g., a video feed) with a destination device (e.g., a physician at another hospital that has agreed to assist in the treatment of the patient) associated with the destination address for which the connection input component was selected.

[0103] In some embodiments, the cloud server 135 is configured to automatically identify the location of the user device and the location of the destination device where the connection input component was selected. The cloud server 135 can also compare the locations to determine whether to generate a condensed representation of the recording. For example, in block 260, the cloud server 135 can generate a condensed representation of the recording because the cloud server 135 determined that each destination address of the set of destination addresses is not co-located with the user device that initiated the consultation broadcast. In this case, the cloud server 135 can automatically determine to generate a condensed representation of the recording to comply with data privacy regulations. As another example, if the set of destination addresses is associated with the same entity as the user device that initiated the consultation broadcast, the cloud server 135 can transmit the complete recording (e.g., without obfuscating any portion of the recording) to the destination device associated with the destination address while still complying with data privacy regulations.

[0104] In some embodiments, the cloud server 135 generates a plurality of other condensed record representations. Each of the plurality of other condensed record representations is associated with another subject. The cloud server 135 transmits the plurality of other condensed record representations to the user device and receives a communication from the user device identifying a selection of a subset of the plurality of other condensed record representations. Each of the sets of destination addresses is represented by one of the condensed record representations. For example, generating the condensed record representation includes determining a jurisdiction of the other subject associated with the condensed record representation, determining data privacy rules governing exchange of subject records within the jurisdiction, and generating the condensed record representation to comply with the data privacy rules. A first other condensed record representation of the plurality of other condensed record representations may include a particular type of data. A second other condensed record representation of the plurality of other condensed record representations may omit or obscure the particular type of data. For example, the particular type of data may be contact information, identifying information such as a name, a social security number, and other suitable information that can be used to uniquely identify the other subject.

[0105] In some implementations, the communication may be received at a central data store. The communication may be sent by a user device operated by a user and may include an identifier of a subject record of the subject of interest. When the communication is received at the central data store, it may cause the central data store to query a set of stored subject records to identify an incomplete subset of the set of subject records. Each subject record of the incomplete subset may be identified and included in the incomplete subset because the subject record was determined to be similar to the subject record of interest along at least one dimension. The similarity between two subject records along a dimension may represent a similarity with respect to data elements of the subject records, such as a similarity with respect to symptoms, diagnoses, treatments, or any other suitable data element. The one or more dimensions along which similarity or dissimilarity is determined may be automatically defined or user defined. Determining a similarity or dissimilarity between the subject record of interest and each subject record of the set of subject records stored in the central data store may include at least the following operations: retrieving the subject record of interest based on an identifier included in the communication, generating a transformed representation of the subject record of interest (or retrieving an existing transformed representation of the subject record of interest), and performing a clustering operation using the transformed representation of the subject record of interest and the transformed representation of each subject record of the set of subject records. The clustering operation may be performed on one or more dimensions (e.g., one or more features of the subject records). For example, the clustering operation may cluster the set of subject records stored in the central data store based on a data element that includes a value representing a symptom of the subject. The transformed representation of the subject record of interest may include a vector representation of a data element that includes a value representing a symptom of the subject. The vector representation of this data element of the subject record and the vector representation of the corresponding data element in each subject record of the set of subject records may be compared to define a cluster of subject records. Each cluster of subject records may define a group of one or more subject records sharing a common characteristic associated with a data element selected as a dimension of similarity.For each cluster of subject records, a Euclidean distance between the transformed representation of the target subject record and other transformed representations of the set of subject records may be calculated. A subject record may be determined to be similar to the target subject record, for example, when the Euclidean distance between the transformed representation of the subject record and the transformed representation of the target subject record is within a threshold.

[0106] IV.B. Updates to Shareable Treatment Plan Definitions Based on Aggregated User Integration FIG. 3 is a flow chart illustrating a process 300 for monitoring user integration of a treatment plan definition (e.g., a decision tree or treatment workflow) and automatically updating the treatment plan definition based on the results of the monitoring. The process 300 can be executed by the cloud server 135 to enable a user device to define a treatment plan for treating a population of subjects with a condition. The user device can distribute the treatment plan definition to user devices connected to an internal or external network. A user device receiving the treatment plan definition can determine whether to integrate the treatment plan definition into a custom rule base. The integration into the custom rule base can be monitored and used to automatically modify the treatment plan definition.

[0107] In block 310, the cloud server 135 stores the interface data that causes the treatment plan definition interface to be displayed when the user device loads the interface data. The treatment plan definition interface is provided to each user device of the set of user devices when the user device accesses the cloud server 135 to navigate to the treatment plan definition interface. In some embodiments, the treatment plan definition interface enables a user to define a treatment plan for treating a population of subjects having a condition (e.g., lymphoma).

[0108] At block 320, the cloud server 135 receives a set of communications. Each communication of the set of communications is received from a user device of the set of user devices and generated in response to an interaction between the user device and the treatment plan definition interface. In some embodiments, the communication includes one or more criteria for defining a population of subject records, for example. Each criterion may be represented by a type of variable. For example, the type of variable may be a value or a variable used as a condition of the criterion. The type of variable of a criterion of a rule may also be any value of a condition that constrains a population of subjects to an incomplete subgroup. For example, the type of variable of a rule defining a population of pregnant women is "IF "subject is pregnant". The criterion may be a filter condition for filtering a pool of subject records. For example, a criterion for defining a population of subject records associated with subjects likely to develop lymphoma may include the filter condition "abnormality in anaplastic lymphoma kinase (ALK)" AND "age 60 or older". The communication may also include a particular type of treatment for a condition. A particular type of treatment may be associated with a particular action (e.g., undergoing surgery) or refraining from a particular action (e.g., reducing salt intake) suggested to treat a symptom associated with a subject represented by the population of subject records.

[0109] In block 330, the cloud server 135 stores the set of rules in a central data store, such as the data registry 140 or any other centralized server in the cloud network 130. Each rule in the set of rules includes one or more criteria and a particular treatment type to be included in the communication from the user device. Illustratively, the rules represent a treatment workflow for treating lymphoma in a subject. The rules include the following criteria (e.g., a condition following an "IF" statement) and the following action (e.g., a particular treatment type defined or selected by the user following a "THEN" statement): "IF "lymph node biopsy indicates lymphoma cells are present" AND "blood test reveals the presence of lymphoma cells", THEN "treat with chemotherapy" AND "active surveillance". Additionally, each rule in the set of rules is stored in association with an identifier corresponding to the user device from which the communication was received.

[0110] In block 340, the cloud server 135 identifies a subset of the set of rules available across entities via the treatment plan definition interface. The subset of rules may include a subset of the set of rules associated with symptoms and distributed to external systems, such as other medical centers, for evaluation. For example, rules may be selected for inclusion in the subset of rules by evaluating characteristics of the rules or identifiers associated with the rules. The characteristics of the rules may include a code or flag that is stored or attached to the stored rules. The code or flag indicates that the rule is generally available to external systems (e.g., utilized by the entity).

[0111] In block 350, for each rule of the subset of rules identified in block 340, the cloud server 135 monitors interactions with the rule. The interactions may include an external entity (e.g., outside the entity associated with the user that defined the treatment plan associated with the rule) integrating the rule into a custom rule base. For example, a user device associated with the external entity (e.g., another hospital) evaluates the rule utilized by the external entity. The evaluation includes determining whether the rule is suitable for integration into a rule set defined by the external entity. The rule may be suitable when a user device associated with the external entity indicates that a treatment workflow defined using the rule is suitable for treating the condition corresponding to the rule. Continuing with the above example, a rule for treating lymphoma may be utilized by an external medical center. A user associated with the external medical center determines that the rule for treating lymphoma is suitable for integration into a rule set defined by the external medical center. In this manner, after the rule is integrated into the custom rule base defined by the external medical center, other users associated with the external medical center can execute the integrated rule by selecting the integrated rule from the custom rule base. Additionally, the cloud server 135 monitors the integration of utilized rules by detecting a signal that is generated, or caused to be generated, when the treatment plan definition interface receives input from a user device associated with an external entity corresponding to the integration of a rule into the custom rule base.

[0112] As another example, a user device associated with an external entity integrates a modified version of a rule specified by an interaction into a custom rule base using a treatment plan definition. The modified version of a rule specified by an interaction is a portion of the rule selected for integration into the custom rule base. Selecting a portion of a rule for integration includes selecting fewer criteria than all criteria included in the rule for integration into the custom rule base. Continuing with the above example, a user device associated with an external entity selects the criterion "IF 'lymph node biopsy indicates lymphoma cells are present'" for integration into the custom rule base, but the user device does not select the criterion "blood test reveals the presence of lymphoma cells" for integration into the custom rule base. Thus, the modified version of a rule specified by an interaction integrated into the custom rule base is "IF 'lymph node biopsy indicates lymphoma cells are present' THEN 'treatment with chemotherapy' AND 'active surveillance'". The criterion "blood test reveals the presence of lymphoma cells" is removed from the rule to create a modified version of a rule specified by an interaction integrated into the custom rule base.

[0113] In block 360, the cloud server 135 may detect that a modified version of the rule specified by the interaction has been integrated into the custom rule base defined by the external entity. Upon detection, the cloud server 135 may update the rule stored in the central data store of the cloud network 130. The rule may be updated based on the monitored interaction. The term "based on" in this example corresponds to "after evaluating" the monitored interaction or "using the results of the evaluation." For example, the cloud server 135 detects that a user device associated with the external entity has integrated a modified version of the rule specified by the interaction. In response to detecting the modified version of the rule specified by the interaction, the cloud server 135 may update the rule stored in the central data store from the existing rule to the modified version of the rule specified by the interaction.

[0114] In some embodiments, the cloud server 135 updates the rules by generating an updated version to be utilized across external entities. Another original version may remain unupdated and is utilized to a user associated with a user device from which one or more communications that identified the criteria and a particular type of action were received. For example, the cloud server 135 updates a rule stored in a central data store, but the cloud server 135 does not update another rule in the set of rules stored in the central data store.

[0115] In some embodiments, the cloud server 135 can update the rules when an update condition is met. The update condition may be a threshold. For example, the threshold may be a number or percentage of external entities that have integrated a modified version of the rule into their custom rule base. As another example, the update condition may be determined using the output of a trained machine learning model. For example, the cloud server 135 can input detection signals received from an external entity into a multi-armed bandit model that automatically determines whether and / or when to utilize a rule and / or whether and when to utilize an updated version of the rule. For example, and by way of non-limiting example only, the rules may be defined as executable code such that, when executed, the rules automatically query the central data store to identify a subset of the set of subject records for further analysis. Additionally, the rules may include one or more treatment protocols for treating subjects associated with the identified subset of subject records. The rules may be defined as workflows that define a subset of the set of subject records and for treating subjects associated with the subset of subject records. For example, a rule may include one or more criteria for filtering subject records from the set of subject records and performing a particular treatment protocol on subjects associated with the remaining subject records (e.g., the subject records remaining after filtering has been performed on the set of subject records). Although the rules are defined by a user of the first entity, the rules may be accepted (e.g., integrated into the rule base of the second entity), modified, or completely rejected by an external user (e.g., a doctor working at another hospital) of a second entity (e.g., the first entity and the second entity are two different medical facilities). In some examples, a feedback signal may be sent to the cloud server 135 each time the external user of the second entity accepts a rule and thus fully integrates the rule into its code base.In another example, a feedback signal may be sent to the cloud server 135 whenever a user of the second entity modifies a rule. In another example, a feedback signal may be sent to the cloud server 135 whenever a user of the second entity completely rejects a rule. In each of the above examples, the feedback signal may include the rule (e.g., a rule identifier) ​​and data indicating whether the rule was accepted, modified, or rejected. The multi-armed bandit model (executable by the cloud server 135) may be configured to intelligently select one of the original rule, the modified rule, or a completely different rule for broadcast to external users of other entities. The selection of the original rule, the modified rule, or the different rule may be based at least in part on the configuration of the multi-armed bandit. In some examples, the multi-armed bandit may be configured by an epsilon-ready search technique. In the epsilon-ready search technique, the multi-armed bandit model may select the original rule for broadcast to external users of other entities with a probability of "1-ε", where ε represents the probability of searching for a new or modified rule. In this way, the multi-armed bandit model can select a modified version of the original rule or an entirely new rule with a defined probability of ε. The multi-armed bandit model can change ε ​​based on feedback signals received from other entities. For example, if the feedback signals indicate that a rule has been modified in a particular manner by different external users more than a threshold number of times, the multi-armed bandit model can learn to select the modified rule in a particular manner for broadcast to the external users instead of broadcasting the original rule.

[0116] In some embodiments, the cloud server 135 identifies multiple rules in the set of rules that include criteria corresponding to the same variable type and identify the same or similar types of treatments. The variable type may be a value or a variable used as a condition of the criteria. The variable type of the criteria of the rule may also be any value of the condition that constrains the population of subjects into subgroups. For example, the variable type of a rule that defines a population of pregnant women is "IF "subject is pregnant". The cloud server 135 identifies a new rule that is a contracted representation of multiple rules when the new rule is sent globally to servers operated by other entities.

[0117] In some embodiments, the cloud server 135 provides another interface configured to receive the set of subject attributes. For example, a user operating a user device accesses the other interface and uses the other interface to select a subject record that includes the set of attributes. The selection of the subject record can cause the cloud server 135 to receive the set of subject attributes. The cloud server 135 identifies (e.g., determines) a particular rule whose criteria are met based on the set of subject attributes. For example, the cloud server 135 evaluates the set of attributes of the subject record against the criteria of the rule stored in the central data store. For example, if the set of attributes includes a data field that includes a value of "is pregnant" and if the rule includes a single criterion of "IF "subject is pregnant", the cloud server 135 identifies the rule. The cloud server 135 updates the other interface to present the particular rule and each particular type of treatment associated with the particular rule.

[0118] In some embodiments, the criteria of the rule is the type of variables associated with a particular demographic variable and / or a particular symptom-type variable.Non-limiting examples of demographic variables include any item of information that characterizes the demographic data of a subject, such as age, sex, ethnicity, race, income level, education level, location, and other suitable items of demographic information.Non-limiting examples of symptom-type variables indicate whether a subject currently or recently (e.g., at last visit, at intake, within 24 hours, within 1 week) experiences a particular symptom (e.g., difficulty in breathing, loss of consciousness, fever above threshold temperature, blood pressure above threshold blood pressure, etc.).

[0119] In some embodiments, the cloud server 135 monitors data in a registry of subject records, such as subject records stored in the data registry 140. For each rule of the subset of rules (identified in block 340), the cloud server 135 monitors data in the registry of subject records. The cloud server 135 identifies a set of subjects for which the criteria of the rule are met and a particular treatment has been previously prescribed for the subject. For each of the set of subjects, the cloud server 135 identifies a reported status of the subject, indicated from or using an assessment or examination. For example, a reported status is any information that characterizes an aspect of the subject's status, such as whether the subject has been discharged from the hospital, whether the subject is alive, a measurement of the subject's blood pressure, the number of times the subject has woken up during a sleep stage, and other suitable status. The cloud server 135 determines an estimated responsiveness metric of the set of subjects to a particular treatment based on the reported status. For example, if a particular action of the rule is to prescribe a medication, the estimated responsiveness metric is a representation of the extent to which the medication has addressed the symptoms or symptoms experienced by the subject. As a non-limiting example, the estimated responsiveness metric for the set of subjects may be an average, weighted average, or any sum of the scores assigned to each subject in the set of subjects. The score may represent or measure the effectiveness of the subject's responsiveness to the treatment. In some examples, the cloud server 135 may generate a score representing the effectiveness of the subject's responsiveness to the treatment by using clustering techniques. For example, and by way of non-limiting example only, the set of subject records may represent subjects who have previously undergone a particular treatment protocol to treat a condition. Each subject record in the set of subject records may be labeled (e.g., by a user) as having one of a positive responsiveness to the particular treatment protocol, a neutral responsiveness to the particular treatment protocol, or a negative responsiveness to the particular treatment protocol.The set of subject records may then be divided into three subsets (e.g., clusters), where a first subset of subject records may correspond to subjects with a positive responsiveness to the particular treatment protocol, a second subset of subject records may correspond to subjects with a neutral responsiveness to the particular treatment protocol, and a third subset of subject records may correspond to subjects with a neutral responsiveness to the particular treatment protocol. The cloud server 135 may convert each subject record of the first subset of subject records to a transformed representation according to the implementations described above. The cloud server 135 may also convert each subject record of the second subset of subject records to a transformed representation using the techniques described above. Finally, the cloud server 135 may convert each subject record of the third subset of subject records to a transformed representation using the techniques described above. In some implementations, determining the predicted responsiveness of the new subject to the particular treatment protocol may include converting the new subject record of the new subject to a new transformed representation. The new transformed representation may be compared in domain space (e.g., Euclidean space) with the transformed representations of each cluster or subset of subject records. If the new transformed representation is closest to the centroid of the transformed representations associated with the first subset, the new subject is predicted to have a positive responsiveness to the particular treatment. If the new transformed representation is closest to the centroid of the transformed representations of the second subset, the new subject is predicted to have a neutral responsiveness to the particular treatment. Finally, if the new transformed representation is closest to the centroid of the transformed representations of the third subset, the new subject is predicted to have a negative responsiveness to the particular treatment protocol. The centroid may be a multidimensional average of the transformed representations associated with the subsets. The cloud server 135 may cause the estimated responsiveness metrics of the subset of the set of rules and the set of subjects to be displayed or otherwise presented in the treatment plan definition interface.

[0120] IV.C. Presenting Treatment Recommendations with Associated Efficacy Using Treatments Prescribed to Similar Subjects 4 is a flow chart illustrating a process 400 for recommending treatments for a subject. The process 400 may be performed by the cloud server 135 to display the recommended treatments for the subject and the effectiveness of each recommended treatment on a user device associated with the medical entity. The recommended treatments may be identified using results of evaluating the effectiveness of treatments previously prescribed for similar subjects.

[0121] At block 410, the cloud server 135 receives input corresponding to a subject record characterizing a subject's behavior. The input is received from a user device associated with the entity. Additionally, the input is received in response to the user device selecting or otherwise identifying the subject record using an interface associated with an instance of a platform configured to manage a registry of subject records. The user device can access the interface by loading interface data stored on a web server (not shown) connected within the cloud network 130. The web server may be included in or executed on the cloud server 135.

[0122] At block 420, the cloud server 135 extracts a set of subject attributes from the subject record received at block 410. The subject attributes characterize the subject's behavior. Non-limiting examples of subject attributes include any information found in an electronic health record, any demographic information, age, sex, ethnicity, recent or past symptoms, symptoms, symptom severity, and any other suitable information that characterizes the subject.

[0123] At block 430, the cloud server 135 generates an array representation of the subject record using the set of subject attributes. For example, the array representation is a vector representation of the values ​​contained in the subject record. The vector representation may be a vector in a domain space, such as Euclidean space. However, the array representation may be any numerical representation of the values ​​of the data fields of the subject record. In some embodiments, the cloud server 135 may perform a feature decomposition technique, such as singular value decomposition (SVD), to generate values ​​that represent the set of subject attributes of the array representation of the subject record.

[0124] At block 440, the cloud server 135 accesses a set of other sequence representations characterizing the plurality of other subjects. An sequence representation in the set of other sequence representations may be a vector representation of a subject record characterizing another subject (e.g., one of the plurality of other subjects).

[0125] In block 450, the cloud server 135 determines a similarity score representing the similarity between the array representation representing the subject and each of the array representations of the other subjects. For example, the similarity score is calculated using a function of the distance (in domain space) between the array representation representing the subject and the array representations representing the other subjects. For example, by way of non-limiting example only, the similarity score may be calculated using a range from "0" to "1", with "0" representing a distance above a defined threshold and "1" representing that the array representations have no distance between them. For example, by way of non-limiting example only, the similarity score may be based on the Euclidean distance between the two array representations (e.g., vectors).

[0126] At block 460, the cloud server 135 identifies a first subset of the plurality of other subjects. A subject may be included in the first subset when the similarity score associated with the subject falls within a predetermined absolute or relative range. Similarly, at block 470, the cloud server identifies a second subset of the plurality of other subjects. However, a subject may be included in the second subset when the similarity score of the subject falls within another predetermined range.

[0127] At block 480, the cloud server 135 retrieves record data for each subject in the first and second subsets of the plurality of other subjects. The record data includes attributes included in the subject record that characterize the subject. For example, the subject record data identifies a treatment the subject received and the subject's responsiveness to the treatment. The responsiveness to the treatment may be represented by text (e.g., "the subject responded positively to the treatment") or a score indicating the degree to which the subject responded positively or negatively to the treatment (e.g., a score from "0" to "1" with "0" indicating a negative responsiveness and "1" indicating a positive responsiveness). In some examples, the treatment responsiveness may indicate the degree to which the subject responded positively to a treatment previously administered to the subject. For example, the treatment responsiveness may be numeric (e.g., a score from "0" to "10") or non-numeric (e.g., a word assigned to represent a responsiveness such as "positive", "neutral", or "negative"). In some examples, the treatment responsiveness for previously treated subjects may be user-defined. In other examples, treatment responsiveness may be determined automatically based on the results of tests or measurements taken from a user. For example, treatment responsiveness may be determined automatically based on values ​​contained in a blood test performed on the subject.

[0128] At block 490, the cloud server 135 generates an output to be presented in an interface on the user device. The output may, for example, indicate one or more treatment recommendations for the subject. The one or more treatment recommendations may be determined, for example, based on treatments received by other subjects in the first and second subsets, treatment responsiveness of the subjects in the first and second subsets, and differences between subject attributes of the subjects in the second subset and subject attributes of the subjects.

[0129] In some embodiments, the cloud server 135 determines that the subject and one of the subjects from the first or second subset are being or have been treated by the same medical entity. The cloud server 135 determines that the subject and another subject from the first or second subset are being or have been treated by a different medical entity. The cloud server 135 can utilize different obfuscated versions of the subject's records via an interface. The cloud-based application can automatically provide the entity with different obfuscated versions of the records based on various constraints imposed on data sharing by data privacy regulations in different jurisdictions. In some embodiments, the cloud server 135 identifies the first and second subsets of subject records by performing a clustering operation on the transformed representation of the set of subject records.

[0130] IV.D. Automatic Disambiguation of Query Results from External Entities 5 is a flow chart illustrating a process 500 for obfuscating query results to comply with data privacy regulations. The process 500 may be executed by the cloud server 135 as an execution rule to ensure that data sharing of subject records with external entities complies with data privacy regulations. A cloud-based application may enable a user device to query the data registry 140 for subject records that satisfy the query constraints. However, the query results may include data records originating from the external entity. In this manner, the process 500 enables the cloud server 135 to provide additional information about a procedure from the external entity to the user device while still complying with data privacy regulations.

[0131] At block 510, the cloud server 135 receives a query from a user device associated with a first entity. For example, the first entity is a medical center associated with a first set of subject records. The query may include a set of symptoms associated with a medical condition, or any other information that constrains the query search of the data registry 140.

[0132] In block 520, the cloud server 135 queries the database using the query received from the user device. In block 530, the cloud server 135 generates a data set of query results corresponding to a set of symptoms and associated with a medical condition. For example, the user device submits a query for subject records of a subject diagnosed with lymphoma. The query results include at least one subject record from a first set of subject records (originating from or created at a first entity) and at least one subject record from a second set of subject records associated with a second entity (e.g., a medical center different from the first entity). Each of the subject records from the first set of subject records and the subject record from the second set of subject records may include a set of subject attributes. The subject attributes may characterize any aspect of the subject.

[0133] In block 540, the cloud server 135 presents (e.g., utilizes or otherwise makes available) to the user device a complete set of subject attributes for the subject records included in the first set of subject records as these records originate from the first entity. Presenting the subject records completely includes making available to the user device the set of attributes included in the subject records for evaluation or interaction using an interface. In block 550, the cloud server 135 also or alternatively utilizes to the user device an incomplete subset of the set of subject attributes for each subject record included in the second set of subject records. Providing the incomplete subset of the set of subject attributes provides anonymity to the subject as the incomplete subset of the subject attributes cannot be used to uniquely identify the subject. For example, providing the incomplete subset may include available 4 of the 10 subject attributes to anonymize the subject associated with the 10 subject attributes. In some embodiments, in block 550, the cloud server 135 utilizes an obscured set of subject attributes for each subject record included in the second subject. Obscuring the set of attributes includes reducing the granularity of the information provided. For example, instead of utilizing a subject attribute of the subject's address, the obscured attribute may be a zip code or the state the subject lives in. Regardless of whether an incomplete subset or an obscured subset is utilized, the cloud server 135 anonymizes the subjects associated with the subject records.

[0134] IV.E. Chatbot Integration with Self-Learning Knowledge Base 6 is a flow chart illustrating a process 600 for communicating with a user using a bot script, such as a chatbot. The process 600 may be performed by the cloud server 135 to automatically link new questions provided by a user to existing questions in a knowledge base to provide answers to the new questions. The chatbot may be configured to provide answers to questions associated with symptoms.

[0135] In block 605, the cloud server 135 defines a knowledge base that includes a set of answers. The knowledge base may be a data structure stored in memory. The data structure stores text representing a set of answers to defined questions. Each answer may be selectable by the chatbot in response to a question received from the user device during a communication session. The knowledge base may be automatically defined (e.g., by retrieving text from a data source and parsing through the text using natural language processing techniques) or may be user-defined (e.g., by a researcher or physician).

[0136] In block 610, the cloud server 135 receives a communication from a particular user device. The communication corresponds to a request to initiate a communication session with a particular chatbot. For example, a doctor or a subject may operate the user device to communicate with the chatbot within the chat session. The cloud server 135 (or a module stored within the cloud server 135) may manage or establish the communication session between the user device and the chatbot. In block 615, the cloud server 135 receives a particular question from the particular user device during the communication session. The question may be a string of text that is processed using natural language processing techniques.

[0137] In block 620, the cloud server 135 queries the knowledge base using at least some words extracted from the particular question. The words may be extracted from the string of text representing the particular question using natural language processing techniques. In block 625, the cloud server 135 determines that the knowledge base does not contain a representation of the particular question. In this case, the received question may be newly posted to the chatbot. In block 630, the cloud server 135 identifies another question representation from the knowledge base. The cloud server 135 may identify the another question representation by comparing the question received from the user device with other question representations stored in the knowledge base. For example, if a similarity is determined based on an analysis of the question representation using natural language processing techniques, the cloud server 135 identifies the other question representation.

[0138] At block 635, the cloud server 135 searches for the answer among other question representations and sets of answers associated in the knowledge base. At block 640, the answer searched at block 635 is sent to the particular user device as an answer to the received question, even if the knowledge base did not contain the representation of the received question. At block 645, the cloud server 135 receives an indication from the particular user device. For example, the indication may be received in response to the user device indicating that the answer provided by the chatbot responded to the particular question.

[0139] In block 650, the cloud server 135 updates the knowledge base to include the representation of the particular question or a different representation of the particular question. For example, storing the representation of the question includes storing keywords included in the question in a data structure. The cloud server 135 can also associate the same or different representation of the particular question with more answers sent to the particular user device.

[0140] In some embodiments, the cloud server 135 accesses a subject record associated with a particular user device. The cloud server 135 determines multiple answers to a particular question. The cloud server 135 then selects an answer from the set of answers. However, the selection of the answer is based at least in part on one or more values ​​included in the subject record associated with the particular user device. For example, the values ​​included in the subject record can represent symptoms recently experienced by the subject. The chatbot may be configured to select an answer that depends on the symptoms recently experienced by the subject. In some examples, the cloud server 135 can access a learning-to-rank machine learning model trained to predict the order of each answer in the set of answers. The learning-to-rank machine learning model may be trained using a training set of answers. Each answer in the training set of answers may be labeled with one or more symptoms and a relevance score for the symptoms. The relevance score may represent the relevance of the associated answer to a given symptom of the one or more symptoms. The relevance score may be user-defined or may be automatically determined based on certain factors, such as the frequency of words (e.g., words for symptoms) in the training answers. The set of training answers may be different from the set of answers used when the chatbot is operating in a production environment. The learning-to-rank machine learning model can learn how to order the set of answers (used in the production environment) in terms of relevance to the symptoms (detected from the subject profile) based on the patterns (e.g., patterns between the training set of answers labeled per symptom of one or more symptoms and the associated relevance scores) learned by the learning-to-rank model. The chatbot can select an answer from the set of answers used in the production environment based on the predicted ordering of the set of answers. In some examples, each answer in the set of answers may be associated with a tag or code that indicates one or more symptoms associated with the answer. The cloud server 135 can compare a value representing a symptom recently experienced by the subject to the tag or code associated with each answer.

[0141] A network environment configured to facilitate intelligent treatment selection for treating subjects diagnosed with V.SMA FIG. 7 is a block diagram illustrating an example of a network environment for deploying a trained artificial intelligence model to facilitate subject-specific identification of a treatment and treatment schedule for treating a subject suffering from SMA, according to some aspects of the present disclosure. The network environment 700 can include a user device 110 and an AI system 702. The user device 110 can interact with the AI ​​system 702 using a network 736 (e.g., any public or private network), which facilitates the exchange of communications between the user device 110 and the AI ​​system 702. The AI ​​system 702 may be another implementation of the AI ​​system 145 described with respect to FIG. 1. The user device 110 can be operated by a user, such as a doctor or other medical professional treating a subject diagnosed with SMA. The user device 110 can send a request to the AI ​​system 702 using an application programming interface (API) 704 to trigger a particular function (e.g., a cloud-based service). Although FIG. 7 shows a single user device 110, it will be appreciated that any number of user devices or other computing devices, such as cloud-based servers, can interact with the AI ​​system 702.

[0142] The AI ​​system 702 can be configured to perform certain predictive functions, such as, for example, predicting suitable candidates for clinical studies, predicting disease progression for a particular subject developing SMA, or predicting a situation treatment schedule specific to a particular subject. The AI ​​system 702 can perform the predictive functions, for example, using an AI model execution system 710. Some data structures (e.g., databases) for storing data can facilitate the predictive functions that the AI ​​system 702 can perform. In some implementations, the data structures can store training data 716, validation data 718, test data 720, subject records from a data registry 722, an AI model 724, treatments 726, treatment schedules 728, clinical studies 730, and subject group identifiers 732. The various components of the AI ​​system 702 can communicate with each other using a communication network 734.

[0143] The AI ​​model training system 708 can facilitate training of the AI ​​model using the training data 716. For example, the AI ​​model training system 708 can execute code (e.g., executed by a processor, such as a physical or virtual CPU of a cloud-based server) that causes the training data 716 to be input into a learning algorithm. The learning algorithm can be executed to detect patterns or correlations between data points included in the training data 716. The detected patterns or correlations can be stored as an AI model, and the AI ​​model is trained to generate outputs that predict outcomes based on the stored patterns or correlations in response to receiving inputs (e.g., of new, previously unseen input data, such as subject records for subjects not included in the training data 716).

[0144] In some implementations, the output of the trained AI model can predict disease progression for a particular subject diagnosed with SMA, as described in more detail with respect to Figures 8 and 11. In other implementations, the output of the trained AI model can predict new or previously unstudied subjects for investigation using new clinical studies and suitable candidate subjects for new clinical studies, as described in more detail with respect to Figures 9 and 12. In other embodiments, the output of the trained AI model can predict treatment selection for a particular subject developing SMA, as described in more detail with respect to Figures 10 and 13.

[0145] The learning algorithms performed by the AI ​​system 702 may include any supervised, unsupervised, semi-supervised, reinforcement, and / or ensemble learning algorithms. Non-limiting examples of learning algorithms that may be performed by the AI ​​system 702 are included in Table 1 below. Selection of a learning algorithm by the AI ​​system 702 to train an AI model can be based, for example, on the type and size of at least a portion of the training data 716, as well as a target prediction outcome targeted to a prediction function that the AI ​​system 702 is capable of performing. The learning algorithms provided in Table 1 can be used for any of the methods described herein. [Table 1]

[0146] Additionally, during the process of training various AI models, the AI ​​model training system 708 can interact with training data 716, validation data 718, and testing data 720. The training data 716 is a data set that is input to the learning algorithm. The learning algorithm detects patterns, correlations, or relationships between data points in the training data 716. However, the patterns, correlations, or relationships (e.g., parameters) detected by the learning algorithm may cause the training data 716 to be overfitted. Overfitting occurs when the analysis performed by the learning algorithm (e.g., that generated the patterns, correlations, or relationships) corresponds exactly or substantially exactly to the training data 716. In this case, the analysis performed by the learning algorithm may not accurately serve as a basis for predicting new, previously unseen input data. Thus, the validation data 718 is a different data set than the training data 716 and is used to correct the patterns, correlations, or relationships to prevent overfitting of the training data 716. If multiple learning algorithms are run on the training data 716, the validation data 718 can be used to identify the learning algorithm that has the best performance on new input data (e.g., input data not included in the training data 716). The validation data 718 can be used to generate error functions that can be evaluated to determine the performance of each learning algorithm on the new input data. For example, patterns, correlations, or relationships detected in the training data 716 by each of the various learning algorithms can be stored in various AI models. The error function of each AI model on the new input data can be evaluated using the validation data 718. The AI ​​model with the smallest error function can be selected. Finally, the test data 720 is another data set independent of each of the training data 716 and the validation data 718. The test data 720 can be input into a selected AI model to test the overall performance of the selected AI model.

[0147] In some implementations, the training data 716, validation data 718, and testing data 720 may be segments across a single larger dataset. For example, the dataset may be segmented into three data subsets. The training data 716 may be one of the three data subsets, the validation data 718 may be another one of the three data subsets, and the testing data 720 may be the last of the three data subsets. In some implementations, the dataset segmented into three or more subsets may include any data or data type. Non-limiting examples of data or data types that may be included in the datasets from which the training data 716, validation data 718, and / or testing data 720 are generated include radiology image data, MRI data, genomic profile data, clinical data (e.g., measurements, treatments, treatment responses, diagnoses, severity, medical history), subject-generated data (e.g., notes entered by a subject with SMA), physician or medical professional generated data (e.g., physician's notes), voice data representing telephone calls between a patient and a physician or other medical professional, administrative data, claims data, health surveys (e.g., Health Risk Assessment (HRS) surveys), third party or vendor information (e.g., out-of-network lab results), public databases related to the subject (e.g., medical journals related to the subject's symptoms), subject demographics, immunizations, radiology reports, pathology reports, utilization information, metadata representing biological samples, social data (e.g., education level, employment status), community specifications, etc. In some examples, at least a portion of the subject records may be initially identified via communications from a device operated by the subject (e.g., received at a caregiver's device and / or a remote server). In some implementations, at least some of the characteristics of the subject record include or are based on one or more photographs (e.g., collected on a device of the subject). In some examples, at least a portion of the subject-specific data was initially identified via and / or received from an electronic medical record corresponding to the subject.

[0148] The AI ​​model execution system 710 may be implemented using executable code that, when executed by a processor (e.g., a physical or virtual CPU in a cloud-based network such as cloud network 130), executes an instance of a particular trained AI model to generate output. The output may predict a particular outcome for the SMA attributable to the AI ​​model.

[0149] For example, and by way of non-limiting example only, the AI ​​model execution system 710 receives a request from the query resolver 706 (originating from the user device 110 operated by a user such as a physician). The request is for predicting disease progression of a particular subject suffering from SMA. The request includes at least a portion of a subject record (or an identifier of the subject record that allows retrieval of the subject record by another component) that characterizes the particular subject. The AI ​​model execution system 710 evaluates the request and selects a trained word-to-vector model (stored in the AI ​​model data store 724) configured to generate a prediction of the subject's disease progression. The AI ​​model execution system 710 retrieves or accesses the word-to-vector model from the AI ​​model data store 724 and then passes input data (e.g., a numerical representation of the particular subject's current condition) to the retrieved AI model. The AI ​​model execution system 710 generates an output (e.g., one or more values ​​in an array, etc.) that can be used to determine the particular subject's disease progression. The prediction functionality described in this example is further described with respect to FIG. 8 and FIG. 11.

[0150] As another illustrative, and merely non-limiting, example, the user device 110 sends a request to the AI ​​system 702 to generate a prediction of which groups of subjects are suitable candidates for enrollment in a new clinical study. The AI ​​system 702 retrieves or accesses the trained feature selection model and the automatic grouping model. The AI ​​system 702 then inputs the set of numerical representations of the subject records into the feature selection model and subsequently into the automatic grouping model to generate a prediction of the groups of subjects that should be suitable candidates for the new clinical study (e.g., the new clinical study stored in the clinical study data store 730). Identifiers of the groups of subjects predicted to be suitable candidates for enrollment in the new clinical study may be stored in the subject group data store 732. In some examples, the AI ​​system 702 can automatically identify the groups of subjects that should be suitable candidates for the clinical study without having to receive a request from the user device 110. In other examples, the AI ​​system 702 can automatically identify the groups of subjects based on common characteristics of the groups of subject records and suggest a new clinical study associated with the common characteristics if one does not already exist. The prediction functionality described in this example is further described with respect to FIG. 9 and FIG.

[0151] As yet another illustrative example, and by way of non-limiting example only, the user device 110 sends a request to the AI ​​system 702 to predict treatment selection and treatment schedule for a particular subject. The AI ​​system 702 retrieves or accesses a trained reinforcement model configured to select an optimal treatment workflow including multi-stage treatments and schedules for multi-stage treatments. The AI ​​system 702 inputs vectors representing the characteristics of the particular subject into the trained reinforcement model to generate an output representing a particular multi-stage treatment and a schedule for performing the multi-stage treatment (among multiple single-stage or multi-stage treatments stored in the treatment data store 726 and the treatment schedule data store 728). The prediction functionality described in this example is further described with respect to FIG. 10 and FIG. 13.

[0152] Certain AI models may exhibit a technical problem of memorizing portions of the training data 716 during the training process. Memorizing portions of the training data 716 may occur when a trained AI model outputs data elements contained in the training data 716 as is in response to receiving input data. Data leakage refers to an AI model outputting data elements from the training data as is in response to the input of new, previously unseen data. In some cases, an AI model memorizes the training data when the AI ​​model is overfitted to the training data. An overfitted AI model memorizes noise contained in the training data (e.g., memorizes data elements from the training data that are not relevant to the task of learning). Thus, an AI model does not generalize predictions on new, previously unseen input data when the AI ​​model exhibits data leakage.

[0153] If the training data includes sensitive or personal data about the subject, a data leak may violate privacy regulations. For example, and by way of a non-limiting example only, the training data 716 includes a subject record that includes a value representing that the subject (as characterized by the subject record) has a genetic mutation associated with early onset of Alzheimer's disease. The value representing the presence of the Alzheimer's genetic mutation is sensitive or personal data. Thus, various privacy statutes and regulations prohibit unauthorized disclosure of a subject's sensitive or personal data (e.g., the Health Insurance Portability and Accountability Act (HIPAA)). However, a technical challenge arises in that if the trained AI model is overfitted to the training data 716, the trained AI model may leak (e.g., unintentionally disclose externally or to an unauthorized user) a value representing that the subject has an Alzheimer's genetic mutation. In some scenarios, a privacy breach may occur if an adversarial user device (e.g., operated by a user who is intentionally attempting to extract sensitive information from the AI ​​model) is able to send inputs to the trained AI model and receive the corresponding outputs generated by the AI ​​model. For example, if an adversarial user device accesses a trained AI model using the public API, the adversarial user device can send inputs to the trained AI model and receive outputs generated by the trained AI model. The adversarial user device can then evaluate various outputs received from the trained AI model to infer confidential or private data about the training data used to train the AI ​​model. Non-limiting examples of confidential or private data that may be inferred include the presence of a particular genetic mutation in a particular subject, the presence or absence of a subject record in the training data, the presence or absence of a particular subject in a particular clinical study, a correlation between a phenotype exhibited by a particular subject and a particular subject's genetic predisposition to develop a particular disease, such as SMA, characteristics of a particular subject's genetic profile, and any other confidential or private data.

[0154] To solve the technical problems related to data leakage discussed above, certain aspects and features of the present disclosure relate to configuring a data leakage detector 712 to detect and prevent data leakage when the AI ​​model execution system 710 executes any of the trained AI models stored in the AI ​​model data store 724. In some implementations, the data leakage detector 712 can execute certain data leakage prevention protocols on the training data 716, the validation data 718, the test data 720, and / or the AI ​​models 724. Executing the data leakage prevention protocols on the training data 716, the validation data 718, the test data 720, and / or the AI ​​models 724 can suppress or prevent leakage of sensitive data by the trained AI models. Non-limiting examples of data leakage prevention protocols that may be performed on the data include encryption of sensitive or personal data contained in subject records, data sanitization, data normalization, robust statistics, adversarial training, differential privacy, federated learning, homomorphic encryption, and other suitable techniques to reduce or prevent leakage of sensitive data characterizing the subject.

[0155] Referring again to FIG. 7, a subject record may include data elements that characterize subject features using multiple dimensions (e.g., hundreds or thousands of feature dimensions). Certain feature dimensions in the subject record may be useful for the task of interest, while other feature dimensions in the subject record may represent noisy data (e.g., features that are not useful for the task of interest). The high dimensionality of the subject records creates technical challenges with inputting the subject records (or numerical representations thereof) as part of the predictive functions provided by the various AI models associated with the AI ​​system 702. Certain aspects and features of the present disclosure relate to a noisy feature detector 714 that provides a solution to the technical challenges discussed above. In some implementations, the noisy feature detector 714 may be configured to convert the high dimensional subject record into a low dimensional subject record by classifying a subset of subject features of a set of subject features included in the subject record as noise. For example, the noisy feature detector 714 may implement a two-class classification model trained to classify subject features as either predictive of the task of interest or noise. It will be appreciated that the noisy feature detector 714 may also be a multi-class classification model capable of classifying subject features of the subject records into one or more of a plurality of classes (e.g., noisy data, useful but not predictive of the target task, and useful and predictive of the target task). Reducing the dimensionality of the subject records improves the computational efficiency of the AI ​​system 702 by reducing the number of feature dimensions of the subject records that the AI ​​model execution system 710 processes when providing a predictive function. Non-limiting examples of techniques for reducing the dimensionality of the subject records include reducing features based on criteria, reducing features based on feature categories, feature selection techniques, eliminating features classified as noise by a trained classifier model, and other suitable techniques.

[0156] VI. Network Environment Configured to Predict Disease Progression in Subjects with SMA Using Artificial Intelligence Techniques FIG. 8 is a block diagram illustrating an example of a network environment for deploying a trained artificial intelligence model to generate an output predicting disease progression for a subject diagnosed with SMA, according to some aspects of the present disclosure. The network environment 800 can include a user device 110 and an AI system 802. The AI ​​system 802 can be similar to the AI ​​system 702 illustrated in FIG. 7, although the components of the AI ​​system 802 can be different from the components of the AI ​​system 702. In some implementations, the AI ​​system 802 can include an API 808, a query resolver 810, a query text string 812, a trained word-to-vector model 814, a progression prediction system 816, and a communication network 818. The components of the AI ​​system 802 illustrated in FIG. 8 can be in addition to, instead of, or part of any of the components of the AI ​​system 702 illustrated in FIG. 7. The API 808 can be the same as the API 704 illustrated in FIG. 7, and the query resolver 810 can be the same as the query resolver 706 illustrated in FIG. 7.

[0157] The AI ​​system 802 can be configured to generate an output that predicts disease progression for a subject diagnosed with SMA. In some examples, the AI ​​system 802 generates the output automatically, without having to be prompted by a request from the user device 110. In other examples, the AI ​​system 802 generates the output in response to receiving a request from the user device 110. For example, the user device 110 (e.g., operated by a physician or other medical professional) can send a request to the AI ​​system 802. The request can be a request for the AI ​​system 802 to perform a predictive function configured to generate a prediction of disease progression that a particular subject is likely to experience. In some examples, the request includes a subject record 804 that characterizes characteristics of the particular subject. In other examples, the request includes an identifier for the particular subject, such that the identifier is later used to retrieve a subject record 804 that characterizes characteristics of the particular subject. Regardless of how the subject record 804 is accessed or retrieved, the subject record 804 can include data elements that are representative of the condition of the particular subject. As non-limiting examples, the status of a particular subject may include text values ​​such as the subject's diagnosis, the SMA type of the diagnosis, the phenotype observed by a physician, any single stage treatments performed on the particular subject, any multi-stage treatments performed on the subject, the time elapsed between any types of treatments, the genetic profile of the particular subject, clinical information characterizing the particular subject, and other suitable text values. Additionally, the status of a particular subject may represent the current status of the particular subject (e.g., the status of the particular subject at or near the time the request was sent by the user device 110).

[0158] The API 808 can be configured to allow the user device 110 to interact with the AI ​​system 802. Thus, the user device 110 can send requests (including the subject record 804) to the AI ​​system 802 using the API 808. The query resolver 810 can receive the request from the API 808, identify a trained AI model that can resolve the request, and then construct a query for the identified AI model. The query resolver 810 can identify that a request to predict disease progression for a particular subject diagnosed with SMA can be resolved by sending input to the word-to-vector model 814 and providing an output to the user device 110.

[0159] In some implementations, when the query resolver 810 receives a request from the user device 110, the query resolver 810 can extract the subject record 804 from the request if the request includes the subject record 804. In examples where the request includes a unique identifier that identifies the subject record 804, the query resolver can extract the identifier of the subject record 804 and search for the subject record 804 from a data source, such as the data registry 722 shown in FIG. 7. In some implementations, the subject record 804 may be anonymized to prevent the AI ​​system 802 from identifying the identity of the subject characterized by the subject record 804. The query resolver 810 of the AI ​​system can then send the retrieved subject record 804 to a query text string 812, which is configured to generate a partial word sequence using one or more features included in the subject record 804.

[0160] For example, and by way of non-limiting example only, subject record 804 includes at least four data elements. The first data element includes a first text value of "SMA Positive," representing a positive diagnosis of SMA. The second data element includes a second text value of "Type III," representing the type of SMA diagnosed. The third data element includes a third text value of "Proximal Muscle Weakness," representing an observable phenotype for a particular subject. The fourth data element includes a third text value of "Proximal Muscle Weakness," representing the first symptom onset experienced by a particular subject and a given time (e.g., time of request received, current month). 1日 ) and a fourth text value of “6 months” representing the time between the first and second data elements. In some examples, each of the four data elements may include or be associated with a tag that indicates an SMA-related data element, and only the four text values ​​contained in these four data elements may be processed by the query text string 812. In other examples, the four data elements may be associated with a health condition of a particular subject, and these four data elements may be processed by the query text string 812. The query text string 812 may convert the four data elements into a partial word sequence “[SMA positive], [Type III], [proximal muscle weakness], [6 months]”. The partial word sequence may be sent to the query resolver 810 for passing to the word-to-vector model 814, or may be sent directly to the word-to-vector model 814.

[0161] The word-to-vector model 814 may be a machine learning model trained to convert text-based word sequences into numerical representations for the purpose of enabling an AI model to process the word sequences. The word-to-vector model may provide a numerical representation for each word of the word sequence. The word embeddings of the words of the word sequence may be aggregated to numerically represent the word sequence. The numerical representations of multiple words in the word sequence may be compared to identify relationships between the multiple words. Additionally, the aggregated numerical representations of words in a word sequence of two or more word sequences may be compared to identify relationships between the two or more word sequences. The word-to-vector model 814 may be trained to learn the numerical representations of words in the word sequence using a neural network. Thus, the word-to-vector model 814 is input with the partial word sequence of “[SMA positive], [Type III], [Proximal muscle weakness], [6 months]”. In some implementations, the word-to-vector model 814 converts the partial word sequence into a numerical representation (e.g., an N-dimensional word vector). The numerical representations of the partial word sequence may then be input to a progression prediction system 816 trained to predict the remaining words in the partial word sequence. The remaining words that the progression prediction system 816 generates as output representing the predicted order of progression of disease-related events, such as phenotypes or symptoms, that a particular subject is predicted to experience.

[0162] In some implementations, the progress prediction system 816 may be a generative sequence model trained to perform specific language-related tasks, such as language modeling and predictive sentence completion. The generative sequence model may model natural English after being trained using all possible English word sequences. The generative sequence model may be trained to assign probabilities to words based on the sentences in which those words appear. Using the assigned probabilities, the generative sequence model may be configured to predict the remaining word or words that complete the subword sequence (e.g., complete the subsentence). For example, the generative sequence model may be trained to predict that the word "hill" is likely to be the next word to complete the subword sequence of "Jack and Jill went up" and that the word "there" is unlikely to be the next word to complete the subword sequence, because English grammar requires that the subword sequence be followed by a noun.

[0163] In the context of predicting disease progression for a particular subject diagnosed with SMA, the progression prediction system 816 can execute the trained generative sequence model to generate a prediction of the next word to complete the particular word sequence, where the predicted next word represents the predicted disease progression for the particular subject. The progression prediction system 816 can be trained using a training dataset that includes a set of word sequences. Each word sequence in the set of word sequences represents a predicted disease-related event, such as a phenotype or symptom, previously experienced by a subject with SMA. Table 2 below provides an example set of word sequences. Each word sequence in Table 2 is a single or multiple word sequence (e.g., a single word such as “[scoliosis]” or a group of multiple words such as “[walking with a cane]”) that represents the progression of a disease event for a subject previously diagnosed with SMA. [Table 2]

[0164] The progression prediction system 816 can be trained to use tracked disease progressions of subjects previously diagnosed with SMA (e.g., as shown in Table 2) to learn correlations between events along the longitudinal dimension of a subject's disease progression. For example, the progression prediction system 816 can learn that subjects who experience loss of ambulation have a high probability (e.g., above a threshold probability) that the disease will progress to a respiratory infection, which may be caused by the weakening of the muscles that support the spine. Thus, when a particular subject's current condition is loss of ambulation, the progression prediction system 816 can predict that the particular subject's disease progression is likely to include a respiratory infection by predicting words that complete at least a partial word sequence of "loss of ambulation." When disease progression is defined as a word sequence, and each word or group of words in a word sequence represents a disease-related event in a progression sequence of disease-related events, the next disease-related event in a particular subject's disease progression can be predicted by predicting the next word that completes a given partial word sequence.

[0165] Continuing with the above non-limiting example of four data elements in the subject record 804, the progression prediction system 816 receives an input partial word sequence (or a numerical representation of a partial word sequence) of "[SMA positive], [Type III], [proximal muscle weakness], [6 months]". The progression prediction system 816 can generate an output partial word sequence that is predicted to complete the input partial word sequence. The output partial word sequence is a sequence of words that is predicted to complete the input partial word sequence of "[SMA positive], [Type III], [proximal muscle weakness], [6 months]" based on the historical disease progression of a previously treated subject with SMA. The output partial word sequence in this non-limiting example is "[weakness of muscles supporting femur], [walks with cane], [difficulty rising from sitting position], [wheelchair bound]". In other words, the output partial word sequence represents that the particular subject's predicted disease progression includes: (1) weakening of the muscles supporting the femur, then (2) needing a cane to assist with walking, then (3) difficulty sitting unassisted, then (4) needing a wheelchair for mobility for the rest of one's life. The query resolver 810 can transmit the predicted disease progression 806 specific to the particular subject characterized by the subject record 804 to the user device 110 for further evaluation by the user.

[0166] VII. A network environment configured to automatically define subject groups for proposing new clinical studies using artificial intelligence techniques Clustering subject records of subjects with SMA includes identifying clusters of subject records that share common subject characteristics. Clustering subject records can also identify groups of subjects that are similar to each other in some aspects, characteristics, or features. However, clustering subject records is technically challenging given the high dimensionality of subject records. For example, subject records can have hundreds of individual subject features (e.g., dimensions). Thus, clustering high dimensional subject records is problematic for certain clustering techniques, such as k-means clustering. Certain aspects and features of the present disclosure provide a technical solution that allows for clustering of high dimensional subject records that characterize subjects with SMA, for example, for the purpose of defining groups of subjects that are suitable candidates for new or existing clinical studies.

[0167] FIG. 9 is a block diagram illustrating an example of a network environment for intelligently defining subject groups for a new or existing clinical study, according to some embodiments of the present disclosure. The network environment 900 can include an AI system 902 and data stores 904-908 for storing high-dimensional subject records characterizing subjects. Although FIG. 9 illustrates three data stores (e.g., data stores 904-908), it will be appreciated that FIG. 9 is exemplary and thus any number of data stores can be included in the network environment 900. The AI ​​system 902 can be similar to the AI ​​system 702 illustrated in FIG. 7, although the components of the AI ​​system 902 can differ from the components of the AI ​​system 702. The components of the AI ​​system 902 illustrated in FIG. 9 can be in addition to, instead of, or part of any of the components of the AI ​​system 702 illustrated in FIG. 7. The API 910 can be the same as the API 704 illustrated in FIG. 7. Further, the feature selection model 912 and the subspace clustering system 914 may be stored in the AI ​​model data store 724 and may be executable by the AI ​​model execution system 710 shown in FIG.

[0168] In some implementations, the AI ​​system 902 can be configured to automatically detect groups of subjects diagnosed with SMA that are candidates for existing clinical studies. In other implementations, the AI ​​system 902 can be configured to generate predictions of new treatment trajectories that did not exist before and identify subjects that are candidate subjects for new clinical studies. An existing or new clinical study can be, for example, a clinical trial designed to study the clinical outcomes of a new treatment or diagnostic test to determine the effectiveness of the new treatment or diagnostic test. For example, an existing clinical study for SMA can be a clinical trial studying the effect of low-dose celecoxib on SMN2 expression in subjects developing SMA.

[0169] The high dimensional subject record data stores 904-908 can store subject records across multiple entities. As a non-limiting example, subject record data store 904 is operated by a medical facility in the United States, subject record data store 906 is operated by a medical research facility in Italy, and subject record data store 908 is operated by a hospital in Canada. The subject records stored in subject record data store 904 characterize a first group of subjects being treated at the medical facility in the United States. Further, the subject records stored in subject record data store 906 characterize a second group of subjects participating in a clinical study performed at the medical research facility in Italy. Finally, the subject records stored in subject record data store 908 characterize a third group of subjects being treated at the hospital in Canada. Whether data stores 904-908 are geographically distributed across facilities or co-located at a single facility, the subject records stored therein can be grouped using AI-based feature selection techniques to define suitable candidate subjects for existing or new clinical studies.

[0170] The feature selection model 912 may be, for example, a sparse logistic regression, a least absolute shrinkage and selection operator (LASSO), a univariate thresholding (e.g., l 0 -norm minimization, l1 The feature selection model 912 may be executable code representing an instance of any AI-based feature selection model, such as eigenvalue (Eq. 10-norm minimization), least angle regression of LASSO, coordinate descent, proximal techniques, Elastic Net, fused or grouped LASSO, and other suitable feature selection techniques. The feature selection model 912 may be trained to identify which incomplete subsets of subject features of a set of subject features of a subject record are relevant to a target task. By way of illustration, the target task is to identify subjects who are candidates for inclusion in a clinical study related to Evrysdi® (risdiplam, F. Hoffman-La Roche AG). Detecting suitability for a clinical study may be a trained property of the feature selection model 912. For example, the feature selection model 912 may be trained using a training dataset of subject records, each of which includes a label of either “enroll” or “do not enroll” in the clinical study. Based on the training of the feature selection model 912, the feature selection model 912 may learn which incomplete subsets of the set of subject features are relevant to the clinical study. For example, the feature selection model 912 can be trained to learn that subjects between the ages of 2 and 25 who have been diagnosed with SMA type II are suitable candidates for clinical studies based on patterns, correlations, and relationships detected within the training dataset. In this manner, the feature selection model 912 can include subject features related to "age" and subject features related to "SMA type" in an incomplete subset of the set of subject features. The incomplete subset of subject features may or may not be considered high dimensional. Once relevant features have been automatically extracted using the feature selection model 912, the incomplete subset of subject features can be input into a subspace clustering system 914.

[0171] The subspace clustering system 914 can be configured to perform a subspace clustering technique to identify clusters of subject records in different subspaces (e.g., a selection of one or more dimensions). Performing a subspace clustering technique allows clusters of subject records to be formed. The clusters can be defined by a subset of subject features (e.g., subject features that represent dimensional aspects of the subject). For example, and by way of non-limiting example only, an incomplete subset of subject features of a subject record includes gene expression levels of 75 genes, including the SMN2 gene, after a procedure has been performed on the subject. The subspace clustering system 914 is trained to cluster subjects across the 75 genes of the incomplete subset of subject features (e.g., across the 75 dimensions). As part of the clustering of subjects across the 75 genes, the subspace clustering system 914 forms three clusters of subjects associated with expression of the SMN2 gene: "SMN2 expression above threshold", "SMN2 expression below threshold", and "no SMN2 expression". For example, the subspace clustering system 914 can identify a cluster of subjects who experienced expression of the SMN2 gene at levels above a threshold, thereby indicating a potentially successful treatment. The identified cluster of subjects can then be associated with a group identifier stored in the subject group identifier system 916. Furthermore, due to the expression of the SMN2 gene at levels above a threshold, the identified cluster of subjects is determined to be suitable for further existing clinical studies. As another example, the subspace clustering system 914 can identify a subcluster of the "no SMN2 expression" cluster. The subcluster includes subjects who experienced an observable improvement in motor function after the treatment was performed and no increase in SMN2 expression was detected after the treatment. If there is no existing clinical study for this subcluster of subjects, the subspace clustering system 914 can generate a proposal where a new clinical study is created to study subjects who experienced an improvement in motor function after the treatment and no increase in SMN2 gene expression after the treatment.

[0172] VIII. Cloud-based applications can select optimal treatments for subjects with SMA given the context of the subject's records FIG. 10 is a block diagram illustrating an example of a network environment for deploying a trained reinforcement learner to select an action, according to some aspects of the disclosure. The network environment 1000 can include an AI system 1002. The AI ​​system 1002 can be similar to the AI ​​system 702 illustrated in FIG. 7, although the components of the AI ​​system 1002 can be different from the components of the AI ​​system 702. The components of the AI ​​system 1002 illustrated in FIG. 10 can be in addition to, instead of, or part of any components of the AI ​​system 702 illustrated in FIG. 7. The API 1008 can be the same as the API 704 illustrated in FIG. 7, and the query resolver 1010 can be the same as the query resolver 706 illustrated in FIG. 7. The action selection system 1032 can be stored in the AI ​​model data store 724 and can be executable by the AI ​​model execution system 710 illustrated in FIG. 7.

[0173] In some implementations, the AI ​​system 1002 can be configured to select an optimal treatment for a particular subject from a group of treatments 1012-1030. The treatments 1012-1030 can represent potential actions that a physician can initiate while treating a particular subject. For example, and by way of non-limiting example only, treatment 1012 can be nusinersen (SPINRAZA), treatment 1014 can be providing a walking cane, treatment 1016 can be providing a wheelchair, treatment 1018 can be providing a suitable diet plan for a subject with weak jaw muscles, treatment 1020 can be onasemnogene abeparvovec xioi (Zolgensma), treatment 1022 can be a specialized mask or breathing device to support weak respiratory muscles, treatment 1024 can be a feeding tube, treatment 1026 can be physical therapy, treatment 1028 can be a back brace, and treatment 1030 can be a leg brace. A treatment may be a multi-step treatment that may occur sequentially over several phases or stages. Although Figure 10 shows treatments 1012-1030, it will be appreciated that any number of treatments may be performed as actions by or at the direction of a treating physician.

[0174] The treatment observation 1034 may be a data store that stores past observations across previously treated subjects of outcomes in response to each of the treatments 1012-1030. For example, the treatment observation for which treatment 1012 was performed on the subject may be that SMN2 gene expression was increased. As another example, the treatment observation for which treatment 1014 was performed may be that the assistance provided by a cane is insufficient to aid the subject in walking given the progression of degeneration of the subject's thigh muscles (e.g., rectus femoris). In some examples, a survival probability associated with each treatment 1012-1030 may be stored in the treatment observation 1034. For each treatment 1012-1030, the survival probability may be a value (e.g., a percentage) representing the probability that the subject will survive after receiving the treatment. In other examples, the survival probability may also include a value representing the quality of life of the subject after receiving the treatment. In some implementations, the survival probability is automatically determined and updated as new treatment observations are stored in the treatment observation data store 1034. For example, the survival probability is the number of subjects who are alive 30 days after undergoing a treatment, such as surgery. In some implementations, the survival probability may be entered by a physician or the subject after an evaluation of the subject's health. In other examples, the treatment observations data store 1034 may also store any side effects associated with each treatment 1012-1030.

[0175] The treatment selection system 1032 can be trained to learn patterns, correlations, or relationships between each treatment 1012-1030 and the treatment observations stored in the data store 1034. The treatment observations associated with each treatment 1012-1030 can represent a reward function associated with the treatment. The reward function can generate a "reward value," such as a score of "5" out of "5," indicating that the treatment had a strong positive response in the subject. The "reward value" can also be a negative value, such as "-3" out of "5," indicating that the treatment had a strong negative response in the subject. In some implementations, the reward value can be an increase in expression of SMN2 in response to receiving gene therapy. The reward function can be designed to balance any short-term treatment observations with the long-term treatment observations. The short-term and long-term treatment observations can be converted to numbers or vectors (e.g., using a word-to-vector model). The short-term and long-term treatment observations can be weighted individually to reflect the balance between short-term and long-term observable outcomes. The action selection system 1032 can be trained to select actions from among the actions 1012-1030, such that the actions are selected to maximize the reward function. The action selection system 1032 may be any reinforcement learning model, such as, for example, model-free reinforcement learning, policy optimization, policy gradients, model-based reinforcement learning, Q-functions, Q-tables, importance sampling, U-curves, deep reinforcement, supervised reinforcement learning with recurrent neural networks, and other suitable reinforcement learning techniques.

[0176] For example, and by way of non-limiting example only, a subject's condition may be characterized by a subject record 1004, and an observable phenotype 1006 may be an SMA phenotype observed in a subject diagnosed with SMA. The subject record 1004 and the phenotype 1006 may represent a current health condition of a particular subject. The subject record 1004 and the phenotype 1006 are input to the AI ​​system 1002. The API 1008 may be configured to enable exchange of particular data between the AI ​​system 1002 and an external system. The query resolver 1010 may transmit the subject record 1004 or the phenotype 1006 of a particular subject to a treatment selection system 1032 for selecting an optimal action. The treatment selection system 1032 may execute to select a treatment from the treatments 1012-1030 based on a reward function. Once a treatment, such as treatment 1018, is selected, the AI ​​system 1002 may transmit the selected treatment 1018 to a computing device for further evaluation.

[0177] IX. Cloud-Based Applications Can Use Artificial Intelligence Techniques to Predict Disease Progression in Subjects with SMA 11 is a flow chart illustrating an example of a process for predicting disease progression in a subject diagnosed with SMA, according to some embodiments of the present disclosure. Process 1100 may be performed by any of the components illustrated in FIG. 1 and FIG. 7-FIG. 10. For example, process 1100 may be performed by AI system 802. Additionally, process 1100 may be executed to execute an AI model that generates an output that predicts the progression of a phenotype, symptom, or other disease-related event for a particular subject diagnosed with SMA.

[0178] Process 1100 begins at block 1105, where the AI ​​system 802 accesses or retrieves a subject record, for example, corresponding to a particular subject (e.g., a subject being treated at a hospital). The subject record (e.g., an electronic medical record or electronic health record) can include any number of features (e.g., data elements including values ​​for vaccinations, medication history, age, demographics, etc.) collected from or on behalf of the subject. The subject record can include a set of features that characterize the subject's appearance. For example, the subject record can include a feature that indicates, among many other features, that the subject has been diagnosed with SMA Type III.

[0179] Non-limiting examples of features that may be included in a subject record include radiological imaging data, MRI data, genomic profile data, clinical data (e.g., measurements, treatments, treatment responses, diagnoses, severity, medical history), subject-generated data (e.g., notes entered by a subject who has SMA), physician or medical professional generated data (e.g., physician's notes), voice data representing telephone calls between a patient and a physician or other medical professional, administrative data, claims data, health surveys (e.g., Health Risk Assessment (HRS) surveys), third party or vendor information (e.g., out-of-network lab results), public databases related to the subject (e.g., medical journals related to the subject's symptoms), subject demographics, immunizations, radiology reports, pathology reports, utilization information, metadata representing biological samples, social data (e.g., education level, employment status), community specifications, and the like.

[0180] In block 1110, the AI ​​system 802 can extract features related to SMA or the diagnosis of SMA in a subject. In some implementations, any feature related to the diagnosis or treatment of a subject with SMA can be tagged as related to SMA. For example, features related to the results of a motor function test, such as the 6-minute walk test or the Wolff Motor Function Test, can be tagged as related to the diagnosis or treatment of SMA. Tagging features in a subject record can include storing a code (e.g., "0000" or "SMA-TAG") in a data element such that the code is detectable and readable by the AI ​​system 802. The code can be interpreted by the AI ​​system 802 as a feature related to SMA characteristics. A user (e.g., a physician) can tag features individually, or features can be tagged automatically as data is entered into the feature.

[0181] In some implementations, features may not be tagged as relevant to SMA diagnosis or treatment, but instead, the AI ​​system 802 may automatically classify which features are relevant to SMA diagnosis or treatment. For example, the AI ​​system 802 may store a classification model trained to recognize features associated with SMA diagnosis or treatment (or any other relationship to SMA). Any classifier model may be used, including, for example, logistic regression, naive Bayes, stochastic gradient descent, K-nearest neighbors, decision tree models, random forest models, support vector machines (SVMs), and any other suitable of the models.

[0182] In block 1115, the AI ​​system 802 can generate partial word sequences using the SMA-related features identified in block 1110. For example, features identified (in block 1110) as corresponding to SMA include the following: ["SMA Type II"], ["4 months since onset of symptoms"], ["Loss of walking at age 2"], ["Currently 3 years old"], and ["Difficulty sitting upright"]. The AI ​​system 802 can execute the query text string 812 to convert the features of the subject record into partial word sequences such as [SMA Type II, 4 months since onset of symptoms, Loss of walking at age 2, Currently 3 years old, Difficulty sitting upright]. A partial word sequence is a sentence that includes the SMA-related features identified in block 1110, separated by commas.

[0183] The partial word sequence is partial because it represents the subject's current health status with respect to the subject's SMA diagnosis. In block 1120, the AI ​​system 802 receives the partial word sequence as input and converts the partial word sequence into a vector representation using a word-to-vector model (e.g., Word2Vec).

[0184] Once the partial word sequence is converted into a vector representation, certain predictive functions can be performed using the partial word sequence. In block 1125, in connection with predicting disease progression (e.g., progression of an SMA phenotype) for a particular subject diagnosed with SMA, the AI ​​system 802 can input the vector representation of the partial word sequence into a trained generative sequence model (e.g., a natural language processing (NLP) model). In block 1130, the generative sequence model can generate a prediction of one or more next words (e.g., a completion word or completion phrase) predicted to follow the partial word sequence (e.g., to complete the partial word sequence). The predicted next words represent the subject's predicted disease progression of SMA phenotypes, symptoms, diagnosis, or treatments over a period of time. The prediction of the next words likely to complete the partial word sequence represents the next SMA phenotype the subject is predicted to exhibit. For example, each next word output by the generative sequence model represents a predicted phenotype, symptom, treatment, and / or disease-related event that the subject is predicted to experience or exhibit. The prediction of the next word is based on a training dataset that includes word sequences that represent a progression of disease-related events, such as predicted changes in phenotype or symptoms in previously treated subjects with SMA.

[0185] In block 1135, matching techniques such as word matching can be performed to match the predicted completed word sequence to existing disease progression to identify previously treated subjects who have experienced the same or similar disease progression. Additionally, matching the predicted completed word sequence to existing disease progression of another subject can also be performed to identify physicians treating other subjects who have exhibited the same or similar disease progression. In block 1140, if the predicted disease progression meets the early treatment condition, the process 1100 can proceed to block 1145. However, if the predicted disease progression does not meet the early treatment condition, the process 1100 proceeds to block 1155. In some implementations, the early treatment condition can be a rule used to evaluate whether the predicted progression of the SMA phenotype indicates a health risk over a future period, such as the next six months. For example, if the predicted progression of the SMA phenotype for a subject diagnosed with SMA is loss of walking in the next four months, the AI ​​system 802 can interpret the predicted progression as meeting the early treatment condition. In this case, in block 1145, the AI ​​system 802 queries a data store, such as the data registry 722, for identifiers of physicians who have previously treated subjects with the same or similar disease progression (e.g., who are employed by the same hospital or have agreed to be searchable for this purpose).

[0186] In block 1150, the AI system 802 can automatically generate a communication (e.g., an email) and send it to a user device associated with the identified physician. The communication can be a request for a communication session that should be initiated between the physician treating the subject and the physician who has previously treated other subjects with a similar disease progression (identified in block 1145). For example, during the communication session, the physicians can discuss the treatment procedures and clinical outcomes performed on other subjects. The information provided by the physician identified in block 1145 can assist the physician treating the subject in the treatment schedule for the subject before symptoms occur according to the predicted progression of the SMA phenotype. When the early treatment conditions are not met (e.g., when the predicted progression of the SMA phenotype is mild or not predicted to occur over several years), the AI system 802 can search for subject records corresponding to subjects sharing a similar or the same predicted disease progression (in block 1155) and display the relevant treatments and treatment schedules on the user device (in block 1160).

[0187] X. Cloud-based applications can automatically define groups of subjects for new clinical studies using artificial intelligence techniques FIG. 12 is a flowchart illustrating an example of a process for intelligently defining groups of subjects for new or existing clinical studies according to some aspects of the present disclosure. Process 1200 can be executed by any of the components shown in FIGS. 1 and 7-10. For example, process 1200 can be executed by the AI system 902. Further, process 1200 can automatically generate low-dimensional subject records and perform subspace clustering on the subject records to identify candidate subjects for new or existing clinical studies.

[0188] Process 1200 begins at block 1210, where the AI ​​system 902 accesses subject records stored in a data registry, e.g., data registry 722. The subject records can be accessed automatically, at regular or irregular time intervals, or in response to a user input that triggers a predictive function, described in more detail with respect to FIG. 12. At block 1220, some (e.g., not all) or all of the subject records stored in the data registry can be converted to a numerical representation (e.g., a vector representation) using various implementations described herein (e.g., described with respect to FIGS. 1-6). The subject records may be converted or vectorized to a numerical representation in advance, or in real time or substantially in real time, by execution of block 1210.

[0189] At block 1230, the AI ​​system 902 can perform AI-based feature selection on the subject records to select a subset of salient features from the numerical representation of the subject records. For example, given the high dimensionality of the subject records (e.g., potentially having hundreds of features), a feature selection model can be trained to detect and select features within the subject records that are important for performing a task of interest, such as identifying candidate subjects for a new or existing clinical study. At block 1240, for each subject record accessed at block 1210, the AI ​​system 902 can generate a low dimensional numerical representation of the automatically selected salient features of the subject record.

[0190] The feature selection performed in block 1230 may be, for example, sparse logistic regression, least absolute shrinkage and selection operator (LASSO), univariate thresholding (e.g., l 0 -norm minimization, l 1The selection may be performed using any AI-based feature selection model, such as linear regression (linear-norm minimization), least angle regression in LASSO, coordinate descent, proximal techniques, elastic net, fused or grouped LASSO, and other suitable feature selection techniques. The AI-based feature selection model may be trained to identify which incomplete subset of subject features of a set of subject features of the subject recording are relevant to or for performing the target task.

[0191] For example, and by way of non-limiting example only, the target task is to identify subjects who are suitable candidates for inclusion in a clinical study related to Evrysdi® (risdiplam, F. Hoffman-La Roche AG). Detecting whether a subject is a suitable candidate for clinical study can be a trained capability of the feature selection model. The feature selection model can be trained using a training dataset of subject records, each of which includes a label of either "suitable" or "unsuitable" for an existing clinical study. Based on learning of patterns, correlations, and relationships during the training process, the feature selection model can learn which incomplete subsets of the set of subject features are relevant for clinical study. For example, the feature selection model can be trained to learn that subjects between the ages of 2 and 25 who have been diagnosed with SMA type II are suitable candidates for clinical study based on patterns, correlations, and relationships detected within the training dataset. In this way, the feature selection model can include subject features related to "age" and subject features related to "SMA type" in the incomplete subset of the set of subject features.

[0192] In block 1250, the AI ​​system 902 can execute a protocol for automatically defining subject groups for a new or existing clinical study. In some implementations, the subject groups can be defined based on clustering of low-dimensional subject records (or numerical representations thereof). Low-dimensional subject records can still be difficult to process using techniques such as k-means clustering. Thus, for example, the low-dimensional subject records can be clustered in subspaces according to various remaining dimensions of the features. Subspace clustering is performed to identify clusters of subject records in different subspaces (e.g., selection of one or more dimensions). Executing the subspace clustering technique allows clusters of subject records to be formed. The clusters can be defined by a subset of subject features (e.g., subject features representing dimensional aspects of the subject).

[0193] In block 1260, the AI ​​system 902 can generate a clinical study efficacy parameter to represent the efficacy of a new or existing clinical study for the automatically defined subject group. In some implementations, the clinical study efficacy parameter may be a numerical value representing the degree to which the characteristics of the subject group (defined in block 1250) are related to the characteristics of a particular existing clinical study. The trained classification model can be used to classify the characteristics associated with the subject as "effective" or "not effective" based on the clinical outcomes included in the clinical study. The output classification can also be associated with a reliability or relevance parameter that is also output by the classification model. If there is no existing clinical study for the subject group, and if the subject group is classified as having characteristics that are likely to be "effective" for the clinical study, the AI ​​system 902 can generate a suggestion for a new clinical study to be created to study subjects in the subject group. In block 1270, the subject group is selected for a new or existing treatment file based on the clinical study efficacy parameter generated in block 1260.

[0194] XI. Cloud-based applications can select optimal treatments for subjects with SMA given the context of the subject's records 13 is a flow chart illustrating an example of a process for deploying an artificial intelligence model to facilitate the selection of a treatment to perform for a subject diagnosed with SMA, according to some embodiments of the present disclosure. Process 1300 may be performed by any of the components illustrated in FIG. 1 and FIG. 7-FIG. 10. For example, process 1300 may be performed by AI system 1002. Additionally, process 1300 may be performed to execute a reinforcement learning model trained to automatically select a treatment to maximize a reward function, such as the amount of improvement in SMN2 expression.

[0195] Process 1300 begins at block 1310, where the AI ​​system 1002 accesses or retrieves a subject record stored in a data registry, such as data registry 722. The subject record can characterize a particular subject diagnosed with SMA. At block 1220, the subject record accessed or retrieved at block 1210 can be converted to a numerical representation (e.g., a vector representation) using various implementations described herein (e.g., described with respect to FIGS. 1-6). The subject record may be converted or vectorized to a numerical representation in advance or in real-time or substantially real-time by execution of block 1210.

[0196] At block 1330, the AI ​​system 1002 can generate a context vector representing the context of the health state of the particular subject. For example, the context vector is a fixed-length vector that can contextualize the state of the subject record of the particular subject in a numerical form. At block 1340, the context vector representing the particular subject can be input to a treatment selection system, the treatment selection system including a reinforcement learner that learns to reinforce a selected action (e.g., treatment) when a reward is received in response to performing the selected action. The treatment selection system can be any reinforcement learning model, such as, for example, model-free reinforcement learning, policy optimization, policy gradient, model-based reinforcement learning, Q-function, Q-table, importance sampling, U-curve, deep reinforcement, supervised reinforcement learning with recurrent neural networks, and other suitable reinforcement learning techniques.

[0197] At block 1350, the treatment selection system may select an action, such as performing gene therapy to increase expression of the SMN protein. The treatment selection system may intelligently select a treatment from among a group of treatments based on a prediction of the reward received. For example, during the training process, the treatment selection system may detect a pattern within the treatment observations indicating that a subject treated with a first treatment (e.g., risdiplam) between the ages of 10 and 20 is likely to experience a 15%-20% increase in expression of the SMN protein, a subject treated with a second treatment (e.g., nusinersen) between the ages of 2 and 10 is likely to experience a 3% increase in expression of the SMN protein, and a subject treated with a third treatment of weekly physical therapy between the ages of 5 and 12 is likely to experience a 23% increase in 6-minute walk test score (indicating a significant increase in motor function). When the subject is 7 years old, the treatment selection system intelligently selects a treatment among the first treatment, the second treatment, and the third treatment based on the predicted reward. The treatment selection system selects a treatment to maximize a potential reward from the behavior. If the reward function is configured to maximize the rate of increase in expression of SMN protein, the treatment selection system may select the second treatment for the 7-year-old subject because this treatment provides the best reward in terms of increased SMN protein expression. However, if the reward function is configured to maximize the increase in a motor function score, e.g., a 6-minute walk test score, the treatment selection may select the third treatment for the 7-year-old subject to maximize the reward.

[0198] Whatever treatment is selected in block 1360, the treatment selection system receives a response signal after the selected treatment is administered. For example, if the selected treatment is administration of nusinersen, the response signal should include a detected increase in the expression of SMN protein in the subject (whenever available after the treatment). As another example, if the selected treatment is weekly physical therapy, the response signal should include a rate of improvement in the subject's 6-minute walk test score (whenever available after the treatment). In block 1370, the treatment observations of the treatment selection system are updated with the response signal.

[0199] XII. Further Considerations Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the methods and / or one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause the one or more data processors to perform some or all of the methods and / or one or more processes disclosed herein.

[0200] The terms and expressions employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown or described, or portions thereof, recognizing that various modifications are possible within the scope of the invention as claimed. Thus, although the invention as claimed has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be reclassified by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0201] The description that follows provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the description that follows of preferred exemplary embodiments is intended to provide those skilled in the art with an enabling description for implementing various embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.

[0202] Specific details are presented in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0203] XIII. Further Examples As used below, any reference to a series of examples should be understood disjunctively as a reference to each of those examples (e.g., "Examples 1-4" should be understood as "Examples 1, 2, 3, or 4").

[0204] Example 1 is a computer-implemented method including: searching a subject record associated with a subject, the subject record including a set of features characterizing the subject, the subject being diagnosed with spinal muscular atrophy (SMA); extracting a subset of the set of features included in the subject record, each feature of the subset of the set of features being associated with an SMA characteristic; generating a partial word sequence by combining the subset of the set of features to result in a sequence of one or more words, each word of the one or more words representing a characteristic of the subset of features; converting the partial word sequence into a numerical representation using a trained word-to-vector model; inputting the numerical representation of the partial word sequence into a natural language processing (NLP) model trained to predict a completion word or completion phrase to complete the partial word sequence; generating a disease progression representing a predicted phenotype or symptom for the subject over a future timeline (e.g., over the next year, the next five years, the next ten years) based on the completion word or completion phrase output by the NLP model; and outputting an indication that the subject is predicted to exhibit one or more SMA phenotypes included in the disease progression.

[0205] Example 2 is the computer-implemented method of Example 1, further including determining that a predicted progression of the one or more SMA phenotypes specific to the subject satisfies an early treatment condition, where satisfying the early treatment condition indicates a recommendation to administer treatment before the subject exhibits one of the one or more SMA phenotypes.

[0206] Example 3 is the computer-implemented method of Examples 1-2, further comprising: identifying an existing disease progression associated with the de-identified subject when the predicted progression of the one or more SMA phenotypes meets an initial treatment condition; where the existing disease progression matches the predicted progression of the one or more SMA phenotypes specific to the subject, the de-identified subject has been diagnosed with SMA; identifying a user to train the de-identified subject associated with the existing disease progression; and sending a communication to a user device associated with the user, the communication requesting a treatment recommendation for the subject.

[0207] Example 4 is the computer-implemented method of Examples 1-3, identifying existing disease progression associated with the de-identified subject when the predicted progression of the one or more SMA phenotypes does not meet an initial treatment condition, the existing disease progression matches the predicted progression of the one or more SMA phenotypes specific to the subject, and the de-identified subject has been diagnosed with SMA, searching a de-identified subject record characterizing the de-identified subject, extracting a treatment schedule from the de-identified subject record, and transmitting the treatment schedule to the user device.

[0208] Example 5 is the computer-implemented method of Examples 1-4, further including matching the completion word or completion phrase associated with the subject to another SMA phenotype or phenotypes associated with another subject previously treated for SMA, searching an anonymized subject record characterizing the other subject, extracting a treatment schedule from the anonymized subject record, and transmitting the treatment schedule to the user device.

[0209] Example 6 is the computer-implemented method of Examples 1-5, in which a completed word or completed phrase is predicted as the next word in a complete word sequence that includes the partial word sequence, and the completed word or completed phrase is indicative of an SMA phenotype.

[0210] Example 7 is the computer-implemented method of Examples 1-6, in which the disease progression is output at the subject's computing device using a chatbot.

[0211] Example 8 is the computer-implemented method of Examples 1-7, wherein the subject record includes data identified in the electronic medical record corresponding to the subject.

[0212] Example 9 is the computer-implemented method of Examples 1-8, wherein the subject record corresponding to the subject includes a diagnosis of SMA Type I, SMA Type II, SMA Type III, or SMA Type IV.

[0213] Example 10 is the computer-implemented method of Examples 1-9, wherein training the NLP model further includes collecting a training dataset including a set of subject records, where each subject record of the set of subject records corresponds to another subject diagnosed with SMA, and where each subject record of the set of subject records includes one or more features representative of the progression of the SMA phenotype over a period of time; running a learning algorithm associated with the generative sequence model using the training dataset, where the learning algorithm detects patterns associated with the progression of the SMA phenotype exhibited by the set of subjects corresponding to the set of subject records; and generating the NLP model in response to running the learning algorithm associated with the generative sequence model using the training dataset.

[0214] Example 11 is the computer-implemented method of Examples 1-10, further including detecting a data leak associated with the NLP model, where the data leak exposes a feature of a set of features included in the subject record that characterizes the subject, and in response to detecting the data leak associated with the NLP model, executing a data leak prevention protocol to prevent or block the exposure of one feature of the set of features included in the subject record.

[0215] Example 12 is the computer-implemented method of Examples 1-11, wherein executing the data leakage prevention protocol includes retraining the NLP model according to a differential privacy model.

[0216] Example 13 is the computer-implemented method of Examples 1-12, further including using a feature selection model to generate a reduced-dimensional subject record characterizing the subject, where the reduced-dimensional subject record removes one or more features from the set of features included in the subject record, where the one or more features are characterized as noise.

[0217] Example 14 is a system including one or more processors and a non-transitory computer-readable storage medium including instructions that, when executed on the one or more processors, cause the one or more processors to perform some or all of one or more of Examples 1-13 disclosed above.

[0218] Example 15 is a computer program product tangibly embodied in a non-transitory computer-readable storage medium that includes instructions configured to cause one or more data processors to execute some or all of one or more of Examples 1-13 disclosed above.

Claims

1. retrieving a subject record associated with a subject, the subject record including a set of features characterizing the subject, the subject having been diagnosed with spinal muscular atrophy (SMA); extracting a subset of the set of features contained in the subject record, each feature of the subset of the set of features being associated with an SMA trait; generating a partial word sequence by combining the subset of the set of features to result in a sequence of one or more words, each word of the one or more words representing a feature of the subset of the set of features; converting the sub-word sequence into a numerical representation using a trained word-to-vector model; inputting the numerical representation of the partial word sequence into a natural language processing (NLP) model that predicts a completion word or phrase to complete the partial word sequence, the NLP model being trained using a set of subject records corresponding to a set of subjects as a training data set, the set of subject records including a set of word sequences, the training of the NLP model being according to patterns present in the set of word sequences in the set of subject records that are associated with a progression of the SMA phenotype exhibited by the set of subjects; providing, by the NLP model, in response to receiving the numerical representation of the partial word sequence, the completion word or phrase for completing the partial word sequence, the completion word or phrase being indicative of a disease progression representative of a predicted progression of one or more SMA phenotypes specific to the subject over a period of time; providing an indication that said subject is predicted to exhibit said one or more SMA phenotypes involved in said disease progression; 4. A computer-implemented method comprising:

2. 2. The computer-implemented method of claim 1, further comprising determining that the predicted progression of the one or more SMA phenotypes specific to the subject satisfies an early treatment condition, the early treatment condition being a rule used to evaluate whether the predicted progression of the one or more SMA phenotypes specific to the subject indicates a health risk over the period of time, and determining that meeting the early treatment condition indicates a recommendation to implement treatment before the subject exhibits one of the one or more SMA phenotypes.

3. When the predicted progression of the one or more SMA phenotypes satisfies the early treatment condition, identifying an existing disease progression associated with a de-identified subject, wherein the existing disease progression is consistent with the predicted progression of the one or more SMA phenotypes specific to the subject, and the de-identified subject has been diagnosed with SMA; identifying a user to train the de-identified subject associated with the existing disease progression; sending a communication to a user device associated with the user, the communication requesting a treatment recommendation for the subject; The computer-implemented method of claim 2 .

4. when the predicted progression of the one or more SMA phenotypes does not meet the early treatment condition; identifying an existing disease progression associated with a de-identified subject, wherein the existing disease progression is consistent with the predicted progression of the one or more SMA phenotypes specific to the subject, and the de-identified subject has been diagnosed with SMA; retrieving a de-identified subject record characterizing the de-identified subject; Extracting a treatment schedule from the de-identified subject record; transmitting the treatment schedule to a user device; A computer implemented method according to claims 2-3.

5. matching the completed word or phrase associated with the subject with another SMA phenotype or phenotypes associated with another subject previously treated for SMA; retrieving anonymized subject records characterizing said other subjects; and extracting a treatment schedule from the de-identified subject record; transmitting the treatment schedule to a user device; The computer-implemented method of any one of claims 1 to 4, further comprising:

6. 6. The computer-implemented method of claim 1, wherein the completed word or phrase is predicted as a next word in a complete word sequence that includes the partial word sequence, the completed word or phrase being indicative of an SMA phenotype.

7. The computer-implemented method of claims 1 to 6, wherein the disease progression is output at the subject's computing device using a chatbot.

8. The computer-implemented method of any preceding claim, wherein the subject record comprises data identified in an electronic medical record corresponding to the subject.

9. 9. The computer-implemented method of claim 1, wherein the subject record corresponding to the subject includes a diagnosis of SMA Type I, SMA Type II, SMA Type III, or SMA Type IV.

10. Training the NLP model collecting a training dataset comprising a set of subject records, each subject record of the set of subject records corresponding to a different subject diagnosed with SMA, and each subject record of the set of subject records comprising one or more features indicative of progression of an SMA phenotype over a period of time; running a learning algorithm associated with a generative sequence model using the training data set, the learning algorithm detecting patterns associated with the progression of an SMA phenotype exhibited by a set of subjects corresponding to the set of subject records; generating the NLP model in response to executing the learning algorithm associated with the generative sequence model using the training data set; The computer-implemented method of any one of claims 1 to 9, further comprising:

11. Detecting a data leak associated with the NLP model, the data leak revealing a feature of the set of features included in the subject record that characterizes the subject; and In response to detecting a data leak associated with the NLP model, executing a data leak prevention protocol to prevent or block the disclosure of the features of the set of features included in the subject record. The computer-implemented method of any preceding claim, further comprising:

12. 12. The computer-implemented method of claim 11, wherein executing the data leakage prevention protocol includes retraining the NLP model according to a differential privacy model.

13. 13. The computer-implemented method of claim 1, further comprising: using a feature selection model to generate a reduced dimensional subject record characterizing the subject, wherein the reduced dimensional subject record removes one or more features from the set of features included in the subject record, the one or more features being characterized as noise.

14. one or more processors; a non-transitory computer-readable storage medium comprising instructions that, when executed on said one or more processors, cause said one or more processors to perform part or all of the computer-implemented method of any one of claims 1 to 13; A system comprising:

15. A computer program product, tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform part or all of the computer-implemented method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Support system for estimating internal state of object system

    JP2019204484A

  • Integrated Disease Management System

    JP2019535085A

  • Decision-making system and method for determining initiation and type of treatment for patients with advanced illnesses

    JP2020521202A