System and methods for suicide risk mitigation

A machine learning-based approach using gene expression analysis in lymphoblastoid cell lines identifies DEGs to improve suicide risk assessment in bipolar disorder, enhancing detection and intervention efficacy.

WO2025220019A1PCT designated stage Publication Date: 2025-10-23CARMEL HAIFA UNIV ECONOMIC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050349
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-14
Filing Date
2025-04-20
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Current methods for identifying suicide risk in psychiatric disorders, particularly bipolar disorder, are limited by subjective interpretation and lack of objective biomarkers, leading to inadequate detection of high-risk patients.

Method used

A method and system using machine learning-based analysis of gene expression data from lymphoblastoid cell lines to identify differentially expressed genes (DEGs) associated with suicide risk, employing a pre-trained model to determine high-risk subjects and generate personalized treatment plans.

Benefits of technology

Enhances the accuracy and reliability of suicide risk assessment, enabling timely and effective interventions by identifying high-risk patients through objective biomarkers, potentially reducing suicide rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050349_23102025_PF_FP_ABST
    Figure IL2025050349_23102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and systems for predicting suicide risk in patients with psychiatric disorders, and more particularly to machine learning-based approaches for analyzing gene expression data to identify patients at high risk of suicide. The present invention represents methods and systems for predicting and mitigating suicide risk in subjects, providing solutions that leverage gene expression data to identify objective biomarkers associated with increased suicide risk, thereby enhancing the accuracy and reliability of suicide risk assessment. The present invention further provides systems and methods that utilize machine learning techniques to analyze genomic data for the purposes of suicide risk prediction and mitigation, providing an improvement of the field of psychiatric care by enabling more timely and effective interventions, as well as more personalized treatment planning.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHODS FOR SUICIDE RISK MITIGATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 633,797, titled “GENES THAT CAN BE USED TO PREDICT SUICIDE IN BIPOLAR DISORDER USING THE TRANSCRIPTOME OF LYMPHOBLAST CELL LINES OF THE PATIENTS”, filed April 14, 2024, which is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present disclosure relates to methods and systems for predicting suicide risk in patients with psychiatric disorders, and more particularly to machine learning-based approaches for analyzing gene expression data to identify patients at high risk of suicide.BACKGROUND OF THE INVENTION

[0003] Suicide is a major public health concern worldwide, with hundreds of thousands of deaths occurring globally each year. Individuals with psychiatric disorders, particularly bipolar disorder, face an elevated risk of suicidal behavior. Approximately 20% of patients with bipolar disorder die by suicide, making it one of the highest-risk groups among psychiatric conditions.

[0004] Current methods for identifying patients at high risk of suicide rely primarily on clinical interviews, psychological assessments, and patient self-reporting. These approaches suffer from significant limitations, including subjective interpretation, patient reluctance to disclose suicidal thoughts, and the inability to capture biological factors that may predispose individuals to suicidal behavior. The lack of objective, reliable biomarkers for suicide risk represents a critical gap in clinical practice.

[0005] Recent advances in genomics and molecular biology have opened new avenues for investigating the biological underpinnings of complex psychiatric disorders. Gene expression profiling of patient-derived cell lines offers a potential approach for identifying biological signatures associated with disease states or clinical outcomes.

[0006] Machine learning techniques are increasingly being applied to analyze large-scale biological datasets and identifying predictive patterns. There is a proven record of successful implementation of ML-based approaches to determine highly relevant biomarkers for diverse medical conditions, as well as for classification and prediction of their occurrence probability. Similar computational approaches are also used to uncover patterns in geneexpression data that correlate with different clinical phenotypes. The integration of genomic information with machine learning algorithms enables the development of predictive models for assessing various health condition risks.SUMMARY OF THE INVENTION

[0007] Accordingly, there is a need for methods and systems for predicting and mitigating suicide risk in subjects. Specifically, there is a need for solutions that can leverage gene expression data to identify objective biomarkers associated with increased suicide risk, as it may potentially enhance the accuracy and reliability of suicide risk assessment. It is further a need for systems and methods that may utilize machine learning techniques to analyze genomic data for the purposes of suicide risk prediction and mitigation, which may provide an improvement of the field of psychiatric care by enabling more timely and effective interventions, as well as more personalized treatment planning. These advancements may provide clinicians with tools to identify high-risk patients who might not be detected through traditional assessment methods, potentially reducing suicide rates and improving overall patient outcomes in individuals with mental health disorders. There is also a need for a kit for determining a subject being at high risk of committing suicide, as well as for methods for treating subjects having high risk of committing suicide.

[0008] To address the aforementioned needs, the following is suggested.

[0009] According to an aspect of the present disclosure, a method of suicide risk mitigation is provided. The method of suicide risk mitigation may include: obtaining target RNA sequencing data from a biological sample of a target subject; extracting expression levels of a target selection of differentially expressed genes (DEGs) from the obtained target RNA sequencing data; inferring a pre-trained machine-learning (ML)-based model on the extracted expression levels, to determine the subject being at high risk of committing suicide; and generating a treatment plan for the target subject to mitigate the risk of committing suicide. Said target selection of DEGs may include genes selected from a group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

[0010] According to another aspect of the present disclosure, a system for suicide risk mitigation is provided. The system may include: at least one non-transitory memory device,wherein modules of instruction code are stored, and at least one processor associated with said at least one memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to: receive target RNA sequencing data obtained from a biological sample of a target subject suffering from the psychiatric disorder; extract expression levels of a target selection of differentially expressed genes (DEGs) from the obtained target RNA sequencing data; infer a pre-trained machine-learning (ML)-based model on the extracted expression levels, to determine the target subject being at high risk of committing suicide; and generate a treatment plan for the target subject, so as to mitigate the risk of committing suicide. Said target selection of DEGs may include genes selected from a group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, EYST, MAN2B2, EUZP1, and MSH3.

[0011] In some embodiments, said target selection of DEGs may include at least eight genes selected from the group consisting of: EYE1, EM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

[0012] In some embodiments, said target selection of DEGs may include at least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, and GJCL

[0013] In some embodiments, said pre-trained ML-based model may be trained via supervised training, based on a training dataset comprising: a first subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have never attempted suicide, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have one or more suicide attempts, said training samples of the second subset being labelled as indicative of a high risk of committing suicide.

[0014] In some embodiments, the method of suicide risk mitigation may further include: selecting, from a group of subjects, a first sub-group of subjects that have never attempted suicide, and a second sub-group of subjects that have one or more suicide attempts; obtaining training RNA sequencing data from biological samples of the subjects of the first and second sub-groups; extracting expression levels of the target selection of DEGs from the obtained training RNA sequencing data, performing supervised training of the ML-based model based on a training dataset comprising: a first subset of training samples comprising the expression levels of the subjects of the first sub-group, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising the expression levels of the subjects of the second sub-group, said training samples of the second subset being labelled as indicative of a high risk of committing suicide.

[0015] In some embodiments, the method of suicide risk mitigation may further include: selecting, from a group of subjects, a first sub-group of subjects that have never attempted suicide, and a second sub-group of subjects that have one or more suicide attempts; obtaining training RNA sequencing data from biological samples of the subjects of the first and second sub-groups; extracting expression levels for an initial pool of DEGs from the obtained training RNA sequencing data; from the initial pool of DEGs, identifying an initial selection of DEGs providing statistically significant representation of the subjects of the first subgroup with respect to the subjects of the second sub-group, based on p-value and fold change metrics; performing supervised training of the ML-based model to classify subjects as pertaining to the first and the second sub-groups, using a leave-one-out cross validation (LOOCV) algorithm, while concurrently assessing, at each iteration of the LOOCV algorithm, a prediction accuracy of various intermediate sub-selections of DEGs, each representing a distinct combination of DEGs selected from the initial selection of the DEGs; and determining the target selection of DEGs, based on the assessed prediction accuracy.

[0016] In some further embodiments, each iteration of the LOOCV algorithm may include: (a) generating a plurality of said intermediate sub-selections of DEGs; (b) forming a plurality of training datasets, each comprising: a first subset of training samples comprising the expression levels of a respective of said plurality of intermediate selections of DEGs that are associated with the subjects of the first sub-group; and a second subset of training samples comprising the expression levels of the respective of said plurality of intermediate selectionsof DEGs that are associated with the subjects of the second sub-group; said training samples of the first and second subsets being labelled by an association with the first and second subgroups of subjects, respectively. Each iteration of the LOOCV algorithm may further include: (c) performing a supervised training of a plurality of instances of the ML-based model to classify subjects as pertaining to the first and the second sub-groups, based on said plurality of the training datasets, respectively; (d) assessing performance of the plurality of instances of the ML-based model, to identify a sub-plurality of intermediate sub- selections of DEGs from the plurality of the intermediate sub-selections of DEGs providing a highest prediction accuracy between the plurality of instances of the ML-based model; (e) determining a list of DEGs that most frequently appear across the intermediate subselections of DEGs of the identified sub-plurality. Said determining the target selection of DEGs may be performed by selecting DEGs that most frequently appear across lists of DEGs obtained at different iterations of the LOOCV algorithm.

[0017] In some embodiments, said identifying the initial selection of DEGs may be performed by selecting, from the initial pool of DEGs, DEGs with a p- value less than 0.05 and the fold change metrics with log2 fold-change greater than or equal to |0.58|.

[0018] In some embodiments, said biological sample may be a peripheral blood sample.

[0019] In some embodiments, extracting expression levels may include measuring mRNA expression, protein expression or both.

[0020] In some embodiments, said biological sample of the target subject may include lymphoblastoid cell line (LCL) derived from peripheral blood mononuclear cells (PBMCs) of the target subject.

[0021] In some embodiments, said pre-trained ML-based model is selected from a group consisting of: Logistic Regression (LR) model; Random Forest (RF) model; K-Nearest Neighbors (kNN) model, Support Vector Machine (SVM) model, and Neural Network (NN) model.

[0022] According to yet another aspect of the present disclosure, a method of suicide risk mitigation for a subject is provided. The method of suicide risk mitigation for a subject may include: determining expression levels of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A,AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3; and comparing the expression levels of at least one gene to a predetermined threshold, wherein the expression levels beyond the predetermined threshold may indicate a suicide risk in said subject, thereby mitigating the suicide risk of said subject.

[0023] According to yet another aspect of the present disclosure, a method of treating a subject is provided. The method may include: determining expression levels of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11- 459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3; and comparing the expression levels of at least one gene to a predetermined threshold; wherein expression beyond the predetermined threshold indicated of a suicide risk in said subject; and treating a subject indicated as having a suicide risk.

[0024] In some embodiments, determining expression levels comprises measuring mRNA expression, protein expression or both.

[0025] In some embodiments, said method of treating a subject may further include measuring, in a target sample, expression of at least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

[0026] In some embodiments, said method of treating a subject may further include measuring in said sample expression of least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, and GJC1.

[0027] According to yet another aspect of the present invention, a kit is provided. Said kit may include at least two detection molecules selected from detection molecule specific to LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features, and advantages thereof, may best be understood by reference to the following detailed description when read with the accompanying drawings in which:

[0029] Fig. 1A is a flowchart illustrating the process of identifying differentially expressed genes (DEGs) and training and testing the machine learning models, according to some embodiments of the present invention;

[0030] Fig. IB is a flowchart illustrating an algorithm of identifying a target selection of DEGs, according to some embodiments of the present invention;

[0031] Fig. 1C is a flowchart illustrating an algorithm of training the ML-based model using the target selection of DEGs, according to some embodiments of the present invention;

[0032] Fig. 2A is a flowchart illustrating the concept of the present invention, according to some embodiments of the present invention;

[0033] Fig. 2B is a heatmap of logio gene count in the initial selection of DEGs between the two groups of ‘SUICIDE’ and ‘NON- SUICIDE’, according to some embodiments of the present invention;

[0034] Fig. 2C is a chart showing gene count of the top 30 genes with the lowest p- values, with red indicating suicidal individuals and blue indicating non-suicidal individuals, according to some embodiments of the present invention;

[0035] Fig. 3A is a diagram displaying the dysregulated KEGG pathways associated with the DEGs between the two groups, according to some embodiments of the present invention;

[0036] Fig. 3B is a chart showing the most prevalent conditions in the DisGeNET database, according to some embodiments of the present invention;

[0037] Fig. 3C is a chart showing prominent GAD disease classes, according to some embodiments of the present invention;

[0038] Fig. 3D is a chart showing significant biological processes, according to some embodiments of the present invention;

[0039] Fig. 3E is a chart showing molecular functions associated with the DEGs, according to some embodiments of the present invention;

[0040] Fig. 4 A is a heat map organized based on the descending order of the subtraction of the average read count of the top predicted genes of the two groups (‘SUICIDE’-‘NON- SUICIDE’), according to some embodiments of the present invention;

[0041] Fig. 4B is a chart showing a gene count for the most dominant predicting DEGs (the target selection), according to some embodiments of the present invention;

[0042] Fig. 5A is a chart showing the variation in ML-based model accuracy for a logistic regression architecture in relation to the number of genes used for the target selection, ranging from 1 to 10 genes, according to some embodiments of the present invention;

[0043] Fig. 5B is a chart showing the accuracy of different alternative architectures of the ML-based model, when specifically using the target selection of 8 genes, according to some embodiments of the present invention;

[0044] Figs. 6A-6F are sets of charts showing ROC curve with cross validation and a confusion matrix of different alternative architectures of the ML-based model, tested on the target selection of 8 genes, according to some embodiments of the present invention;

[0045] Fig. 7A is a table showing details of BD patients that were used in the research, according to some embodiments of the present invention;

[0046] Fig. 7B is a table showing clinical details of unknown labelled BD patients and the results of the suicide risk prediction provided by different architectures of the ML-based model, according to some embodiments of the present invention;

[0047] Fig. 8 is a block diagram, depicting a computing device which may be included in a system for suicide risk mitigation, according to some embodiments of the present invention;

[0048] Fig. 9 is a block diagram, depicting a system for suicide risk mitigation, according to some embodiments of the present invention;

[0049] Fig. 10A is a flow diagram, depicting a method of suicide risk mitigation, according to some embodiments of the present invention;

[0050] Fig. 10B is a flow diagram, depicting a method of suicide risk mitigation, according to some embodiments of the present invention;

[0051] Fig. 10C is a flow diagram, depicting a method of treating a subject, according to some embodiments of the present invention.

[0052] It will be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, whereconsidered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION OF THE PRESENT INVENTION

[0053] One skilled in the art will realize the invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting of the invention described herein. Scope of the invention is thus indicated by the appended claims, rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.

[0054] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the present invention. Some features or elements described with respect to one embodiment may be combined with features or elements described with respect to other embodiments. For the sake of clarity, discussion of same or similar features or elements may not be repeated.

[0055] Although embodiments of the invention are not limited in this regard, discussions utilizing terms such as, for example, “processing,” “computing,” “calculating,” “determining,” “establishing”, “analyzing”, “checking”, “choosing”, “selecting”, “omitting”, “training”, “deciding”, “raising” or the like, may refer to operation(s) and / or process(es) of a computer, a computing platform, a computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer’s registers and / or memories into other data similarly represented as physical quantities within the computer’s registers and / or memories or other information non-transitory storage medium that may store instructions to perform operations and / or processes.

[0056] Although embodiments of the invention are not limited in this regard, the terms “plurality” and “a plurality” as used herein may include, for example, “multiple” or “two or more”. The terms “plurality” or “a plurality” may be used throughout the specification todescribe two or more components, devices, elements, units, parameters, or the like. The term “set” when used herein may include one or more items.

[0057] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Additionally, some of the described method embodiments or elements thereof can occur or be performed simultaneously, at the same point in time, concurrently, or iteratively and repeatedly.

[0058] In embodiments of the present invention, some steps of the claimed method may be performed using machine-learning (ML)-based models or may include actions performed on ML-based models. ML-based models may be configured or “trained” for a specific task, e.g., classification or regression.

[0059] In some embodiments, ML-based models may be artificial neural networks (ANN).

[0060] A neural network (NN) or an artificial neural network (ANN), e.g., a neural network implementing a machine learning (ML) or artificial intelligence (Al) function, may refer to an information processing paradigm that may include nodes, referred to as neurons, organized into layers, with links between the neurons. The links may transfer signals between neurons and may be associated with weights. A NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training a NN for the specific task may involve adjusting these weights based on examples. Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN. Typically, the neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights. A processor, e.g., CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.

[0061] In some embodiments, ML-based models may be logistic regression models. Logistic regression is a statistical model that may be used for binary classification tasks. It may predict the probability of a binary outcome based on one or more predictor variables. The logistic regression model may use a logistic function to model the probability of the dependent variable. The logistic function, also known as the sigmoid function, may produce an output between 0 and 1, representing the probability of the occurrence of the event. Themodel may be trained by adjusting the coefficients of the predictor variables to minimize the difference between the predicted probabilities and the actual outcomes. Training may involve using optimization techniques such as gradient descent. The logistic regression model may be represented by a mathematical equation, where the coefficients are applied to the predictor variables, and the result is transformed using the logistic function to produce the probability. A processor, e.g., CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.

[0062] It shall be understood that, depending on specific embodiments of the present invention, other known architectures of ML-based models may be used herein.

[0063] It should be obvious for the one ordinarily skilled in the art that various ML-based models can be implemented without departing from the essence of the present invention. It should also be understood that in some embodiments ML-based model may be a single ML- based model or a set (ensemble) of ML-based models realizing as a whole the same function as a single one. Hence, in view of the scope of the present invention, the abovementioned variants should be considered equivalent.

[0064] The concept of the present invention is further explained in greater detail, beginning with a description of the research that underpins the present invention.

[0065] The research investigated the genetic signatures associated with a high risk of suicide in Bipolar disorder (BD) patients through RNA sequencing analysis of lymphoblastoid cell lines (LCLs). By identifying differentially expressed genes (DEGs) and their enrichment in pathways and disease associations, the research uncovered insights into the molecular mechanisms underlying suicidal behavior. LCL gene expression analysis revealed significant enrichment in pathways related to primary immunodeficiency, ion channel, and cardiovascular defects. Notably, genes such as LCK, KCNN2, and GRIA1 emerged as pivotal in these pathways, suggesting their potential roles as biomarkers. Machine learning models trained on a subset of the patients and then tested on other patients demonstrated high accuracy in distinguishing low and high-risk of suicide in BD patients. Moreover, the study explored the genetic overlap between suicide-related genes and several psychiatric disorders. This comprehensive approach enhanced the understanding of the complex interplay between genetics and suicidal behavior, laying the groundwork for future prevention strategies.

[0066] As known, suicide is a deeply complex and tragic phenomenon that continues to be a significant public health concern worldwide. Annually, the global incidence of suicide surpasses 700,000 fatalities exhibiting variations observed among distinct age cohorts and geographical regions. Suicide often unfolds within the framework of psychiatric work. Research findings show that more than 90% of individuals who succumb to suicide have a diagnosed psychiatric illness and most individuals who attempt suicide also experience such disorders. Known statistics shows that a significant number of individuals diagnosed with BD may undergo at least one episode of self-harm in their lifetime, while approximately 20% died by suicide. Thus, BD carries the highest risk of suicidal behavior among other psychiatric disorders and currently, there is no way to identify those BD patients that are at a high risk of suicide.

[0067] The risk associated with suicide in BD varies depending on the disease attributes and the phase or stage of the illness. Existing studies suggest that BD II has a higher suicide rate than BD I. This observation emphasizes the multifaceted severity of BD II with major depressive episodes, carrying the highest risk of suicide among individuals with BD. Following closely are BD patients with mixed episodes, where both depression and mania are present, posing a significant risk as well. Conversely, during manic episodes, marked by elevated mood and energy levels, the patient has the lowest risk of suicide. Additionally, individuals with rapid cycling BD face an increased risk compared to those without rapid cycling patterns. When considering sex differences, certain studies suggest that men exhibit higher rates of fatal suicide attempts, whereas women have more non-fatal suicide attempts. However other research indicates no sex disparity in suicide attempts.

[0068] In patients with BD, a history of prior suicide attempts emerges as one of the most robust predictors of both fatal and non-fatal suicide attempts. While a significant portion, approximately 56% of those who die by suicide do so on their first attempt the majority of individuals who attempt suicide non-fatally do not ultimately die by suicide. Subsequently, a few studies have also highlighted the strong association between a first-degree family history of suicide and suicide deaths among individuals with BD. Thus, it may confirm the presence of both a family history and a personal history of suicide attempts as a critical suicide risk factor in BD. Other factors include early life trauma such as childhood abuse or stress, which are linked to higher rates of suicide attempts in individuals with BD. Moreover, the combination of childhood abuse and drug usage drastically increases the likelihood ofsuicide attempts. Psychosocial stressors, including interpersonal conflicts and occupational issues, have also been associated with increased suicide risks.

[0069] The study of BD has been revolutionized by various cellular model systems, especially using the induced pluripotent stem cell (iPSC) model which allows researchers to explore the cellular and molecular mechanisms underlying the disorder in neuronal cell types derived from patients themselves. The iPSC-based neuronal models from BD patients have reported mitochondrial dysfunction, neuron hyper-excitability, disruption of calcium signaling etc. These BD studies revealed key electrophysiological differences in hippocampal neurons of BD patients, particularly between those responsive (LR) and non- responsive (NR) to lithium treatment. An elevated fast after-hyper-polarization (AHP) has shown to be a distinct trait in BD neurons, which has been implicated in various neurological disorders. The study led to the development of computational models for the prediction of lithium treatment. Given the high risk of suicide among patients with BD, there is a strong need to understand the neurobiological and genetic factors of suicide, an iPSC model may shed more light on the related mechanism.

[0070] Genome-wide association studies (GW AS) have paved the way in identifying genetic loci associated with suicide across psychiatric disorders. Notable genes including 5- HTT, SLC6A4, and the 5-HT1 to 5-HT7 series involved in serotonergic neurotransmission, along with tryptophan hydroxylase genes TPH1 and TPH2, which are identified for their crucial roles in the neurotransmission dysfunctions tied to various suicidal behaviors. Research also extends to genes such as BDNF and its receptor NTRK2, and others like COMT and MAPK1, exploring their links to suicide risk. Furthermore, genetic variations in genes like AKT1P, GSK3B, and the Alpha2a-adrenergic receptor (ADRA2) are associated with impulsivity, a significant risk factor in BD suicide. Additionally, the PENK gene, along with IL-7 and TMX3, involved in stress and anxiety responses, were linked to suicidal behavior in BD patients. Comprehensive multi-ancestry GW AS analysis has pinpointed genes such as DRD2, FURIN, NLGN1, SOX5, PDE4B, and CACNG2 as involved in suicidal behavior, alongside identifying potential biomarkers within the immunoglobulin gene family for BD. Despite these advancements, a classification based solely on GWAS data remains challenging.

[0071] The research underpinning the present invention has sought biomarkers for suicidal tendencies in BD patients to construct machine learning (ML) algorithms that will enable agood prediction of the patients who are at risk of attempting suicide. To do this, LCLs from 20 patients were grown and their RNA has been extracted. Six of the patients died by suicide; blood samples were collected before their deaths. Seven BD patients have been monitored over a long period and have not attempted suicide over many years. The remaining seven patients have either attempted suicide or have a family history of suicide (and one did not have either). Interestingly, when analyzing enrichment for biological pathways, it was found that brain-related and psychiatric disease-related pathways, despite the source of the RNA being from LCLs. The patients were then divided into two groups of those who died by suicide and those who were at a low risk since they were monitored over the years. Then machine learning (ML) algorithms were applied to train on DEGs derived from the RNA sequencing data. Using cross-validation schemes, the research demonstrated that prediction is possible with a very low error rate using RNA sequencing data of the patient's LCLs.

[0072] Table 1, shown in Fig. 7A, provides details about the BD patients used in the research. The research included conducting RNA sequencing on samples from 20 patients diagnosed with BD. Of these, the status of 13 patients-whether they have died by suicide or have tried to attempt suicide was known. For this study, the RNA sequencing data from these 13 patients were used to train and test ML-based models.

[0073] All participants provided informed consent before participating in the research. Approval for the research was obtained from the Research Ethics Board of the University of Cagliari, Italy. The participants, diagnosed with BD, were categorized into subtypes based on those patients who have died by suicide and those who are at a low risk of suicide.

[0074] In the research, LCLs derived from Peripheral blood mononuclear cells (PBMCs) were obtained from both groups of patients. Briefly, PBMCs were isolated from blood samples using BD vacutainer according to the manufacturer's guidelines. Following which they were exposed to Epstein-Barr virus (EBV) harvested from the B95-8 cell line. Within a week post-infection, the cells transformed and began to aggregate, forming large clusters over time.

[0075] The LCLs were cultured in T25 tissue culture flasks, using a complete RPMI medium composed of RPMI 1640 (Biological Industries, Cat no: 01-100-1A), IX Anti-Anti (Thermofisher Scientific, Cat no: 15240062), 1% Glutamax (Thermofisher Scientific, Cat no: 35050061), 1% Sodium pyruvate (Thermofisher Scientific, 11360070), and 15% heat- inactivated fetal bovine serum (Sigma, Cat no: F9665). Media changes were performedevery other day. Passaging of the LCL was performed to maintain optimal cell growth and prevent over confluency. Cells were sub cultured when reaching a density of 200,000 cells / ml to sustain a logarithmic growth.

[0076] Total cellular RNA was extracted from 5 million LCLs using 800 pl of TRIzolTM reagent (Thermofisher Scientific, Cat no. 15596026) and subsequently chilled on ice for 5 minutes before long-term storage at -80° Celsius. The total RNA was carefully isolated using the zymo RNA clean & concentrator kit, according to manufacturer's protocol. Following the extraction, the quality and integrity of the RNA samples were integrity number (RIN) within 7-8, confirming both good quality and suitability for sequencing.

[0077] RNA sequencing and analysis Libraries of RNAs extracted from LCLs derived from BD patients were prepared using the TruSeq RNA Library Prep Kit v2 (Illumina), adhering to the manufacturer's guidelines. Quality control was conducted on the raw FASTQ files using FastQC (v. 0.11.5). Reads were aligned to the human genome (GRCh38.104) and quantified with STAR (v2.7.9a). Differential gene expression analysis was carried out using DESeq2 (v 1.34.0). To mitigate false positives in identifying DEGs, a false discovery rate (FDR) analysis was applied using the Benjamini-Hochberg (BH) procedure to adjust p- values for multiple hypothesis testing. Genes were considered significant DEGs if they met the criteria of a -valuc <0.05 and a log2 fold-change >|0.58|, with a total of 843 genes meeting these thresholds. The step-by-step process for identifying DEGs is illustrated in Fig. 1A. The top 50 genes were selected as predictors for further machine-learning analysis (depending on embodiments of the present invention, the selection of top 843 genes and / or selection of top 50 genes may be referred herein as an “initial” selection of DEGs providing statistically significant representation of the subjects of the ‘NON-SUICIDE’ sub-group with respect to the subjects of the ‘SUICIDE’ sub-group (also referred herein as ‘first’ and ‘second’ sub-groups, respectively)).

[0078] In the research, five distinct supervised classifiers were assessed for their ability to predict suicidal and non-suicidal outcomes. Logistic Regression (LR) is a fundamental statistical model utilized for binary classification, leveraging a logistic function to estimate probabilities and optimize parameters through maximizing data likelihood, though its performance may falter with non-linear data relationships. The K-Nearest Neighbors (K- NN) algorithm, a non-parametric method, classifies new instances by using the majority vote from the 'K' nearest training samples, but struggles with large or high-dimensional datasetsdue to the heavy computational demand of distance calculations. Support Vector Machine (SVM) creates a hyperplane in high-dimensional space to separate classes with optimal margins. It adapts well to nonlinear boundaries using the kernel trick, although tuning its hyperparameters can be challenging and resource-intensive. Random Forest (RF) enhances robustness and accuracy by building multiple decision trees and integrating their outcomes, which are suitable for various tasks but potentially complex and difficult to interpret. Some previous researches also showed that SVM and RF were the most effective machine-learning algorithms in distinguishing between lithium responders and non-responders among BD patients. Inspired by the human brain, neural Networks (NN) excel in pattern recognition and learning complex nonlinear relationships, making them powerful yet resource-heavy and challenging to manage due to their propensity for over-fitting. Each algorithm offers unique strengths and limitations, with the choice often depending on data characteristics and specific problem needs. In the present research, each of these classifiers has been carefully selected to address different aspects of the suicide prediction analysis, taking into account their respective strengths and limitations according to the specific requirements of the data and the task at hand. Logistic regression, random forest, and neural network classifiers gave the best performance with high accuracy (>96%).

[0079] The feature selection process for suicide prediction, specifically, an algorithm of identifying the target selection of DEGs, is shown in Fig. IB. Fig. 1C shows the continuation of the algorithm of Fig. IB, specifically, an algorithm of training the ML-based model using the target selection of DEGs, according to some embodiments of the present invention.

[0080] In particular, the research employed an intensive approach that combines leave-one- out cross-validation (LOOCV) (the algorithm, known in the art) with an exhaustive feature selection process to identify the most crucial features for the prediction, especially useful for smaller dataset. Specifically, LOOCV was utilized for the abovementioned dataset, which consists of 13 samples (corresponding to 13 patients). This method involved training the ML-based model on 12 samples and testing on the remaining one, cycling through iteratively so that each sample served as the test set once, resulting in 13 total iterations. Concurrently, during each LOOCV iteration, all possible combinations of up to 10 genes were examined, taken from the initial top 50 DEGs (also referred herein as ‘intermediate’ sub-selections of DEGs, each representing a distinct combination of DEGs selected from the initial selection of the DEGs), assessing the LR model's accuracy with these subsets on the left-out sample,where the outcome was binary - 100% or 0% accuracy based on correct or incorrect predictions. Successful combinations led to a tally of each gene's appearance frequency, and after 13 iterations, the top 20 genes were identified based on this frequency. As shown in Fig. IB, this process was repeated to yield 13 lists of the top 20 genes, from which the results were further distilled to pinpoint the top 10 consistently appearing genes (also referred herein as the target selection of DEGs). Finally, as shown in Fig. 1C, to evaluate the ME-based models of various abovementioned architectures, these top 10 genes (the target selection of DEGs) were used to split the full dataset into training and testing sets multiple times (50 splits with a 50-50% ratio), and the mean accuracy from these tests provided a robust assessment of the model's effectiveness.

[0081] The effectiveness of the approach was validated by testing five supervised classification algorithms using the specified input features (the target selection of DEGs) for binary classification. The outcomes of the model predictions were classified into four categories: true positives (tp), false positives (fp), true negatives (tn), and false negatives (fn). Evaluation metrics included accuracy, the Receiver Operating Characteristic (ROC) curve, and the confusion matrix, as shown in Figs. 5A-5B and 6A-6F. To ensure robustness, cross-validation methods were applied, with results being averaged or aggregated for the confusion matrix across iterations. Performance was assessed using the Area Under the Curve (AUC) of the ROC and metrics from the confusion matrix such as Accuracy and Precision.

[0082] Accuracy of ROC was calculated as:tp Precision= - — ; tp+JP tpRecall= - — ; tp+fa where, y' represents the model's prediction, while y denotes the actual class.

[0083] The research focused on developing biomarkers that are easily and cheaply extracted from biological samples that will aid in identifying BD patients who are at an increased risk of committing suicide.

[0084] Various aspects of the research underpinning the present invention are further described in greater detail.

[0085] To identify biomarkers for suicide risk, sequencing of RNA extracted from LCLs of 20 BD patients were performed. Of these 6 patients died by suicide ('SUICIDE' sub-group), 7 have been followed over the years and have never attempted suicide, nor do they have a family history of suicide ('NON-SUICIDE' sub-group), and the classification of the remaining 7 patients having a family history of suicide and some have attempted non-fatal suicide ('UNKNOWN' sub-group). Differential gene expression analysis was conducted between the first two groups of BD subjects ('SUICIDE' and 'NON-SUICIDE'), leading to identifying 843 DEGs with a fold change greater than 1.5 (log2 fold-change >|0.58|) and a -valuc of less than 0.05. These are presented in Fig. 2B as a heatmap. Fig. 2B displays the gene counts for the 30 most significantly DEGs based on the lowest p- values. DEGs between 'SUICIDE' and 'NON-SUICIDE' LCLs exhibited enrichment in pathways related to "Primary immunodeficiency" such as LCK, CD8A, AICDA, ZAP70, and CIITA. Specifically, LCK, AICDA, ZAP70, and CIITA demonstrated higher gene counts in 'SUICIDE' LCLs compared to 'NONSUICIDE' LCLs, indicating increased expression levels associated with suicidal behaviour. However, CD8A shows a contrasting pattern, with higher gene counts observed in 'NON-SUICIDE' LCLs compared to 'SUICIDE' LCLs, suggesting a distinct expression profile in non-suicidal individuals. DEGs between the 'SUICIDE' and 'NON-SUICIDE' groups were enriched in Immunoglobulin genes, such as IGHV1-3, IGHV4-31, IGHV3-33, IGHVL46, IGHV4-59, IGHV4-61 were enriched in 'NON-SUICIDE' samples whereas IGKV4-1 was enriched in 'SUICIDE' samples. 0.8% (470 genes) of the total human genes are immunoglobulin genes. However, among the DEGs, 2.3% (19 out of 843 DEGs) are immunoglobulin genes. Furthermore, when examining DEGs with minimum counts of 10, 100, and 1000, the proportion of immunoglobulin-related genes increases to 2.3%, 3.1%, and 5.2%, respectively. This pattern indicates a significant enrichment of immunoglobulin-related genes in the DEGs.

[0086] Figs. 2A-2C show a flow chart, genes and pathways that distinguish the two subgroups of 'SUICIDE' and 'NON-SUICIDE' patients. Fig. 2A is a flowchart illustrating the concept of the present invention. Fig. 2B represents a heatmap of logio gene count of the DEGs between the two groups of 'SUICIDE' and 'NON-SUICIDE'. The heat map was organized based on the descending order of the subtraction of the average read count of thetwo groups ('SUICIDE'-'NON-SUICIDE'). Fig. 2C shows the gene count of the top 30 genes with the lowest p-values, with red indicating suicidal individuals and blue indicating non- suicidal individuals.

[0087] Further analysis of the DEGs included exploring KEGG pathways, disease associations, and functional annotations, as depicted in Figs. 3A-3E. The KEGG pathway analysis highlights the intricate biochemical and physiological networks potentially involved in the underlying conditions studied, as shown in Fig. 3A. For example, pathways such as primary immunodeficiency, phenylalanine, tyrosine, tryptophan biosynthesis, and hematopoietic lineage are significantly dysregulated between BD patients from the two groups ('SUICIDE' vs. 'NON-SUICIDE'). Additionally, the Intestinal Immune Network for IgA Production and Primary Immunodeficiency pathways suggest a significant role of immune system functions in the pathology, potentially linking systemic immune responses to neurological and psychiatric conditions.

[0088] In the Disease DisGeNET analysis, a wide array of conditions linked to the DEGs was identified, indicating a broad impact of these genetic variations across various diseases, as shown Fig. 3B. The list includes autoimmune and inflammatory diseases such as Libman- Sacks Disease and Systemic Lupus Erythematosus, as well as conditions related to abnormal blood clotting, like thrombosis. Importantly and surprisingly, many brain-related pathways came up in the analysis, despite the origin of the cells (LCLs). These pathways include BD Schizophrenia, Cerebral infraction Left and Right hemispheres, Subcortical infraction, Autistic Disorder, neuroinflammatory diseases like Gliosis, Astrocytosis, and more.

[0089] The Disease GAD (Genetic Association Database) list encompasses a wide range of disorders, highlighting the complex interplay between genetics and various health conditions. Diseases such as Hypertension, Asthma, and Coronary Artery Disease reflect chronic conditions that significantly impact public health and are influenced by both genetic predispositions and lifestyle factors. Autoimmune and inflammatory diseases like Systemic Lupus Erythematosus and Rheumatoid Arthritis, as well as metabolic disorders such as Type 2 Diabetes, underscore the genetic basis of immune system dysfunction and metabolic dysregulation. Furthermore, the inclusion of conditions like schizophrenia and substance use disorder illustrates the genetic factors contributing to psychiatric disorders and addictive behaviors, demonstrating the broad spectrum of diseases that genetic research can potentially address. In the research, an in-depth analysis of diseases characterized by specificup-regulated genetic or molecular markers has been conducted, focusing on how these enhancements influence disease pathology. This exploration has been pivotal in identifying key biological pathways and potential therapeutic targets, enhancing our understanding of their underlying mechanisms.

[0090] The GAD Disease class encompasses a broad spectrum of medical categories, reflecting the diverse genetic underpinnings and environmental influences on various health conditions, as shown in Fig. 3C. Interestingly, the data indicates that neurological and developmental pathways are dysregulated even in the white blood cells of the patients. Another interesting pathway that was dysregulated was "AGING" as accelerated aging was reported in BD and according to the obtained data, this may be even further implied in BD patients who died by suicide. "PSYCH" was one of the top dysregulated pathways in the LCLs. The "IMMUNE" pathway is among the top 3 dysregulated pathways. Immune pathways have been shown repeatedly to be dysregulated in BD. The most dysregulated pathway was "CARDIOVASCULAR". Associations between BD and heart defects were previously shown. The present research shows that these pathways are even more dysregulated in BD patients, who are at a high risk of suicide.

[0091] The research also included a detailed analysis across five functional annotation categories for whom the DEGs between the groups were enriched including UP KW BIOLOGICAL PROCESS, UP KW CELLULAR COMPONENT, UP KW MOLECULAR FUNCTION, UP KW PTM, and UP SEQ FEATURE as shown in Figs. 3D and 3E. Specifically, within the UP KW BIOLOGICAL PROCESS category, significant dysregulation in processes such as cell adhesion was identified, which is vital for cellular assembly and integrity, and various immune-related functions, including adaptive and innate immunity, underscoring the genes' pivotal roles in systemic defence mechanisms, as shown in Fig. 3D. Additionally, this category highlighted the involvement of genes in essential developmental and reparative processes such as angiogenesis and osteogenesis, crucial for maintaining vascular and skeletal health. In the present study, the UP KW MOLECULAR FUNCTION revealed critical roles of various proteins in cellular mechanisms, with a particular emphasis on ion channels, sodium channels, and voltage-gated channels, which are essential for electrical signalling and cellular communication and have been shown to alter in neurons derived from BD patients, as shown in Fig. 3E. Proteins such as cytokines and growth factors were also identified as changed in LCLs of suicide victims. Theseproteins are important for cell signalling and regulatory processes that influence cell growth, differentiation, and immune responses, suggesting implications for suicide. Additionally, the identification of host cell receptors for virus entry, integrins, and receptors underscores the relevance of cell-environment interactions in suicide, which facilitates crucial processes such as viral entry, cellular adhesion, and signal transduction across cellular membranes.

[0092] The research initiated with the goal of developing biomarkers that could be quickly and easily implemented in clinical settings. Consequently, it was essential to establish a robust protocol that facilitates the prediction of BD patients who are at a high risk of suicide using RNA-seq datasets, as illustrated in Fig. 2A. As described above, the top 50 genes were selected from a pool of 843 DEGs to serve as predictors for machine learning algorithms. Hence the research involved implementation of a rigorous approach to identify the most dominant and predictive genes using LOOCV and exhaustive feature selection. The algorithm involved training the model on 12 samples and testing on the remaining one, cycling through each sample as the test set, completing 13 iterations. During each iteration, all combinations of up to 10 genes from the initial 50 were tested, assessing the Logistic Regression model's accuracy on the left-out sample, which was binary either 100% or 0% based on prediction accuracy. This allowed to identify the top 20 genes appearing most frequently across iterations. After repeating this process, the initial selection of genes was narrowed down to the top 50 genes consistently identified across the 13 lists. The gene counts and heat maps for these top 50 genes are shown in Figs. 4A and 4B. Further, the selection was narrowed down the top 10 genes to train and test all ML models of different architectures, as shown in Figs. IB and 1C.

[0093] After identifying the top 10 genes (LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1), the research proceeded with training and testing various machine learning (ML) models. The study employed five types of supervised classification algorithms: Logistic Regression (LR), Random Forest (RF), K-Nearest Neighbors (K-NN), Support Vector Machine (SVM), and Neural Network (NN). Cross- validation methods were used to avoid over-fitting (as shown in Fig. IB). The models were tested using subsets varying from 1 to 10 predictors (the size of the target selection was varied from 1 to 10 DEGs), as illustrated in Fig. 5A. Notably, it was found that 8 features provided high accuracy of more than 95%.

[0094] The target selection of DEGs was narrowed down to 8 genes with high predictive power of suicidal and non-suicidal outcomes, as indicated in Fig. 5B. Here, we divided the data with 50% for training and 50% for testing, a process repeated 50 times to ensure the robustness of our ML algorithms. The accuracy results were as follows: Logistic Regression: 98.3%+5.0%, Random Forest: 96.0%+10.8%, K-Nearest Neighbors: 77 ,0%+30.7%, SVM: 90.7%±20.6%, and Neural Network: 99.7%±2.3%. Both Logistic Regression and Neural Network showed higher accuracy rates compared to the other ML models. Figs. 6A-6F display the performance of these classifiers through the area under the ROC curves and confusion matrices, with the following accuracies: Logistic Regression: 0.89+0.00, Random Forest: 0.99+0.00, K-Nearest Neighbors: 0.99+0.00, SVM: 0.61+0.22, and Neural Network: 0.67+0.20.

[0095] The trained models were also used to assess the status of 7 BD patients whose conditions were not previously identified, yet some of their clinical data suggests that they are at a high risk of suicide, such as a first-degree relative who died of suicide and previous suicide attempts. For this classification, the majority decision by each of the classifiers ('Combined Results') was taken. The analysis revealed that patients ROU7776B and 21868 are at a low risk of suicide, whereas the other five BD patients are classified as at a high risk. These findings are detailed in Table 2 (Fig. 7B). Interestingly, the BD patient with no previous suicide attempts and no family history was classified as a low risk for suicide. The other one that was classified as a low -risk had no previous attempts and one family member who died by suicide. The rest of the patients were classified as a high-risk for suicide. These patients had previously attempted suicide or had a few family members who died by suicide (with one patient 14501 having just one family member who died by suicide).

[0096] The research has unearthed distinct gene expression signatures intricately linked to suicide and various facets of its phenomenology by utilizing RNA sequencing data from LCLs of BD patients who died by suicide compared to BD patients who are at a low risk of suicide. After identifying genes and pathways linked with suicide, they were further used as features to train classification algorithms in order to construct a predictor for suicide risk in BD using biomarkers that are easily and cheaply available without much stress to the patient. As mentioned above, first, the initial selection of 843 genes was determined. Then, for the initial selection, several pathways, disease, and functional annotation analyses were applied to identify possible mechanisms of suicide. In this study, the biological samples were LCLs- immune-related cells. Accordingly, it was expected to find immune-related changes and it was not expected to find brain-related implications. For brain-related mechanisms, it was assumed that iPSC work related to differentiating the cells into neurons would be required. It was surprising, therefore, to find many neuronal pathways that are altered between patients at a high risk of suicide and patients with the low risk. This suggests that when differentiating neurons, it is possible to find neurophysiological changes in neurons derived from patients at a high risk of suicide compared to those at a low risk.

[0097] Among the primary findings, DEGs were enriched for the primary immunodeficiency pathway between LCL from BD patients who died by suicide and patients who are at a low risk of suicide. For example, genes such as LCK (log2 fold change of 0.81), AICDA (log2 fold change of 1.329), and CIITA (log2 fold change of 0.975) exhibited increased expression levels in LCL from BD patients who died by suicide compared to those at a low risk of suicide. These genes are critical for immune system functioning; LCK is involved in T-cell receptor signaling. CIITA (Class II Major Histocompatibility Complex Transactivator) regulates MHC (Major Histocompatibility complex) class II genes. AICDA is crucial for B-cell antibody response. The differential expression of these genes might reflect an underlying dysregulation in the immune response, which could be contributing to the pathophysiology of suicide. On the other hand, the CD8A gene (log2 fold change of 2.008), which is vital for T-cell mediated cytotoxicity, showed lower expression in LCLs from people who died by suicide, suggesting a potential protective role of this gene against suicide. This pattern of gene expression adds an intriguing layer to our understanding of how genes associated with both high and low risk can be linked to immune dysregulation and psychiatric conditions, particularly in the context of suicide in BD. Continued surveillance of these genes is crucial, as it will enable further exploration into the complex relationships between immune function disturbances and suicide.

[0098] Further supporting the link between immune function and the brain, the study also highlighted significant alterations in cell adhesion and ion channel pathways in LCLs of BD patients who died by suicide compared to the patients at a low risk. These pathways are known to influence neuronal connectivity and excitability. From the cell adhesion pathway, genes like NTNG1 and NRXN1 showed increased expression, with log2 fold changes of 2.27 and 1.42, respectively in LCLs of BD patients who died by suicide compared to BD patients with a low risk of suicide. Both genes are crucial for axonal guidance and synapticstability. Additionally, CD8A was upregulated with a log2 fold change of 2.008 and this gene also participates in the primary immunodeficiency pathway. The gene CLDN10 is a key component of the blood-brain barrier (BBB) tight junctions and was down-regulated with a log2 fold change of -2.16 in LCLs of people who died by suicide. Dysregulation in cell adhesion genes can lead to altered excitatory / inhibitory balance within the neural circuit.

[0099] Ion channels are crucial in modulating neuronal excitability and signal transduction, with growing evidence linking their dysregulation to psychiatric disorders. The study highlights significant gene expression differences in BD patients who died by suicide compared to those with a low risk of suicide. Specifically, voltage-dependent calcium channel genes such as CACNB2 (Calcium Voltage-Gated Channel Auxiliary Subunit Beta 2) showed a notable increase (log2 fold change of 1.318) in BD patients who died by suicide. In contrast, CACNA1A (Calcium Voltage-Gated Channel Sub-unit Alphal A) exhibited a log2 fold change of 0.955, predominantly in patients with a low risk of suicide. These genes mediate calcium influx into a wide variety of electrically excitable cells, including cardiac, muscle, neurons, and sensory cells. In a similar pattern, SCN11A and SCN1B genes, which encode for voltage-gated sodium channels, showed differential expression. SCN11A (sodium voltage-gated channel alpha subunit 11) was up-regulated (log2 fold change of 1.355) in BD patients who died by suicide, while SCN1B (sodium voltage-gated channel beta subunit 1) showed a higher expression (log2 fold change of 2.33) in patients with a low risk of suicide. This differential expression of calcium and sodium channel genes suggests distinct roles, potentially categorizing them as either low or high predictors of suicide risk. Moreover, genes encoding for calcium-activated potassium channels such as KCNN2 (Potassium Calcium- Activated Channel Subfamily N Member 2) and KCNMB 1 (Potassium Calcium-Activated Channel Subfamily M Regulatory Beta Subunit 1) were also highly expressed in BD patients who died by suicide with a log2 fold change of 3.228 and 1.988 respectively. Additionally, the serotonin receptor gene HTR3B also showed a significant upregulation (log2 fold change of 3.4) in BD patients who died by suicide, reaffirming the well- established link between serotonin signaling disruptions and both depressive and suicidal behaviors. These insights reveal a complex interplay between genetic factors and ion channel function, further elucidating the multifaceted nature of suicide risk in BD patients.

[0100] Likewise, these findings support the established link between ion channel dysfunctions and their role in both cardiovascular and neurological disorders, highlightingtheir critical roles in cellular excitability and signaling. This connection has been well- documented, with ion channel abnormalities known to contribute to conditions such as cardiac arrhythmias and various neurological maladies, suggesting a shared pathophysiological foundation. In the study, it was observed that genes such as KCNN2 (log2 fold change of 3.228), KCNMB 1 (log2 fold change of 1.98), SCN 11 A (log2 fold change of 1.35), CACNB2 (log2 fold change of 1.318) and glutamate ionotropic receptor AMPA type subunit 1 also known as GRIA1 (log2 fold change of 2.85), showed increased expression in LCLs from BD patients who died by suicide, while a few ion channels like SCN IB (log2 fold change of 2.33) and KCNE4 (log2 fold change of 4.06 ) were more prominently expressed in BD patients with a low risk of suicide. This pattern of differential gene expression underscores the significant role these channels play in influencing both neuronal and cardiovascular functions, which in turn affect mood and behavior. The data points to a complex interaction between genetic factors that impact both neurological health and cardiovascular stability, reinforcing the potential of these genes as dual biomarkers for assessing the risk of suicide as well as cardiovascular diseases.

[0101] It was also identified that genes associated with astrocytosis and gliosis are relevant to suicide risk. Astrocytosis and gliosis involve the proliferation and hypertrophy of astrocytes, usually in response to central nervous system pathologies like neurodegenerative disease, tumor growth, and trauma. The research found five genes to be differentially expressed, among which BDNF and ITGA5 were notably upregulated in BD patients who are at a low risk of suicide, with log2 fold changes of 2.04 and 1.30, respectively. BDNF is pivotal for neuronal survival and development, influencing the growth, differentiation, and synaptic plasticity of neurons. It also plays a critical role in neuronal and synaptic regeneration and repair following injury. On the other hand, ITGA5 is essential for cell adhesion to the extracellular matrix, supporting cell migration during development and wound healing, and is involved in angiogenesis. The roles of neurotrophins like BDNF and integrins like ITGA5 are crucial in guiding neuron positioning, growth, and maturation, highlighting why their low expression might be observed in suicide samples. However, the GRM8 gene was found to be elevated in LCLS of BD patients who died by suicide (log2 fold change of 1.38), and encodes for glutamate metabotropic receptor 8 and plays a significant role in glutamate neurotransmission. This gene's function is particularly pertinent in BD, where glutamate system dysregulation is often implicated. Additionally, theGRM8 gene's involvement in processes like astrocytosis can be linked to altered glutamatergic signaling which may influence the proliferation and hypertrophy of astrocytes during neuroinflammatory responses.

[0102] Intriguingly, beyond their hematopoietic lineage, LCLs exhibit a remarkable expression profile reminiscent of neuronal activity, as evidenced by RNA sequencing studies. This unexpected convergence underscores the potential of LCLs as a valuable model for probing the genetic underpinnings of neuropsychiatric and neurodevelopmental disorders. In the analysis, several genes, including SHANK2 (log2 fold change 1.96), CACNG5 (log2 fold change 4.93), ANK3 (log2 fold change 2.04), and NRG3, (log2 fold change 4.29) were found to be significantly upregulated in LCLs of BD patients who died by suicide compared to patients who are at a low risk of suicide. SHANK2 plays a crucial role in synaptic transmission and has been associated with autism spectrum disorders and schizophrenia. CACNG5 encodes a subunit of the voltage-dependent calcium channels and has been previously known to be linked to schizophrenia. ANK3 is crucial for encoding ankyrin-G, a protein involved in the structural stability of neurons by anchoring integral membrane proteins to the cytoskeleton. Previous studies have shown that variations in this gene are associated with both BD and schizophrenia. NRG3, encoding neuregulin 3, is a part of the neuregulin family. Neuregulins act through binding to the ERBB family initiating important signaling pathways that influence cell growth, and differentiation. Genetic studies have linked variations in the NRG3 gene to schizophrenia and BD. This concurrent association underscores the potential role of these genes as high-risk factors for suicide in BD. The upregulation of these genes not only reinforces their critical function in neuropsychiatric pathologies but also highlights them as promising biomarkers for assessing suicide risk.

[0103] To delve deeper, the research enhanced the understanding of the genetic factors that influence BD, particularly in distinguishing between 'SUICIDE' and 'NON-SUICIDE' tendencies, by refining the research methodology through a robust ML-based protocol. Out of 843 DEGs, 50 potential biomarkers were selected (the initial selection of DEGs) to be analyzed via machine learning techniques. By applying algorithm described with reference to Fig. IB (and also with reference to Fig. 9 below) the study established a list of the top 10 markers (the target selection of DEGs), which showcased consistently high predictive accuracy. These genes are LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D,SNX8, TBC1D4, and GJC1. Importantly, these top 10 genes differ from the initial 843 DEGs, showing clearer gene count distinctions between 'SUICIDE' and 'NON-SUICIDE' cases. The abovementioned methodology involved evaluating the performance of machine learning models ranging from 1 to 10 predictors to validate the robustness of these biomarkers.

[0104] The findings of the research were particularly significant in the evaluation of predictive performance across various machine learning models. When narrowed down to an optimized set of eight specific genes from our top 10 (the target selection of DEGs), enhanced prediction accuracies were observed. The accuracy levels achieved by different models were as follows: Logistic Regression at 98.3%±5.0%, Random Forest at 96.0%±10.8%, K-Nearest Neighbors at 77.0%±30.7%, Support Vector Machine at 90.7%±20.6%, and Neural Networks at an impressive 99.7%±2.3%. These findings emphasize both the robustness and reliability of our chosen genetic markers and underscore their clinical effectiveness in distinguishing between 'SUICIDE' and 'NON-SUICIDE' BD patients. The trained models were used finally to assess 7 BD patients, previously undiagnosed but identified as high-risk due to factors like family history of suicide and personal attempts. Using a majority decision from the classifiers, the analysis indicated that two patients, ROU7776B and 21868, are at a low risk of suicide, while the remaining five are at high risk. Interestingly, Rou7776B has never attempted suicide, nor did he have any family history of suicide. Patient 21868 attempted suicide once and had one family member who died by suicide. Most high-risk patients either had multiple personal attempts or multiple family members with a history of suicide (except for patient 14501, who had one only one family member). This implies that these findings could contribute to the development of targeted therapeutic strategies and improve preventive measures in BD management, as well as in the treatment of other conditions associated with suicide risks. These aspects are crucial for enhancing patient outcomes and could potentially lead to personalized treatment approaches based on genetic predispositions.

[0105] It shall be understood that, although the research was conducted on patients suffering from bipolar disorder (BD), the suggested methods and systems may be efficiently applicable to other subjects, including those suffering from differing psychiatric disorders, as well as those not suffering from or diagnosed with any type of psychiatric disorder. It shall be clear to a person skilled in the art that concentrating the research on a narrow cohortof subjects — those suffering from BD, including individuals who have attempted suicide and those who have never attempted it, while being monitored for a substantial amount of time — allowed for the efficient determination of features (the target selection of genes (DEGs)) outlining the distinction between the two groups. However, once the target selection of genes is determined, it can then be effectively applied to a much larger cohort, potentially including subjects not suffering from bipolar disorder. Furthermore, this may be relevant to both ML training and inference stages. For example, the ML-based model may be trained on a training dataset formed from biological samples of different subjects within such a large and diverse cohort, including training samples comprising the expression levels of the target selection of DEGs from subjects who have at least one suicide attempt and / or a family record of suicides, and training samples comprising the expression levels of the target selection of DEGs from subjects who have never attempted suicide and / or have no family record of suicide. Accordingly, such a pre-trained ML-based model may further be inferred on the expression levels of the target selection of DEGs of the target subject to evaluate / classify / determine their risk of committing suicide.

[0106] It shall further be understood that although the research was conducted using lymphoblastoid cell lines (LCL) for extracting DEGs data and using it for training and inferring the ML-based model, this aspect shall not be considered limiting the scope of the present invention. Thus, in some general embodiments, the sample may be a blood sample. In some particular embodiments, the sample may be selected from blood plasma, whole blood, blood serum or peripheral blood mononuclear cells. In some particular embodiments, the sample may be plasma. In some particular embodiments, the blood may be peripheral blood. In some particular embodiments, the sample may be lymphoblastoid cell line (LCL).

[0107] Reference is now made to Fig. 8, which is a block diagram depicting a computing device, which may be included within an embodiment of the system for suicide risk mitigation, according to some embodiments.

[0108] Computing device 1 may include a processor or controller 2 that may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing or computational device, an operating system 3, a memory device 4, instruction code 5, a storage system 6, input devices 7 and output devices 8. Processor 2 (or one or more controllers or processors, possibly across multiple units or devices) may be configured to carry out methods described herein, and / or to execute or act as the various modules, units,etc. More than one computing device 1 may be included in, and one or more computing devices 1 may act as the components of, a system according to embodiments of the invention.

[0109] Operating system 3 may be or may include any code segment (e.g., one similar to instruction code 5 described herein) designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 1, for example, scheduling execution of software programs or tasks or enabling software programs or other modules or units to communicate. Operating system 3 may be a commercial operating system. It should be noted that an operating system 3 may be an optional component, e.g., in some embodiments, a system may include a computing device that does not require or include an operating system 3.

[0110] Memory device 4 may be or may include, for example, a Random- Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short-term memory unit, a long-term memory unit, or other suitable memory units or storage units. Memory device 4 may be or may include a plurality of possibly different memory units. Memory device 4 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM. In one embodiment, a non-transitory storage medium such as memory device 4, a hard disk drive, another storage device, etc. may store instructions or code which when executed by a processor may cause the processor to carry out methods as described herein.

[0111] Instruction code 5 may be any executable code, e.g., an application, a program, a process, task, or script. Instruction code 5 may be executed by processor or controller 2 possibly under control of operating system 3. For example, instruction code 5 may be a standalone application or an API module that may be configured to extract expression levels of a target selection of differentially expressed genes (DEGs) from the obtained target RNA sequencing data; infer a pre-trained machine-learning (ML)-based model on the extracted expression levels, to determine the subject being at high risk of committing suicide; and generate a treatment plan for the target subject to mitigate the risk of committing suicide, as further described herein. Although, for the sake of clarity, a single item of instruction code 5 is shown in Fig. 8, a system according to some embodiments of the invention may includea plurality of executable code segments or modules similar to instruction code 5 that may be loaded into memory device 4 and cause processor 2 to carry out methods described herein.

[0112] Storage system 6 may be or may include, for example, a flash memory as known in the art, a memory that is internal to, or embedded in, a micro controller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disk (BD), a universal serial bus (USB) device or other suitable removable and / or fixed storage unit. Various types of input and output data may be stored in storage system 6 and may be loaded from storage system 6 into memory device 4 where it may be processed by processor or controller 2. In some embodiments, some of the components shown in Fig. 8 may be omitted. For example, memory device 4 may be a non-volatile memory having the storage capacity of storage system 6. Accordingly, although shown as a separate component, storage system 6 may be embedded or included in memory device 4.

[0113] Input devices 7 may be or may include any suitable input devices, components, or systems, e.g., a detachable keyboard or keypad, a mouse and the like. Output devices 8 may include one or more (possibly detachable) displays or monitors, speakers and / or any other suitable output devices. Any applicable input / output (RO) devices may be connected to Computing device 1 as shown by blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device or external hard drive may be included in input devices 7 and / or output devices 8. It will be recognized that any suitable number of input devices 7 and output device 8 may be operatively connected to Computing device 1 as shown by blocks 7 and 8.

[0114] A system according to some embodiments of the invention may include components such as, but not limited to, a plurality of central processing units (CPU) or any other suitable multi-purpose or specific processors or controllers (e.g., similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.

[0115] Reference is now made to Fig. 9 which is a block diagram, depicting system 100 for suicide risk mitigation, according to some embodiments of the present invention.

[0116] The present invention is based on the surprising finding that patients suffering from a psychiatric disorder have a distinguished gene expression signature indicative of increased suicide risk.

[0117] According to some embodiments of the invention, system 100 may be implemented as a software module, a hardware module, or any combination thereof. For example, system 100 may be or may include computing devices such as element 1 of Fig. 8. Furthermore, system 100 may be adapted to execute one or more modules of instruction code (e.g., element 5 of Fig. 8) to request, receive, analyze, calculate and produce various data.

[0118] As further described in detail herein, system 100 may be adapted to execute one or more modules of instruction code (e.g., element 5 of Fig. 8) in order to perform steps of the claimed method.

[0119] As shown in Fig. 9, arrows may represent flow of one or more data elements to and from system 100 and / or among modules or elements of system 100. Some arrows have been omitted in Fig. 9 for the purpose of clarity.

[0120] As described above, the process of evaluating a suicide risk of a target subject, may start with receiving biological sample 10A of the target subject. In some embodiments, said target subject may be patient suffering from psychiatric disorder, e.g., bipolar disorder (BD). In some embodiments, sample 10A may be a blood sample. In some embodiments, sample 10A is selected from blood plasma, whole blood, blood serum or peripheral blood mononuclear cells of the target subject. In some embodiments, sample 10A may be plasma. In some embodiments, the blood may be a peripheral blood. In some embodiments, sample 10A is lymphoblastoid cell line (LCL).

[0121] In some embodiments, sample 10A may be used as an input to RNA extraction stage 10, in which target RNA sequencing data 10B may be extracted from sample 10A (e.g., as described above, in relation to the conducted research).

[0122] In some embodiments, system 100 may include DEG analysis module 110. In some embodiments, DEG analysis module 110 may be configured to extract expression levels 110A of the target selection of differentially expressed genes (DEGs), from the obtained target RNA sequencing data 10B. In some embodiments, DEG analysis module 110 may be configured to extract said expression levels by measuring, in obtained target RNA sequencing data 10B, mRNA expression, protein expression or both.

[0123] In some embodiments, said target selection of DEGs said target selection of DEGs may include genes selected from a group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302,RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3, which were determined as highly dysregulated between ‘SUICIDE’ and ‘NON-SUICIDE’ groups of patients, as shown in Fig. 4B and described above.

[0124] In some embodiments, said target selection of DEGs may include at least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11- 459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3, as it has been established during the research described above that the set of 8 genes may provide high reliability of the suicide risk prediction.

[0125] In some particular embodiments, said target selection of DEGs may include at least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, and GJC1, which are the genes that demonstrated, in the conducted research, the highest relevance to suicide risk prediction, by being that most frequently appearing across lists of DEGs obtained at different iterations of the LOOCV algorithm (as described with reference to Figs. IB and 1C above).

[0126] In some embodiments, system 100 may further include ML-based model 120 and training module 140.

[0127] In accordance to the research described above, in some embodiments, ML-based model 120 may be trained via supervised training, based on a training dataset including: a first subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have never attempted suicide, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have one or more suicide attempts, said training samples of the second subset being labelled as indicative of a high risk of committing suicide. It shall be understood that, depending on embodiments, ML-based model 120 may be trained using semi- supervised training instead of supervised training, hence, in the context of the present disclosure this types of ML-based training shall be considered equivalent).

[0128] In some embodiments, the RNA sequencing data of subjects used for the training dataset may be from patients suffering from a psychiatric disorder. Depending on the embodiment, this disorder may be the same as the one from which the target patient is suffering, or it may be a different disorder.

[0129] In some embodiments, biological sample 10A of the target subject, as well as the biological samples of the subjects used for the training dataset, may include lymphoblastoid cell line (LCL) derived from peripheral blood mononuclear cells (PBMCs) of the target subject and the subjects used for the training dataset, respectively.

[0130] In some embodiments, training module 140 may be connected to external database 20A of historical data about subjects (e.g., patients having psychiatric disorders). Training module 140 may be further configured to access database 20A to select, from a group of subjects, a first sub-group of subjects that have never attempted suicide, and a second sub-group of subjects that have one or more suicide attempts. Module 140 may be further configured to obtain, from database 20A, training RNA sequencing data from biological samples of the subjects of the first and second sub-groups. Training module 140 may be further configured to extract expression levels of the target selection of DEGs from the obtained training RNA sequencing data (e.g., via DEG analysis module 110). Training module 140 may be further configured to perform supervised training of ML-based model 120 based on a training dataset including: a first subset of training samples comprising the expression levels of the subjects of the first sub-group, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising the expression levels of the subjects of the second sub-group, said training samples of the second subset being labelled as indicative of a high risk of committing suicide. This way, ML-based model 120 may be pretrained to perform suicide risk prediction, same as discussed above with reference to the conducted research.

[0131] Additionally, in some embodiments, training module 140 may be also configured to perform the procedure of determining the target selection of DEGs (e.g., as discussed with reference to Figs. IB and 1C above).

[0132] In particular, training module 140 may be further configured to access database 20A to select, from a group of subjects, a first sub-group of subjects that have never attempted suicide, and a second sub-group of subjects that have one or more suicide attempts. Training module 140 may be further configured to obtain, from database 120A,training RNA sequencing data from biological samples of the subjects of the first and second sub-groups. Training module 140 may be further configured to extract (e.g., via DEG analysis module 110) expression levels for an initial pool of DEGs from the obtained training RNA sequencing data. Training module 140 may be further configured to, identify, from the initial pool of DEGs, an initial selection of DEGs (e.g., same as abovementioned top 50 genes from 843 genes) providing statistically significant representation of the subjects of the first sub-group with respect to the subjects of the second sub-group, based on p-value and fold change metrics. E.g., in some embodiments, training module 140 may be configured to identify the initial selection of DEGs by selecting, from the initial pool of DEGs, DEGs with a p-value less than 0.05 and the fold change metrics with log2 fold-change greater than or equal to |0.58|.

[0133] Module 140 may be further configured to perform supervised training of ML- based model 120 to classify subjects as pertaining to the first and the second sub-groups, using a leave-one-out cross validation (LOOCV) algorithm, while concurrently assessing, at each iteration of the LOOCV algorithm, a prediction accuracy of various intermediate subselections of DEGs, each representing a distinct combination of DEGs selected from the initial selection of the DEGs.

[0134] In some embodiments, training module 140 may be configured to perform, at each iteration of the LOOCV algorithm, the following: (a) generate a plurality of such intermediate sub-selections of DEGs; and (b) form a plurality of training datasets, each including: a first subset of training samples comprising the expression levels of a respective of said plurality of intermediate selections of DEGs that are associated with the subjects of the first sub-group; and a second subset of training samples comprising the expression levels of the respective of said plurality of intermediate selections of DEGs that are associated with the subjects of the second sub-group; the training samples of the first and second subsets may be labelled by an association with the first and second sub-groups of subjects, respectively.

[0135] Training module 140 may be further configured to perform, at each iteration of the LOOCV algorithm, the following: (c) perform a supervised training of a plurality of instances of the ML-based model (e.g., ML-based model 120) to classify subjects as pertaining to the first and the second sub-groups, based on said plurality of the training datasets, respectively; (d) assess performance of the plurality of instances of the ML-basedmodel, to identify a sub-plurality of intermediate sub-selections of DEGs from the plurality of the intermediate sub-selections of DEGs providing a highest prediction accuracy between the plurality of instances of the ML-based model; and (e) determine a list of DEGs that most frequently appear across the intermediate sub-selections of DEGs of the identified subplurality. Training module 140 may be further configured to determine the target selection of DEGs performed by selecting DEGs that most frequently appear across lists of DEGs obtained at different iterations of the LOOCV algorithm.

[0136] Module 140 may be further configured to determine the target selection of DEGs, based on the assessed prediction accuracy.

[0137] In some embodiments, ML-based model 120 may have at least one of the following architectures: Logistic Regression (LR) model; Random Forest (RF) model; K- Nearest Neighbors (kNN) model, Support Vector Machine (SVM) model, and Neural Network (NN) model.

[0138] In some embodiments, system 100 may be further configured to infer pre-trained ML-based model 120 on extracted expression levels 110A of the target subject, to determine if the target subject is at high or low risk of committing suicide, e.g., by classifying the target subject by pertinence to one of the following classes: the class labelled as indicative of a low risk of committing suicide (also referred herein as ‘NON-SUICIDE’), and the class labelled as indicative of a high risk of committing suicide (also referred herein as ‘SUICIDE’).

[0139] Accordingly, ML-based model 120 may be configured to output predicted suicide risk 120A. It shall be understood that, depending on specific embodiments, suicide risk 120A may be represented as discrete value, showing pertinence to respective classes, or as a probability of pertinence to the respective class, which may be considered as a probability of a risk of committing a suicide.

[0140] In some embodiments, system 100 may further include treatment plan generation module 130. Treatment plan generation module 130 may be configured to receive the predicted suicide risk 120A and to generate, based on suicide risk 120A, treatment plan 130A for the target subject, so as to mitigate the risk of committing suicide (e.g., if it is necessary for the particular case, that is, if suicide risk 120A indicates high risk of committing suicide). In some embodiments, treatment plan generation module 130 may be further configured to receive treatment history 30A and / or other clinical characteristics 40A of the target subject, and generate treatment plan 130A further based thereon. Clinicalcharacteristics 40A may, e.g., include age, BMI, albumin levels, complete blood counts (CBC), C-reactive protein results etc. Treatment history 30A may, e.g., include diagnosed diseases, previously prescribed drugs and their observed effect on the target subject’s condition etc. It shall be understood that treatment plan generation module 130 may operate through some pre-determined logic and algorithms, based on treatment plans and approaches, approved by medical institutions to be applied for the respective health conditions.

[0141] In some embodiments , treatment plan 130A may include immediate interventions such as safety planning, which involves identifying other warning signs, coping strategies, and support networks. Additionally, the plan may incorporate pharmacotherapy, such as the administration of antidepressants or mood stabilizers, and psychotherapy, including cognitive-behavioral therapy (CBT) specifically tailored for suicide prevention. If the patient is already undergoing treatment, the current plan (e.g., described in treatment history 30A) may be adjusted to enhance monitoring and support, such as increasing the frequency of therapy sessions, involving family members in the treatment process, and limiting access to means of self-harm. The current treatment plan may be further adjusted by adjusting dosages of previously prescribed drugs (e.g., antidepressants, such as Selective Serotonin Reuptake Inhibitors (SSRIs), Serotonin-Norepinephrine Reuptake Inhibitors (SNRIs), Atypical Antidepressants, Tricyclic Antidepressants (TCAs) and other known types if antidepressants), and / or prescribing drugs with different active ingredients (which may be represented in treatment plan 130A). The treatment plan may also include regular reassessment of the patient's risk factors and protective factors to ensure that the interventions remain effective and responsive to the patient's needs. Treatment plan generation module 130 may be further configured to evaluate clinical characteristics 40A when generating / adjusting pharmacotherapy aspects of treatment plan 130A.

[0142] In some embodiments, treatment plan generation module 130 may be further configured to generate a report indicating the predicted suicide risk category for the target patient (e.g., ‘high’, ‘medium’, ‘low’ etc.), e.g., based on the probability of the target patient pertaining to ‘ SUICIDE’ or ‘NON- SUICIDE’ group. Treatment plan generation module 130 may be further configured to specify actual actions of the treatment plan adjustment (dosage, changing type of drugs etc.).

[0143] Referring now to Fig. 10A, a flow diagram is presented, depicting a method of suicide risk mitigation, by at least one processor (e.g., processor 2 of Fig. 1), according to some embodiments.

[0144] As shown in step S1005, the at least one processor (e.g., such as processor 2 of Fig. 8) may receive a target RNA sequencing data (e.g., target RNA sequencing data 10B, shown in Fig. 9) from a biological sample of a target subject (e.g., biological sample 10A, shown in Fig. 9).

[0145] As shown in step S1010, the at least one processor (e.g., such as processor 2 of Fig. 8) may extract expression levels of a target selection of differentially expressed genes (DEGs) (e.g., expression levels 110A, shown in Fig. 9) from the obtained target RNA sequencing data (e.g., target RNA sequencing data 10B, shown in Fig. 9), wherein said target selection of DEGs comprises genes selected from a group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3. Step S1010 may be carried out by DEG analysis module 110 (as described, e.g., with reference to Fig. 9).

[0146] As shown in step S1015, the at least one processor (e.g., such as processor 2 of Fig. 8) may infer a pre-trained machine-learning (ML)-based model (e.g., ML-based model 120, shown in Fig. 9) on the extracted expression levels (e.g., expression levels 110A, shown in Fig. 9), to determine the subject being at high risk of committing suicide (e.g., by estimating suicide risk 120A, as described with reference to Fig. 9). Step S 1015 may be carried out by ML-based model 120 (as described, e.g., with reference to Fig. 9).

[0147] As shown in step S1020, the at least one processor (e.g., such as processor 2 of Fig. 8) may generate a treatment plan (e.g., treatment plan 130A, described with reference to Fig. 9) for the target subject to mitigate the risk of committing suicide. Step S1020 may be carried out by treatment plan generation module 130 (as described, e.g., with reference to Fig. 9).

[0148] Referring now to Fig. 10B, a flow diagram is presented, depicting a method of suicide risk mitigation (e.g., for a subject suffering from a psychiatric disorder susceptible to having a suicide risk), according to some embodiments.

[0149] Since the determined biomarkers are highly dysregulated between ‘SUICIDE’ and ‘NON- SUICIDE’ groups of patients, as it was determined during research, it is further suggested that, in some embodiments, other methods, not involving ML may be applied.

[0150] As shown in step S2005, the method of suicide risk mitigation may include determining (e.g., using a processor, such as processor 2 of Fig. 8) expression levels (e.g., same as determining expression levels 110A, discussed with reference to Fig. 9) of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11- 459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

[0151] As further shown in step S2010, the method of suicide risk mitigation may further include comparing (e.g., using a processor, such as processor 2 of Fig. 8) the expression levels of at least one gene to a predetermined threshold, wherein expression beyond the predetermined threshold indicated of a suicide risk in said subject, thereby mitigating the suicide risk of said subject.

[0152] Referring now to Fig. 10C, a flow diagram is presented, depicting a method of treating a subject (e.g., a subject suffering from a psychiatric disorder susceptible to having a suicide risk), according to some embodiments of the present invention.

[0153] As shown in step S3005, the method of treating a subject may include determining (e.g., using a processor, such as processor 2 of Fig. 8) expression levels (e.g., same as determining expression levels 110A, discussed with reference to Fig. 9) of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

[0154] As further shown in step S3010, the method of treating a subject may further include comparing (e.g., using a processor, such as processor 2 of Fig. 8) the expression levels of at least one gene to a predetermined threshold, wherein expression beyond the predetermined threshold may be indicative of a suicide risk in said subject, thereby mitigating the suicide risk of said subject.

[0155] As shown in step S3015, the method of treating a subject may further include treating a subject indicated as having a suicide risk (e.g., treatment may be performedaccording to a new or adjusted treatment plan, e.g., as described in relation to treatment plan 130A, shown in Fig. 9).

[0156] In some embodiments, the methods of the invention are performed ex-vivo. In some embodiments, the diagnostic aspects of the methods of the invention are performed ex-vivo. It will be understood by a skilled artisan that all steps of the invention that include administering a therapeutic agent to a subject may require in-vivo action of the therapeutic.

[0157] As used herein, the terms “treatment” or “treating” of a disease, disorder, or condition encompasses alleviation of at least one symptom thereof, a reduction in the severity thereof, or inhibition of the progression thereof. Treatment need not mean that the disease, disorder, or condition is totally cured. To be an effective treatment, a useful composition herein needs only to reduce the severity of a disease, disorder, or condition, reduce the severity of symptoms associated therewith, or provide improvement to a patient or subject’s quality of life.

[0158] In some embodiments, a sample from the subject is provided. In some embodiments, the providing comprises withdrawing a sample from the subject. In some embodiments, the sample is a bodily fluid. Bodily fluids include for example, blood, plasma, urine, lymph, stool, saliva, semen, and breast milk. In some embodiments, the sample is any one of blood, urine and saliva. In some embodiments, the sample is a blood sample. In some embodiments, the sample is selected from blood plasma, whole blood, blood serum or peripheral blood mononuclear cells. In some embodiments, the sample is plasma. In some embodiments, the blood is peripheral blood. In some embodiments, the sample is lymphoblastoid cell line (LCL). In some embodiments, providing comprises drawing a blood sample. In some embodiments, the sample is processed before the measuring. In some embodiments, protein is extracted from the sample before the measuring. In some embodiments, nucleic acids are extracted from the sample before the measuring.

[0159] In some embodiments, the measuring occurs in the sample. Examples of such in situ measuring include for example by ELISA. In some embodiments, the measuring occurs in a composition comprising material extracted from the sample, such as by PCR with extracted nucleic acids. In some embodiments, the measuring expression comprises measuring mRNA expression. In some embodiments, the measuring expression comprises measuring protein expression. In some embodiments, the measuring expression comprises measuring mRNA and protein expression.

[0160] In some embodiments, the measuring is measuring expression of at least one molecule that regulates expression of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3. In some embodiments, the measuring is measuring expression of a plurality of molecules that regulate expression of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11- 459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3. In some embodiments, the measuring is measuring expression of at least 2, at least 3, at least 4, at least 5, at least 6, 7, or at least 8 molecules that regulates expression of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3. Each possibility represents a separate embodiment of the invention.

[0161] The nucleotide and amino acid sequences of these factors can be found in the Entrez gene and Uniprot database, for example, and each factor’s accession number is provided hereinbelow. The human LYL1 gene can be found under Entrez gene # 4066 and the protein data (e.g., amino acid sequences) can be found under Uniprot ID P12980. The human LM07 gene can be found under Entrez gene # 4008 and the protein data can be found under Uniprot ID Q8WWI1. The human DDAH2 gene can be found under Entrez gene # 23564 and the protein data can be found under Uniprot ID O95865.The human PAG1 gene can be found under Entrez gene # 55824 and the protein data can be found under Uniprot ID Q9NWQ8. The human LCK gene can be found under Entrez gene # 3932 and the protein data can be found under Uniprot ID P06239. The human AP001372.2 (also known as LIPT2) gene can be found under Entrez gene # 387787 and the protein data can be found under Uniprot ID A6NK58. The human INPP5D gene can be found under Entrez gene # 3635 and the protein data can be found under Uniprot ID Q92835. The human SNX8 gene can be found under Entrez gene # 29886 and the protein data can be found underUniprot ID Q9Y5X2. The human TBC1D4 gene can be found under Entrez gene # 9882 and the protein data can be found under Uniprot ID 060343. The human GJC1 gene can be found under Entrez gene # 10052 and the protein data can be found under Uniprot ID P36383. The human AP000347.2 gene (also known as RGL4 or ral guanine nucleotide dissociation stimulator like 4) can be found under Entrez gene # 266747 and the protein data can be found under Uniprot ID E7EPT8. The human CD302 gene can be found under Entrez gene # 9936 and the protein data can be found under Uniprot ID Q8IX05. The human RGS17 gene can be found under Entrez gene # 26575 and the protein data can be found under Uniprot ID Q9UGC6. The human TBC1D22B gene can be found under Entrez gene # 55633 and the protein data can be found under Uniprot ID Q9NU 19 (also known as TBC1 domain family member 22B). The human HCG11 gene can be found under NCBI reference sequence NR_026790.1. The human XPO6 gene can be found under Entrez gene # 23214 and the protein data can be found under Uniprot ID Q96QU8 (also known as Exportin-6). The human DMAC 1 gene can be found under Entrez gene # 90871 and the protein data can be found under Uniprot ID Q96GE9. The human S1PR2 gene can be found under Entrez gene # 9294 and the protein data can be found under Uniprot ID 095136 (also known as Sphingosine 1 -phosphate receptor 2). The human PPFIA3 gene can be found under Entrez gene # 8541 and the protein data can be found under Uniprot ID 075145 (also known as Liprin-alpha-3). The human SLC7A6 gene can be found under Entrez gene # 9057 and the protein data can be found under Uniprot ID Q92536 (also known as YLAT2 or Y+L amino acid transporter 2). The human LAPTM5 gene can be found under Entrez gene # 7805 and the protein data can be found under Uniprot ID Q13571 (also known as Lysosomal- associated transmembrane protein 5). The human AC020951.1 transcript can be found under Ensembl ID: ENST00000624361. The human RP11-459E5.1 transcript can be found Ensembl ID: ENSG00000253125. The human TNFRSF1A gene can be found under Entrez gene # 7132 and the protein data can be found under Uniprot ID P19438 (also known as Tumor necrosis factor receptor superfamily member 1A). The human AC026202.3 transcript can be found under Ensembl ID: ENST00000439325 and the protein data can be found under Uniprot ID F8WER9. The human HMGN 1P38 gene can be found under Entrez gene # 643790. The human LYST (Lysosomal-trafficking regulator) gene can be found under Entrez gene # 1130 and the protein data can be found under Uniprot ID Q99698. The human MAN2B2 gene can be found under Entrez gene # 23324 and the protein data can befound under Uniprot ID Q9Y2E5 (also known as Epididymis-specific alpha-mannosidase). The human LUZP1 gene can be found under Entrez gene # 7798 and the protein data can be found under Uniprot ID Q86V48 (also known as Leucine zipper protein 1). The human MSH3 gene can be found under Entrez gene # 4437 and the protein data can be found under Uniprot ID P20585.

[0162] In some embodiments, the threshold is an expression level above a predetermined expression level. In some embodiments, the threshold is a predetermined number of standard deviations above a baseline expression. In some embodiments, the baseline expression is the average expression in control subjects not in risk of suicide. In some embodiments, the baseline expression is the average expression in control subjects having low or no risk of suicide. In some embodiments, the threshold is 0.25, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, or 3 standard deviations above a baseline. Each possibility represents a separate embodiment of the invention.

[0163] In some embodiments, mRNA from the genes is measured. In some embodiments, the measuring comprises the step of obtaining nucleic acid molecules from the sample. In some embodiments, the nucleic acids molecules are selected from mRNA molecules, DNA molecules and cDNA molecules. In some embodiments, the cDNA molecules are obtained by reverse transcribing the mRNA molecules. In some embodiments, the expression is determined by measuring mRNA levels of the genes. Methods for mRNA extraction are well known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al., Current Protocols of Molecular Biology, John Wiley and Sons (1997). Methods for RNA extraction from paraffin embedded tissues are disclosed, for example, in Rupp and Locker, Lab Invest. 56:A67 (1987), and De Andres et al., BioTechniques 18:42044 (1995).

[0164] Further, methods, reagents, and assays for measuring expression levels of these factors are well known in the art and are commercially available.

[0165] Numerous methods are known in the art for measuring expression levels of a one or more gene such as by amplification of nucleic acids (e.g., PCR, isothermal methods, rolling circle methods, etc.) or by quantitative in situ hybridization. Design of primers for amplification of specific genes is well known in the art, and such primers can be found or designed on various websites such as http: / / bioinfo.ut.ee / primer3-0.4.0 / or https: / / pga.mgh.harvard.edu / primerbank / for example.

[0166] The skilled artisan will understand that these methods may be used alone or combined. Non-limiting exemplary method are described herein.

[0167] RT-qPCR: A common technology used for measuring RNA abundance is RT- qPCR where reverse transcription (RT) is followed by real-time quantitative PCR (qPCR). Reverse transcription first generates a DNA template from the RNA. This single- stranded template is called cDNA. The cDNA template is then amplified in the quantitative step, during which the fluorescence emitted by labeled hybridization probes or intercalating dyes changes as the DNA amplification process progresses. Quantitative PCR produces a measurement of an increase or decrease in copies of the original RNA and has been used to attempt to define changes of gene expression in cancer tissue as compared to comparable healthy tissues.

[0168] RNA-Seq: RNA-Seq uses recently developed deep- sequencing technologies. In general, a population of RNA (total or fractionated, such as poly(A)+) is converted to a library of cDNA fragments with adaptors attached to one or both ends. Each molecule, with or without amplification, is then sequenced in a high-throughput manner to obtain short sequences from one end (single-end sequencing) or both ends (pair-end sequencing). The reads are typically 30-400 bp, depending on the DNA-sequencing technology used. In principle, any high-throughput sequencing technology can be used for RNA-Seq. Following sequencing, the resulting reads are either aligned to a reference genome or reference transcripts or assembled de novo without the genomic sequence to produce a genome-scale transcription map that consists of both the transcriptional structure and / or level of expression for each gene. To avoid artifacts and biases generated by reverse transcription direct RNA sequencing can also be applied.

[0169] Microarray: Expression levels of a gene may be assessed using the microarray technique. In this method, polynucleotide sequences of interest (including cDNAs and oligonucleotides) are arrayed on a substrate. The arrayed sequences are then contacted under conditions suitable for specific hybridization with detectably labeled cDNA generated from RNA of a test sample. As in the RT-PCR method, the source of RNA typically is total RNA isolated from a tumor sample, and optionally from normal tissue of the same patient as an internal control or cell lines. RNA can be extracted, for example, from frozen or archived paraffin-embedded and fixed (e.g., formalin-fixed) tissue samples. For archived, formalin- fixed tissue cDNA-mediated annealing, selection, extension, and ligation, DASL-Hluminamethod may be used. For a non-limiting example, PCR amplified cDNAs to be assayed are applied to a substrate in a dense array. Microarray analysis can be performed by commercially available equipment, following manufacturer's protocols, such as by using the Affymetrix GenChip technology, or Incyte's microarray technology.

[0170] In some embodiments, protein expression from the genes is measured. In some embodiments, the expression, and the level of expression, of proteins or polypeptides of interest can be detected through immunohistochemical staining of tissue slices or sections. Additionally, proteins / polypeptides of interest may be detected by Western blotting, ELISA or Radioimmunoassay (RIA) assays employing protein- specific antibodies.

[0171] Alternatively, protein levels can be determined by constructing an antibody microarray in which binding sites comprise immobilized, preferably monoclonal, antibodies specific to a plurality of proteins of interest. Methods for making monoclonal antibodies are well known (see, e.g., Harlow and Lane, 1988, Antibodies: a laboratory manual, Cold Spring Harbor, N.Y., which is incorporated in its entirety for all purposes). In one embodiment, monoclonal antibodies are raised against synthetic peptide fragments designed based on genomic sequence of the cell. With such an antibody array, proteins from the cell are contacted to the array, and their binding is assayed with assays known in the art.

[0172] In some embodiments, other clinical characteristics are considered in determining suitability. A skilled artisan will be aware that when determining drug suitability, a physician may consider a subject’s full medical history. Such other clinical characteristics include, but are not limited to, age, BMI, albumin levels, complete blood counts (CBC), and C-reactive protein results.

[0173] As used herein, the terms “administering,” “administration,” and like terms refer to any method which, in sound medical practice, delivers a composition containing an active agent to a subject in such a manner as to provide a therapeutic effect. One aspect of the present subject matter provides for intravenous administration of the therapeutic agent to a patient determined to be suitable for treatment. One aspect of the present subject matter provides for rectal administration of the therapeutic agent to a patient determined to be suitable for treatment. Other suitable routes of administration can include parenteral, subcutaneous, oral, rectal, intramuscular, enema, intra-intestinal or intraperitoneal.

[0174] The dosage administered will be dependent upon the age, health, and weight of the recipient, kind of concurrent treatment, if any, frequency of treatment, and the nature of the effect desired.

[0175] In another aspect, there is provided a kit comprising at least two detection molecules selected from detection molecule specific to LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

[0176] In some embodiments, the kit is for use in performing a method of the invention. In some embodiments, the kit is for treatment of a subject suffering from a psychiatric disorder susceptible to having a suicide risk. In some embodiments, the kit is for suicide risk mitigation for a subject suffering from a psychiatric disorder. In some embodiments, the kit is for determining suicide risk and treatment of a subject suffering from a psychiatric disorder.

[0177] In some embodiments, the kit comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 detection molecules. Each possibility represents a separate embodiment of the invention. In some embodiments, the kit comprises no more than 8, 9, 10, 15, 20, 25, 50, 75, 100, 200, 300, 500 or 100 detection molecules. Each possibility represents a separate embodiment of the invention. In some embodiments, the detection molecule detects protein. In some embodiments, the detection molecule detects RNA. In some embodiments, the detection molecule is an antibody. In some embodiments, the detection molecule is a hybridization probe. In some embodiments, the detection molecule is a pair of primers. In some embodiments, the primers are PCR primers. In some embodiments, the detection molecule is a nucleic acid sequence complementary to an mRNA that encodes the target protein.EXAMPLES

[0178] Generally, the nomenclature used herein, and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al.,"Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.

[0179] As can be seen in the provided description, the present invention represents methods and systems for predicting and mitigating suicide risk in subjects, providing solutions that leverage gene expression data to identify objective biomarkers associated with increased suicide risk, thereby enhancing the accuracy and reliability of suicide risk assessment. The present invention further provides systems and methods that utilize machine learning techniques to analyze genomic data for the purposes of suicide risk prediction and mitigation, providing an improvement of the field of psychiatric care by enabling more timely and effective interventions, as well as more personalized treatment planning. These advancements may provide clinicians with tools to identify high-risk patients who might not be detected through traditional assessment methods, potentially reducing suicide rates and improving overall patient outcomes in individuals with mental health disorders. The present invention further provides a kit for determining a subject being at high risk of committing suicide, as well as methods for treating subjects having high risk of committing suicide.

[0180] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Furthermore, all formulas described herein are intended as examples only and other or different formulas may be used. Additionally, some of the described method embodiments or elements thereof may occur or be performed at the same point in time.

[0181] While certain features of the invention have been illustrated and described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention.

[0182] Various embodiments have been presented. Each of these embodiments may of course include features from other embodiments presented, and embodiments not specifically described may include various features described herein.

Claims

CLAIMS1. A method of suicide risk mitigation, the method comprising: obtaining target RNA sequencing data from a biological sample of a target subject; extracting expression levels of a target selection of differentially expressed genes (DEGs) from the obtained target RNA sequencing data, inferring a pre-trained machine-learning (ML)-based model on the extracted expression levels, to determine the subject being at high risk of committing suicide; and generating a treatment plan for the target subject to mitigate the risk of committing suicide; wherein said target selection of DEGs comprises genes selected from a group consisting of:LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, EYST, MAN2B2, EUZP1, and MSH3.

2. The method of claim 1, wherein said target selection of DEGs comprises at least eight genes selected from the group consisting of:EYE1, EM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

3. The method of claim 2, wherein said target selection of DEGs comprises at least eight genes selected from the group consisting of:LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, and GJCL4. The method according to any one of claims 1-3, wherein said pre-trained ML- based model is trained via supervised training, based on a training dataset comprising: a first subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samplesof a sub-group of subjects that have never attempted suicide, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have one or more suicide attempts, said training samples of the second subset being labelled as indicative of a high risk of committing suicide.

5. The method according to any one of claims 1-4, further comprising: selecting, from a group of subjects, a first sub-group of subjects that have never attempted suicide, and a second sub-group of subjects that have one or more suicide attempts; obtaining training RNA sequencing data from biological samples of the subjects of the first and second sub-groups; extracting expression levels of the target selection of DEGs from the obtained training RNA sequencing data, performing supervised training of the ML-based model based on a training dataset comprising: a first subset of training samples comprising the expression levels of the subjects of the first sub-group, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising the expression levels of the subjects of the second sub-group, said training samples of the second subset being labelled as indicative of a high risk of committing suicide.

6. The method according to any one of claims 1-5, further comprising: selecting, from a group of subjects, a first sub-group of subjects that have never attempted suicide, and a second sub-group of subjects that have one or more suicide attempts; obtaining training RNA sequencing data from biological samples of the subjects of the first and second sub-groups; extracting expression levels for an initial pool of DEGs from the obtained training RNA sequencing data,from the initial pool of DEGs, identifying an initial selection of DEGs providing statistically significant representation of the subjects of the first sub-group with respect to the subjects of the second sub-group, based on p-value and fold change metrics; performing supervised training of the ML -based model to classify subjects as pertaining to the first and the second sub-groups, using a leave-one-out cross validation (LOOCV) algorithm, while concurrently assessing, at each iteration of the LOOCV algorithm, a prediction accuracy of various intermediate sub- selections of DEGs, each representing a distinct combination of DEGs selected from the initial selection of the DEGs; and determining the target selection of DEGs, based on the assessed prediction accuracy.

7. The method of claim 6, wherein each iteration of the LOOCV algorithm comprises:(a) generating a plurality of said intermediate sub- selections of DEGs;(b) forming a plurality of training datasets, each comprising: a first subset of training samples comprising the expression levels of a respective of said plurality of intermediate selections of DEGs that are associated with the subjects of the first sub-group; and a second subset of training samples comprising the expression levels of the respective of said plurality of intermediate selections of DEGs that are associated with the subjects of the second sub-group; said training samples of the first and second subsets being labelled by an association with the first and second sub-groups of subjects, respectively;(c) performing a supervised training of a plurality of instances of the ML -based model to classify subjects as pertaining to the first and the second sub-groups, based on said plurality of the training datasets, respectively;(d) assessing performance of the plurality of instances of the ML-based model, to identify a sub-plurality of intermediate sub- selections of DEGs from the plurality of the intermediate sub-selections of DEGs providing a highest prediction accuracy between the plurality of instances of the ML-based model;(e) determining a list of DEGs that most frequently appear across the intermediate sub- selections of DEGs of the identified sub-plurality; andwherein said determining the target selection of DEGs is performed by selecting DEGs that most frequently appear across lists of DEGs obtained at different iterations of the LOOCV algorithm.

8. The method according to any one of claims 6-7, wherein said identifying the initial selection of DEGs is performed by selecting, from the initial pool of DEGs, DEGs with a p-value less than 0.05 and the fold change metrics with log2 fold-change greater than or equal to |0.58|.

9. The method according to any one of claims 1-8, wherein said biological sample is a peripheral blood sample.

10. The method according to any one of claims 1-9, wherein said extracting expression levels comprises measuring mRNA expression, protein expression or both.

11. The method according to any one of claims 1-10, wherein said biological sample of the target subject comprises lymphoblastoid cell line (LCL) derived from peripheral blood mononuclear cells (PBMCs) of the target subject.

12. The method according to any one of claims 1-11, wherein said pre-trained ML- based model is selected from a group consisting of: Logistic Regression (LR) model; Random Forest (RF) model; K-Nearest Neighbors (kNN) model, Support Vector Machine (SVM) model, and Neural Network (NN) model.

13. A system for suicide risk mitigation, the system comprising: at least one non- transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with said at least one memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to: receive target RNA sequencing data obtained from a biological sample of a target subject suffering from the psychiatric disorder; extract expression levels of a target selection of differentially expressed genes (DEGs) from the obtained target RNA sequencing data,infer a pre-trained machine-learning (ML) -based model on the extracted expression levels, to determine the target subject being at high risk of committing suicide; and generate a treatment plan for the target subject, so as to mitigate the risk of committing suicide; wherein said target selection of DEGs comprises genes selected from a group consisting of:LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

14. The system of claim 13, wherein said target selection of DEGs comprises at least eight genes selected from the group consisting of:LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

15. The system of claim 14, wherein said target selection of DEGs comprises at least eight genes selected from the group consisting of:LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, and GJC1.

16. The system according to any one of claims 13-15, wherein said pre-trained ML- based model is trained via supervised training, based on a training dataset comprising: a first subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have never attempted suicide, said training samples of the first subset being labelled as indicative of a low risk of committing suicide; and a second subset of training samples comprising expression levels of the target selection of DEGs extracted from RNA sequencing data obtained from biological samples of a sub-group of subjects that have one or more suicide attempts, said training samples of the second subset being labelled as indicative of a high risk of committing suicide.

17. The system according to any one of claims 13-16, wherein said biological sample of the target subject comprises lymphoblastoid cell line (LCL) derived from peripheral blood mononuclear cells (PBMCs) of the target subject.

18. The system according to any one of claims 13-17, wherein said sample is a peripheral blood sample.

19. The system according to any one of claims 13-18, wherein said at least one processor is configured to extract said expression levels by measuring mRNA expression, protein expression or both.

20. The system according to any one of claims 13-19, wherein said pre-trained ML- based model is selected from a group consisting of: Logistic Regression (LR) model; Random Forest (RF) model; K-Nearest Neighbors (kNN) model, Support Vector Machine (SVM) model, and Neural Network (NN) model.

21. A method of suicide risk mitigation for a subject, the method comprising: determining expression levels of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3; and comparing the expression levels of at least one gene to a predetermined threshold; wherein expression beyond the predetermined threshold indicated of a suicide risk in said subject, thereby mitigating the suicide risk of said subject.

22. A method of treating a subject, the method comprising: determining expression levels of at least one gene selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XPO6, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3; and comparing the expression levels of at least one gene to a predetermined threshold; wherein expression beyond the predetermined threshold indicated of a suicide risk in said subject; andtreating a subject indicated as having a suicide risk.

23. The method of claim 22, wherein said determining expression levels comprises measuring mRNA expression, protein expression or both.

24. The method according to any one of claims 22-23, comprising measuring in a target sample expression of at least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

25. The method according to any one of claims 22-24, comprising measuring in a target sample expression of at least eight genes selected from the group consisting of: LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, and GJC1.

26. A kit comprising at least two detection molecules selected from detection molecule specific to LYL1, LM07, DDAH2, PAG1, LCK, AP001372.2, INPP5D, SNX8, TBC1D4, GJC1, AP000347.2, CD302, RGS17, TBC1D22B, HCG11, XP06, DMAC1, S1PR2, PPFIA3, SLC7A6, LAPTM5, AC020951.1, RP11-459E5.1, TNFRSF1A, AC026202.3, HMGN1P38, LYST, MAN2B2, LUZP1, and MSH3.

Citation Information

Patent Citations

  • Method and kit for determining benefit of chemotherapy

    US20220290251A1

  • Methods of predicting the risk of having lower-extremity artery disease in patients suffering from type 2 diabetes

    WO2019238554A1

  • Determination of epigenetic modifications by nanopore sequencing

    WO2019239218A2