Device for generating an artificial intelligence system of prediction of an effect of a drug on a patient, device for predicting an effect of a drug on a patient
An AI system using deep learning and computational biology predicts drug effects on patients by integrating molecular and clinical data, addressing discrepancies in drug development and enhancing prediction accuracy for personalized treatments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SCIENTA LAB
- Filing Date
- 2026-01-08
- Publication Date
- 2026-07-16
AI Technical Summary
Current drug development methods face discrepancies between preclinical studies and clinical trial results, are costly and time-consuming, and fail to account for diverse genetic backgrounds and individual variations among patients, leading to inefficiencies in drug development.
A device using an artificial intelligence system trained with deep learning models and computational biology to predict drug effects on patients by integrating molecular-level data with clinical outcomes, incorporating patient-specific gene expression data to provide personalized drug response predictions.
The system enhances the accuracy of drug performance predictions in clinical settings, reduces failure rates in clinical trials, and streamlines the drug development process by providing more reliable predictions before costly trials, allowing for better resource allocation and personalized treatment approaches.
Smart Images

Figure EP2026050323_16072026_PF_FP_ABST
Abstract
Description
DEVICE FOR GENERATING AN ARTIFICIAL INTELLIGENCE SYSTEM OF PREDICTION OF AN EFFECT OF A DRUG ON A PATIENT, DEVICE FOR PREDICTING AN EFFECT OF A DRUG ON A PATIENTFIELD OF INVENTION
[0001] The present invention relates to the field of computer-assisted drug effect prediction.
[0002] More precisely, the invention concerns a device for generating an artificial intelligence system of prediction of an effect of a drug on a patient, device for predicting an effect of a drug on a patient and a computer-implemented method for generating an artificial intelligence system of prediction of an effect of a drug on a patient.STATE OF THE ART
[0003] The present disclosure addresses several critical challenges in drug development and personalized medicine.
[0004] A first difficulty met in the field of drug development lies in the existing discrepancies between preclinical studies and clinical trial results, that render the assessment of the effect of drugs, as well as decision-making about which drug candidates to advance in the development pipeline challenging. Besides, clinical trials are costly and time-consuming.
[0005] In addition, current development methods often fail to adequately account for the diverse genetic backgrounds and individual variations among patients.
[0006] One goal of the invention is to improve the situation.SUMMARY OF THE INVENTION
[0007] A first aspect of this invention thus relates to a device for generating an artificial intelligence system of prediction of an effect of a drug on a patient, said device comprising:at least one input configured to receive:• a first training dataset comprising a first training sub-dataset,said first training sub-dataset comprising a plurality of first training samples, each first training sample being associated to a subject from a first plurality of subjects and a perturbation from a plurality of perturbations; wherein each first training sample comprises (i.e., for each subject in said first plurality of subjects) first molecular profiling data, comprising at least first transcriptomic profiling data, measured from a biological sample of the subject before application of the perturbation, and second molecular profiling data, measured after application of said perturbation on said biological sample;• a second training dataset comprising a plurality of second training samples associated to a second plurality of subjects, each second training sample comprising (i.e., for one subject of said second plurality of subjects) subject biological data and clinical scores associated to the subject; wherein said subject biological data comprises at least second molecular profiling data, comprising at least second transcriptomic profiling data, acquired from a biological sample of said subject;• an initial architecture of said artificial intelligence system, said initial architecture comprising:o a first sub-architecture configured to receive as input molecular profiling data from a biological sample of a patient comprising at least input transcriptomic profiling data and at least one input perturbation associated to at least one target gene, and to output a post-perturbation molecular profiling data prediction;o a second sub -architecture configured to receive said post-perturbation molecular profiling data prediction and to output at least one clinical score; at least one processor configured to generate said artificial intelligence system of prediction via a first training process of said first sub-architecture, comprising afirst phase using (i.e., based on) said first training sub-dataset and a second training process of said second sub-architecture using said second training dataset; at least one output configured to provide as output said artificial intelligence system for prediction.
[0008] One goal of this invention is to develop a more accurate, efficient, and personalized approach to predicting drug performance in clinical settings.
[0009] Advantageously, the trained artificial intelligence system 31 allows predicting an effect of a drug on a specific patient by leveraging deep learning models and computational biology models. The trained artificial intelligence system solves the problem of discrepancies between preclinical studies and clinical trial results.
[0010] By modeling drug effects at the gene level and propagating these effects to predict clinical outcomes, it provides a more accurate representation of how a drug candidate will perform in real-world patient and patient populations.
[0011] In addition, on the contrary of current methods that often fail to adequately account for the diverse genetic backgrounds and individual variations among patients, the trained artificial intelligence system addresses this limitation by incorporating patientspecific gene expression data, and as such, patient heterogeneity, to predict personalized drug responses. In other words, the solution of the present disclosure as potential for personalized outcome predictions based on individual gene expression profiles.
[0012] The solution proposed in the present disclosure bridges the gap between molecular level changes and patient-observable effects.
[0013] Besides, the trained artificial intelligence system aims at enhancing the accuracy of predicting drug performance in clinical settings. By integrating molecular-level data with clinical outcome predictions, it offers a more comprehensive and precise assessment of drug efficacy.
[0014] By providing more reliable predictions of drug performance before costly clinical trials, the trained artificial intelligence system according to the invention can potentially reduce the time and resources spent on candidates unlikely to succeed, therebystreamlining the drug development process. In other words, the artificial intelligence system according to the invention offers a more comprehensive assessment of drug candidates before they enter costly clinical trials.
[0015] By offering a more robust method for predicting drug performance, the invention provides pharmaceutical companies with better tools for making informed decisions about which drug candidates to advance in the development pipeline.
[0016] Examples of expected benefits of the present disclosure include:- Reduced failure rates in clinical trials;- More efficient allocation of resources in drug development;- Improved patient outcomes through more personalized treatment approaches;- Accelerated development of effective therapeutics;- Potential cost savings in the drug development process.
[0017] In some embodiments, the first training dataset further comprises a second training sub-dataset,said second training sub-dataset, being associated to a predefined perturbation of a biological state, comprising a plurality of third training samples from a third plurality of subjects, each third training sample comprising first specific molecular profiling data, comprising at least first specific transcriptomic profiling data, and second specific molecular profiling data, comprising at least second specific transcriptomic profiling data, measured from a biological sample of a subject from said third plurality of subjects, before and after application of said predefined perturbation on said biological sample; wherein said first training process comprises a second phase of fine-tuning learning based on said second training sub-dataset.
[0018] Advantageously, the second phase of fine-tuning learning renders the artificial intelligence system more precise and focused on specific drugs modulating specific target genes.
[0019] The artificial intelligence system presents a modular design with two subarchitectures allowing for prediction of multiple clinical outcomes.
[0020] In some embodiments, input transcriptomic profiling data from said biological sample comprises a vector comprising values representative of an expression level of a set of genes of the corresponding biological sample before said input perturbation of a biological sample, said set of genes comprising said at least one target gene, and wherein said post-perturbation transcriptomic profiling data prediction comprises a vector comprising values representative of a synthetic expression level of said set of genes of said patient after said perturbation of a biological sample.
[0021] In some embodiments, said expression level comprises levels of RNA transcripts.
[0022] In some embodiments, said action of said drug is an activation of said at least one target gene or an inhibition of said at least one target gene.
[0023] In some embodiments, each perturbation of the plurality of perturbation, input perturbation or predefined perturbation, is associated to a one or more target gene, each perturbation is represented as a perturbation vector having a predefined number of components each associated to a gene, each component comprising a perturbation value comprised between a minimum value and a maximum value, said minimum value corresponding to a completely inhibited state of said corresponding gene, and said maximum value corresponding to an activated state of said corresponding gene, said activated state being beyond a predefined threshold. Advantageously, with this definition, the perturbations can be modeled in a flexible and simple manner.
[0024] In some embodiments, each perturbation of the plurality of perturbation, input perturbation or predefined perturbation is representative of a synthetic administration dose of said drug to a subject or a patient. By synthetic administration, it is meant a synthetic administration among an oral administration, an intravenous administration, or other routes of administration (sublingual, transdermal, inhalation).
[0025] In some embodiments, the at least one clinical score comprises at least one among: a synthetic activity level of a disease, a synthetic inflammation biomarker, a synthetic blood cell count.
[0026] Examples of scores of activity level of diseases that the artificial intelligence system can output comprise the DAS28 score (Disease Activity Score) in the case of rheumatoid arthritis, ESSDAI (EULAR Sjogren’s syndrome (SS) disease activity index) in the case of Sjogren’s disease (SjD), SLED Al index (Systemic Lupus Erythematosus Disease Activity Index) in the case of Lupus.
[0027] Examples of inflammation biomarkers include sedimentation rate, CRP.
[0028] Examples of blood cell counts comprise lymphocyte counts, leukocyte counts, platelet counts.
[0029] In some embodiments, each of said first sub -architecture and said second subarchitecture is a foundation model based on generative model.
[0030] In some embodiments, the first sub-architecture and / or said second subarchitecture is based on a Transformer architecture.
[0031] In some embodiments, each first training sample, each second training sample and each third training sample further comprises proteomic profiling data, genomic profiling, cell composition and / or metabolic profiling measured from the corresponding biological sample.
[0032] All possible combinations of the previously described embodiments are part of the present disclosure.
[0033] A second aspect of the invention pertains to a device for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene, by using the artificial intelligence system obtained with the device previously described and comprising:at least one input configured to receive:• input molecular profiling data of said patient comprising at least input transcriptomic profiling data from a biological sample of said patient;• a current perturbation of said at least one target gene, said current perturbation being representative of an action of said drug on said at least one target gene at a given time instant;at least one processor configured to:• provide, as input to said artificial intelligence system, said input molecular profiling data and said current perturbation,• obtain, as output, a patient post-perturbation molecular profiling data prediction and at least one clinical score, said at least one clinical score being representative of said effect of said drug on said patient at said given time instant,- at least one output configured to provide said patient post-perturbation molecular profiling data and at least one clinical score.
[0034] In some embodiments, the effect of said drug on said patient is determined during a predefined timeframe comprising at least two administrations of said drug, wherein:said at least one input is further configured to receive an absorption phase for said patient, said predefined timeframe, a predefined interval, a predefined administration interval representing the interval between two successive administrations;said at least one processor is configured to implement a clock on the base of said predefined interval and an autoregressive model, by iteratively feeding the postperturbation molecular profiling data prediction obtained at a current (ith) iteration as input molecular profiling data for the artificial intelligence model at a successive (i+lth) iteration together with a successive perturbation, wherein the successive perturbation is obtained by modifying the current perturbation using said absorption phase,wherein a number of iterations is determined on the base of the predefined timeframe and said predefined interval.
[0035] Advantageously, the trained artificial intelligence system according to the invention allows predicting the effect of drugs on patients over time, both short-term and long-term. By providing time-spanning predictions of drug performance before costly clinical trials, the trained artificial intelligence system according to the invention can potentially reduce the time and resources spent on candidates unlikely to succeed, therebystreamlining the drug development process. In other words, the artificial intelligence system according to the invention offers a more comprehensive assessment of drug candidates before they enter costly clinical trials.
[0036] A third aspect of the invention pertains to a computer-implemented method for generating an artificial intelligence system of prediction of an effect of a drug on a patient, said method comprising:receiving:• a first training dataset comprising a first training sub -dataset,said first training sub-dataset comprising a plurality of first training samples, each first training sample being associated to a subject from a first plurality of subjects and a perturbation from a plurality of perturbations; wherein each first training sample comprises (e.g. for each subject in said first plurality of subj ects) first molecular profiling data, comprising at least first transcriptomic profiling data, measured from a biological sample of said subject before application of said perturbation, and second molecular profiling data, measured after application of said perturbation on said biological sample; • a second training dataset comprising a plurality of second training samples associated to a second plurality of subjects, each second training sample comprising (e.g. for one subject of said second plurality of subjects) subject biological data and clinical scores associated to said subject; wherein said subject biological data comprises at least second molecular profiling data, comprising at least second transcriptomic profiling data, acquired from a biological sample of said subject;• an initial architecture of said artificial intelligence system, said initial architecture comprising:o a first sub-architecture configured to receive as input molecular profiling data from a biological sample of a patient comprising at least input transcriptomic profiling data and at least one input perturbation associated to at least one target gene, and to output a post-perturbation molecular profiling data prediction;o a second sub-architecture configured to receive said post-perturbation molecular profiling data prediction and to output at least one clinical score;generating said artificial intelligence model via a first training process of said first sub-model comprising a first phase based on said first training sub-dataset; providing as output said artificial intelligence system of prediction.
[0037] In addition, the disclosure relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the methods compliant with any of the above execution modes.
[0038] The present disclosure further pertains to a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the methods compliant with any of the above execution modes.
[0039] The present disclosure further relates to a non-transitory program storage device (i.e. computer-readable storage medium), readable by a computer, tangibly embodying a program of instructions executable by the computer to perform the computer-implemented methods, compliant with the present disclosure.
[0040] Such a non-transitory program storage device can be, without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor device, or any suitable combination of the foregoing. It is to be appreciated that the following, while providing more specific examples, is merely an illustrative and not exhaustive listing as readily appreciated by one of ordinary skill in the art: a portable computer diskette, a hard disk, a ROM, an EPROM (Erasable Programmable ROM) or a Flash memory, a portable CD-ROM (Compact-Disc ROM).DEFINITIONS
[0041] In the present invention, the following terms have the following meanings:
[0042] The terms “adapted” and “configured” are used in the present disclosure as broadly encompassing initial configuration, later adaptation or complementation of thepresent device, or any combination thereof alike, whether effected through material or software means (including firmware).
[0043] The term “processor” should not be construed to be restricted to hardware capable of executing software, and refers in a general way to a processing device, which can for example include a computer, a microprocessor, an integrated circuit, or a programmable logic device (PLD). The processor may also encompass one or more Graphics Processing Units (GPU), whether exploited for computer graphics and image processing or other functions. Additionally, the instructions and / or data enabling to perform associated and / or resulting functionalities may be stored on any processor-readable medium such as, e.g., an integrated circuit, a hard disk, a CD (Compact Disc), an optical disc such as a DVD (Digital Versatile Disc), a RAM (Random-Access Memory) or a ROM (Read-Only Memory). Instructions may be notably stored in hardware, software, firmware or in any combination thereof.
[0044] “Machine learning (ML)” designates in a traditional way computer algorithms improving automatically through experience, on the ground of training data enabling to adjust parameters of computer models through gap reductions between expected outputs extracted from the training data and evaluated outputs computed by the computer models.
[0045] A “control sample” refers to a biological sample serving as a reference or baseline to compare against experimental or treated samples.BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure l is a block diagram representing schematically a particular mode of a device 1 for generating an artificial intelligence system 31 of prediction of an effect of a drug on a patient.
[0047] Figure 2 is a block diagram representing schematically a particular mode of an architecture of an artificial intelligence system of prediction 31 of an effect of a drug on a patient.
[0048] Figure 3 is a flow chart showing successive steps executed with the device 1 for generating an artificial intelligence system 31 of prediction of an effect of a drug on a patient illustrated on Figure 1.
[0049] Figure 4 is a block diagram representing schematically a particular mode of a device 2 for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene.
[0050] Figure 5 is a flow chart showing successive steps executed with the device 2 for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene, illustrated on Figure 4.
[0051] Figure 6 is an example of results that can be obtained with the device 2 for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene, illustrated in Figure 5.
[0052] Figure 7 is a flow chart showing successive steps executed with the device 2 for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene illustrated in Figure 5 in a process of predicting an effect of a drug on a patient during a predefined timeframe.
[0053] Figure 8 is an example of results that can be obtained with the device for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene illustrated in Figure 5.ILLUSTRATIVE EMBODIMENTS
[0054] The present description illustrates the principles of the present disclosure. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the disclosure and are included within its scope.
[0055] All examples and conditional language recited herein are intended for educational purposes to aid the reader in understanding the principles of the disclosureand the concepts contributed by the inventor to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions.
[0056] Moreover, all statements herein reciting principles, aspects, and embodiments of the disclosure, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0057] Thus, for example, it will be appreciated by those skilled in the art that the block diagrams presented herein may represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, and the like represent various processes which may be substantially represented in computer readable media and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0058] The functions of the various elements shown in the figures may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which may be shared.
[0059] It should be understood that the elements shown in the figures may be implemented in various forms of hardware, software or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more appropriately programmed general -purpose devices, which may include a processor, memory and input / output interfaces.
[0060] The present disclosure will be described in reference to a particular functional embodiment of a device 1 for generating an artificial intelligence system of prediction 31 of an effect of a drug on a patient, as illustrated on Figure 1.
[0061] The device 1 is adapted to produce an artificial intelligence system 31 of prediction of an effect of a drug on a patient, which is configured to receive, as an input, molecular profiling data from a biological sample of a patient comprising at least inputtranscriptomic profiling data and at least one input perturbation associated to at least one target gene (i.e., the at least one gene targeted by said drug), and generate as output a post-perturbation molecular profiling data prediction and at least one clinical score.
[0062] Notably, the device 1 is configured to train an initial architecture 20 using a first training dataset 21 and a second training dataset 22 to obtain the artificial intelligence system 31. Said initial architecture 20 may be a foundation model previously trained on a vast amount of generic data (e.g., using freely available datasets), or a randomly initialized artificial intelligence architecture (i.e., untrained model).
[0063] As will be seen further, advantageously the trained artificial intelligence system 31 allows predicting an effect of a drug on a specific patient by leveraging deep learning models and computational biology models. The trained artificial intelligence system 31 solves the problem of discrepancies between preclinical studies and clinical trial results.
[0064] By modeling drug effects at the gene level and propagating these effects to predict clinical outcomes, it provides a more accurate representation of how a drug candidate will perform in real-world patient populations.
[0065] In addition, on the contrary of current methods that often fail to adequately account for the diverse genetic backgrounds and individual variations among patients, the trained artificial intelligence system 31 addresses this limitation by incorporating patientspecific gene expression data to predict personalized drug responses.
[0066] Besides, the trained artificial intelligence system 31 aims at enhancing the accuracy of predicting drug performance in clinical settings. By integrating molecular-level data with clinical outcome predictions, it offers a more comprehensive and precise assessment of drug efficacy.
[0067] Examples of drug classes of which an effect can be predicted by the trained artificial intelligence system 31 according to the invention include: TNF inhibitors, Inductors of cell death, Inhibitors of cell costimulation, Inhibitors of chemotaxis and homing, Inhibitors of inflammatory cytokines, Inhibitors of TLR / IL1 signaling, Inhibitors of JAK / STAT.
[0068] Those drug classes are meant to treat one of the diseases listed below: Rheumatoid arthritis, Psoriatic arthritis, Systemic lupus erythematosus (SLE), Atopic dermatitis, Psoriasis, Ulcerative colitis, Crohn’s disease, Spondylarthritis, Alopecia areata, Autoimmune vasculitis, Sjogren’s disease, Hi dradenitis suppurativa, Eosinophilic esophagitis, Polymyalgia rheumatica, Giant cell arteritis, Multiple sclerosis, Myasthenia gravis.
[0069] The input data for the device 1 may be the first training data set 21, the second training data set 22 and an initial untrained architecture 20 of the artificial intelligence system 31.
[0070] The device 1 for generating a trained artificial intelligence system 31 of prediction of an effect of a drug on a patient, (i.e. trained so as to set the artificial intelligence system parameters) is associated with a device 2, represented on Figure 4. The device 2 is a device for predicting an effect of a drug on a patient, the drug modulating an expression of at least one target gene, from inputted molecular profiling data of the patient comprising at least input transcriptomic profiling data from a biological sample of the patient and from an inputted current perturbation of the at least one target gene, by using the trained artificial intelligence system 31 obtained from the device 1, which will be subsequently described. The current perturbation is representative of an action of the drug on the at least one target gene at a given time instant.
[0071] Though the presently described devices 1 and 2 are versatile and provided with several functions that can be carried out alternatively or in any cumulative way, other implementations within the scope of the present disclosure include devices having only parts of the present functionalities.
[0072] Each of the devices 1 and 2 is advantageously an apparatus, or a physical part of an apparatus, designed, configured and / or adapted for performing the mentioned functions and produce the mentioned effects or results. In alternative implementations, any of the device 1 and the device 2 is embodied as a set of apparatus or physical parts of apparatus, whether grouped in a same machine or in different, possibly remote, machines. The device 1 and / or the device 2 may e.g. have functions distributed over a cloudinfrastructure and be available to users as a cloud-based service, or have remote functions accessible through an API.
[0073] The device 1 for training and the device 2 for predicting an effect of a drug on a patient may be integrated in a same apparatus or set of apparatus, and intended to same users. In other implementations, the structure of the device 2 may be completely independent of the structure of the device 1, and may be provided for other users. For example, the device 2 may comprise the trained artificial intelligence system 31 available to operators for prediction of an effect of a drug on a patient, wholly set from previous training effected upstream by other players with the device 1.
[0074] In what follows, the modules are to be understood as functional entities rather than material, physically distinct, components. They can consequently be embodied either as grouped together in a same tangible and concrete component, or distributed into several such components. Also, each of those modules is possibly itself shared between at least two physical components. In addition, the modules are implemented in hardware, software, firmware, or any mixed form thereof as well. They are preferably embodied within at least one processor of the device 1 or of the device 2.
[0075] The device 1 comprises a module 11 for receiving at least the first training dataset 21, the second training dataset 22 and the initial architecture 20 of the artificial intelligence system 31, stored in one or more local or remote database(s). The latter can take the form of storage resources available from any kind of appropriate storage means, which can be notably a RAM or an EEPROM (Electrically-Erasable Programmable Read-Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk).
[0076] As will be described in detail later, the initial architecture 20 comprises two blocks (i.e., first and second sub-architectures), with different inputs and outputs, and therefore at least two training datasets may be required, notably at least one for each block. In the present disclosure, the first training dataset 21 is configured to train the first sub-architecture 16 while the second training dataset 22 is configured to train the second sub-architecture 18.
[0077] More in details, in the present disclosure, the first training dataset 21 comprises at least one first training sub-dataset. The first training sub-dataset comprises a plurality of first training samples, each first training sample being associated to a subject from a first plurality of subjects and a perturbation from a plurality of perturbations.
[0078] Each perturbation of the plurality of perturbation is representative of an effect of an administration of a dose of said drug to a subject or a patient. In other words, said perturbations may each represent the effect of a drug. In one example, the effect of the drug is an activation of at least one target gene or an inhibition of at least one target gene. In an example, the perturbations are known from experiments, such as CRISPR experiments, where gene expressions are directly perturbed. In another example, in the case of drugs, gene perturbations associated with a specific drug can be obtained from literature studies.
[0079] Each first training sample is associated with one subject from which a biological sample has previously been collected. For instance, the biological sample may be collected from one subject and the biological sample itself may be exposed to the drug without administration of the drug directly to the subject. Alternatively, the drug can be administered to the subject and the biological sample can be further collected from the patient after the administration. Each first training sample may comprise a pair of transcriptomic profiling data representing the biological state of the biological sample before and after the administration of the perturbation. As the perturbations are allowed to be selected plurality of perturbations for this first training sub-dataset, so that the available amount of data is vast, the type of information contained is not specific to a particular type of perturbation or to one perturbation, otherwise said the first training subdataset is a generic perturbation dataset.
[0080] Advantageously, the biological samples used in the first training dataset are derived directly from subjects, ensuring that the data reflects real -world perturbations that the subjects’ cells may actually experienced. The use of samples of cells obtained from subjects, instead of cell linage of the shell, allows to enhance the accuracy of the predictions, as it accounts for the unique genetic and molecular profiles of the subjects. Unlike generic biological samples, which do not capture the full complexity of patientsvariability, these biological samples provide a more precise representation of how perturbations affect real subject cells. In one embodiment, the first training sample may comprise first molecular profiling data measured from a biological sample of the subject before application of the perturbation (i.e., administration of said drug causing the perturbation) and second molecular profiling data, measured after application of the perturbation on the biological sample. In other words, a biological sample of one subject before application of the perturbation is a biological sample where no perturbation had been applied, therefore in its biological state. The biological sample from which are measured the molecular profiling data “before” and after the perturbation may not be the same biological sample (e.g., same ensemble of cells) by only be two biological samples of one same subject. For instance, a biological sample before and after perturbation can correspond to a cell originating from a same source cell, on which a perturbation has been applied (for instance, a drug has been administered). In one example, the first and second molecular profiling data comprise respectively at least first and second transcriptomic profiling data. In some examples, each first and second molecular profiling data further comprises at least one of the following: proteomic profiling data, genomic profiling, cell composition, metabolic profiling, information relative to types and corresponding levels of epigenetics alterations or measured from the corresponding biological sample. From the present disclosure, it has to be understood that, all collection and measurements of the biological samples have been performed previously to the implementation of the methods and devices of the present invention.
[0081] As mentioned before, each perturbation of the plurality of perturbation may be associated to one or more target genes (i.e., one drug). Mathematically, each of these perturbations may be represented as a perturbation vector having a predefined number of components each associated to a gene of a panel of genes (or set of genes). Said panel of genes comprises both said one or more target genes than other genes of the biological sample. Each component of said perturbation vector may comprise a perturbation value comprised between a minimum value and a maximum value, said minimum value corresponding to a completely inhibited state of said corresponding gene, and said maximum value corresponding to an activated state of said corresponding gene. In oneexample, the activated state is represented by the value of the corresponding component exciding a predefined activation threshold.
[0082] For instance, the at least first transcriptomic profiling data can be single-cell RNA-seq data gathered from transcriptomic cell atlases, or bulk RNA-seq data gathered from IMDs cohorts.
[0083] Examples of other molecular profiling data comprised in the molecular profiling data include:- epigenetics data obtained with ATAC-seq or methylation levels measurements, - genomics data obtained by DNA Next-Generation sequencing,- proteomics data obtained by mass spectrometry, protein micro-array of affinity-based methods,- cell compositions obtained by flow cytometry,- metabolomic data obtained by mass spectrometry or Nuclear Magnetic Resonance spectroscopy.
[0084] In some embodiments, the first training dataset 21 only comprises the first subdataset. For instance, the first sub-dataset may comprise data obtained by the PerturbSeq method. Data obtained from this method may be found in public databases such as GSE90546 or GSE146194, available on the Gene Expression Omnibus portal (GEO) and may be included in the first sub-dataset. Those databases comprise sequencing data obtained with the RNA-seq method. The sequencing data corresponds to, on one hand, control cells that did not undergo any perturbation (i.e., biological samples before application of a perturbation) and on the other hand, perturbed cells (i.e., biological sample after application of said perturbation) of which one or several genes were perturbed by the CRISPR method. Typically, the first sub-dataset is generated to comprises at least one million of sequencing data corresponding to perturbed cells. Sequencing data obtained from control cells, corresponding to a first biological sample, on one hand, and sequencing data obtained from perturbed cells, corresponding to a second biological sample different from the first biological sample but belonging to a same subject, can be paired.
[0085] In some embodiments, the first training dataset 21 further comprises a second training sub-dataset. The second training sub-dataset is associated to a predefined perturbation (i.e., one predefined drug) and comprises a plurality of third training samples from a third plurality of subjects. Said third plurality of subjects may be comprised in the first plurality of subjects. Each third training sample comprises a pair constituted of first specific molecular profiling data and second specific molecular profiling data. The first specific molecular profiling data and the second specific molecular profiling data are molecular profiling data previously measured from the biological sample(s) of one subject from the third plurality of subjects, respectively before and after application of the predefined perturbation on the biological sample. Alternatively, the first set of specific molecular profiling data is obtained from a first biological sample that was not perturbed, while the second set of molecular profiling data is obtained from a second biological sample, distinct from the first, that was subjected to perturbation, (e.g., the first and second biological sample being obtained from the same subject). As for the first training subset, in one example, the first and second specific molecular profiling data comprise respectively at least first and second specific transcriptomic profiling data. In some examples, each first specific and second specific molecular profiling data further comprises at least one of the following: proteomic profiling data, genomic profiling, cell composition, metabolic profiling, information relative to types and corresponding levels of epigenetics alterations or measured from the corresponding biological sample.
[0086] For instance, the second training sub-dataset may comprise data obtained with the DrugSeq method, which may be retrieved also on the Gene Expression Omnibus portal (GEO). The DrugSeq method aims at measuring evolution of the RNA data of a biological sample after administration of a specific treatment to this biological sample. The second training sub-dataset consists of two types of data: molecular profiling data measured from control samples that did not receive any treatment, and data from samples that were exposed to the specific treatment. It will be further seen that, in those embodiments, the training of the architecture 20 will include a pre-training phase based on the first training sub-dataset, followed by a fine-tuning phase based on the second training sub-dataset.
[0087] Therefore, according to this example, the first training dataset 21 comprises both data that may be obtained with the Perturb Seq method and with the DrugSeq method.
[0088] In the present disclosure, the second training dataset 22 comprises a plurality of second training samples associated to a second plurality of subjects. Said second plurality of subject may be identical or not to the first plurality of patients. Each second training sample comprises, for one subject of the second plurality of subjects, subject biological data and clinical scores associated to said subject. The subject biological data of each second training sample may comprise at least second molecular profiling data, comprising at least second transcriptomic profiling data, acquired from a biological sample of the subject. In some examples, the second molecular profiling data of the second training dataset 22 further comprises at least one of the following: proteomic profiling data, genomic profiling, cell composition, metabolic profiling, information relative to types and corresponding levels of epigenetics alterations or measured from the corresponding biological sample. The second training dataset 22 comprises biological samples derived directly from corresponding patients (i.e., second plurality of subjects), ensuring that the clinical scores are accurately associated with the biological data. As for the first training dataset, obtaining data from actual subjects provide more complete information than using cell linage which allows to obtain a trained model with enhanced accuracy of predictions.
[0089] The clinical scores of the second training dataset 22 may be one of the following information measures from the corresponding subject: an activity level of a disease, an inflammation biomarker, a blood cell count. Depending on the disease the drug is meant to treat, the inflammation biomarker measure from one subject may vary. The biomarkers measured may include, but are not limited to, C-Reactive Protein (CRP) and Erythrocyte Sedimentation Rate (ESR) for conditions such as Rheumatoid Arthritis (RA), Psoriatic Arthritis (PsA), and Spondylarthritis (SpA). Additional biomarkers include Anti-Nuclear Antibodies (ANA), Anti-dsDNA, and complement proteins (C3, C4) for Systemic Lupus Erythematosus (SLE); fecal calprotectin and Interleukin (IL)-6 for Ulcerative Colitis and Crohn’s Disease; IL- 17 and IL-23 for Psoriasis and Psoriatic Arthritis; and AntiNeutrophil Cytoplasmic Antibodies (ANCA) for autoimmune vasculitis. Otherbiomarkers may be the neurofilament light chain (NfL) for Multiple Sclerosis (MS), Anti-Acetylcholine Receptor (AChR) antibodies for Myasthenia Gravis, and eosinophil counts for Eosinophilic Esophagitis. More in detail, the clinical score is a quantitative indicator of a patient’s overall health status. It provides a global and systemic assessment of the patient that reflects the patient’s condition as a whole. As such, the clinical score captures aspects of health that are not localized to a single tissue or biological sample. Importantly, the clinical score can only be determined through direct evaluation of the patient (for example, via clinical examination, functional test, and / or the like) and cannot be inferred from an isolated biological sample. For instance, a patient biopsy represents only a localized snapshot of a specific tissue and does not reflect the patient’s systemic health state. Consequently, the clinical score as defined in the present disclosure cannot be derived from a biopsy or any other patient sample.
[0090] In some embodiments, the second training dataset 22 comprises a diseasespecific multimodal dataset, combining bulk RNA-seq data or single-cell RNA-seq data with molecular profiling data and associated patient clinical scores of interest. The second training dataset 22 comprises at least 100 second training samples of paired biological data and clinical scores data.
[0091] In some embodiments, as illustrated on Figure 2, the architecture 20 to be trained in the device 1, comprises a first sub-architecture 16 and a second sub-architecture 18.
[0092] More in detail, the first sub -architecture 16 is configured to receive as input molecular profiling data (previously obtained) from a biological sample of a patient 32 and at least one input perturbation 33 associated to at least one target gene, and to output a post-perturbation molecular profiling data prediction 41. The post-perturbation molecular profiling data prediction 41 comprises a prediction of the most likely postperturbation transcriptomic profiling data following the application of the input perturbation. In other words, the first sub-architecture 16 is configured to predict the effect of a drug modulating the expression of at least one target gene at the biological level. The first sub-architecture 16, by receiving molecular profiling data and a perturbation 33 associated with at least one target gene, predicts post-perturbation gene expression while considering the interrelations between genes within regulatorynetworks. In this model, the effect of a perturbation on a target gene is represented as propagating to other connected genes through these networks. This approach allows for a comprehensive modeling of how genetic interactions influence the expression response to the perturbation, thereby enhancing the accuracy of the predictions.
[0093] According to one embodiment, the input molecular profiling data comprises at least one input transcriptomic profiling data. For instance, the at least input transcriptomic profiling data can be bulk RNA expression data of subjects, or single-cell RNA expression data of subjects.
[0094] The second sub-architecture 18 is configured to receive the post-perturbation molecular profiling data prediction 41 and to output at least one clinical score. In other words, based on the post-perturbation molecular profile data prediction, that can include post-perturbation bulk RNA, clinical scores of interest can be predicted to assess drug effect at the clinical level. In brief, the second sub-architecture 18 advantageously allows to map gene expression changes to clinically relevant metrics.
[0095] In some embodiments, the at least one clinical score comprises at least one among: a synthetic activity level of a disease, a synthetic inflammation biomarker, a synthetic blood cell count. As described above, the clinical score represents the global and systemic assessment of the patient’s health.
[0096] In some embodiments, each of the first sub-architecture 16 and the second subarchitecture 18 is a foundation model based on a generative model.
[0097] In some embodiments, the first sub-architecture 16 and / or the second subarchitecture 18 is based on a Transformer architecture.
[0098] For example, the first sub-architecture 16 is a Transformer model with twelve attention layers, each one with eight attention heads and a hidden embedding size of 512. The initial embedding size is 512.
[0099] In this same example, the second sub -architecture 18 consists of a sequence of a Transformer model and a multi-layer perceptron layer. The Transformer model comprises 12 attention layers, each one with 8 attention heads, and a hidden embedding size of 512.The initial embedding size is 512. The multi-layer perceptron model consists of 3 layers of size 768, 100 and 50. In those embodiments, the second sub -architecture 18 is configured to receive as input a vector comprising synthetic gene expression levels and outputs at least one clinical score.
[0100] In the previous example where the first sub-architecture is a Transformer model and the second sub-architecture is a sequence of a Transformer model and a multi-layer perceptron layer, the input data of the first sub-architecture are two vectors whose components are respectively:- gene expression levels associated to (e.g. acquired from) a biological sample before that any perturbation is applied;- (gene) perturbation values.
[0101] In the previous example where the first sub-architecture is a Transformer model and the second sub-architecture is a sequence of a Transformer model and a multi-layer perceptron layer, the input data of the second sub -architecture is a vector constituted of predicted gene expression levels from the first sub -architecture 16 and its output data is at least one clinical score.
[0102] In some embodiments, the device 1 further comprises a module 12 configured for preprocessing data from the first training dataset 21 and from the second training dataset 22. In addition, the module 12 may further be configured to perform augmentation of data from first training dataset 21 and from the second training dataset 22 to increase the volume of training data.
[0103] The device 1 further comprises a module 13 configured to train the architecture 20 using the first training dataset 21 and the second training dataset 22 received by the module 11.
[0104] In the embodiments where the architecture 20 comprises a first sub -architecture 16 and a second sub -architecture 18, the first sub-architecture 16 and the second subarchitecture 18 are trained separately and independently. Advantageously, using two distinct sub-architectures that are trained separately makes it possible to obtain components that are each highly optimized for a specific task. The first sub -architectureis specialized in predicting post-perturbation transcriptomic data, while the second subarchitecture is dedicated to translating these post-perturbation gene expression data into a clinical score. By decomposing the problem into two specialized learning tasks, each sub-architecture can focus on capturing the relevant patterns and representations for its respective objective. Together, these complementary and specialized sub-architectures lead to more accurate and robust clinical score predictions. In contrast, relying on a single unified architecture would force one model to simultaneously learn heterogeneous tasks, thereby reducing overall performance.
[0105] In those embodiments, the module 13 is configured to train the first subarchitecture 16 by implementing a stochastic gradient descent algorithm to predict postperturbation gene expression levels as precisely as possible, for instance using a loss function of Mean Squared Error type.
[0106] In the embodiments wherein the first training dataset 21 comprises only the first training sub-set, the module 13 is configured to train the first sub -architecture 16 by using a stochastic gradient descent algorithm and minimizing a loss function that can be of Mean Squared Error type.
[0107] In the embodiments wherein the first training dataset 21 comprises both first and second training sub-set, the module 13 is configured to first implement a first training phase of the first sub -architecture 16, based on the first training sub-dataset, and then a second training phase of the first sub-architecture 16, the second training phase being a fine-tuning learning phase, based on the second training sub-dataset.
[0108] In the fine-tuning learning phase implemented on the first sub-architecture 16 is configured to partially or fully retrain the first sub-architecture 16 trained in the first training phase on a more specific target task thanks to the use of the second training subdataset which is specific to one perturbation. In other words, the starting point is the first architecture 16 pre-trained based on the first training sub-dataset. The fine-tuning learning phase consists of re-training the pre-trained first architecture 16 so that, after the finetuning learning phase, the performances of the fine-tuned first architecture 16 in predicting post-perturbation gene expression levels after a predefined perturbation (i.e.corresponding to one predefined drug) are increased. The second training sub-dataset comprises typically at least 10 third training samples. The module 13 can typically implement a stochastic gradient descent algorithm to minimize a loss function of Mean Squared Error type.
[0109] The module 13 is configured to train the second sub -architecture 18 by implementing a stochastic gradient optimization algorithm to predict the clinical score(s) of the second training dataset 22 and minimizing a loss function such as a Mean Squared Error function. The prediction of the clinical score(s) consists of a regression task.
[0110] Once the training is completed, the module 13 is configured to output the trained artificial intelligence system for prediction 31. The trained artificial intelligence system 31 may then be stored in one or more local or remote database(s) 10. The latter can take the form of storage resources available from any kind of appropriate storage means, which can be notably a RAM or an EEPROM (Electrically-Erasable Programmable Read-Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk).
[0111] In its automatic actions, the device 1 may for example execute the following process (Figure 3):- receiving the first training data set 21 , the second training data set 22 and the initial untrained architecture 20 (step 51),- training the initial architecture 20 using the first training data set 21 and the second training data set 22 (step 52).
[0112] The present invention also relates to a device 2 for predicting an effect of a drug on a patient, where the drug modulates an expression of a target gene, using the trained artificial intelligence system of prediction 31 obtained from the device 1, as described above. In other words, the device 2 is configured to perform inference using the previously trained, and optionally fine-tuned on the specific case of the patient, artificial intelligence system obtained from device 1. The device 2 will be described in reference to a particular function embodiment as illustrated in Figure 4.
[0113] The device 2 is adapted to receive as input the artificial intelligence system of prediction 31, already trained, input molecular profiling data 32 of a patient and an input perturbation associated to at least one target gene 33.
[0114] The device 2 comprises a module 15 for receiving the trained artificial intelligence system 31, the input molecular profiling data 32 and the at least one input perturbation 33, stored in one or more local or remote database(s) 10. The latter can take the form of storage resources available from any kind of appropriate storage means, which can be notably a RAM or an EEPROM (Electrically-Erasable Programmable Read-Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk). In advantageous embodiments, the trained artificial intelligence system of prediction 31 and all its parameters have been previously generated by a system including the device 2 for training. Alternatively, the trained artificial intelligence system of prediction 31 and its parameters are received from a communication network.
[0115] More in details, the at least one input perturbation associated to at least one target gene 33 may be represented as a (input) perturbation vector having a predefined number of components each associated to a gene, each component comprising a perturbation value comprised between a minimum value and a maximum value, said minimum value corresponding to a completely inhibited state of said corresponding gene, and said maximum value corresponding to an activated state of said corresponding gene. For instance, the perturbation vector can be obtained and built from literature data and known drug mechanisms of action. For instance, in the case of the Rituximab drug, it is expected that the genes CD20 and CD 19 are highly under expressed in bulk RNA-seq profiling data.
[0116] The input molecular profiling data 32 of the patient may comprise at least an input transcriptomic profiling data from a biological sample of the patient (previously obtained). Transcriptomic data is generally obtained by analyzing the RNA molecules expressed in a biological sample. This provides insights into gene expression and cellular activity. This input transcriptomic profiling data may be for example be obtained from single-cell RNA sequencing (scRNA-seq), bulk RNA-sequencing or CITE-seq (Cellular Indexing of Transcriptomes and Epitopes by Sequencing). In particular, CITE-seq is a technique combining scRNA-seq with antibody-based protein detection tosimultaneously measure both gene expression and surface protein levels of individual cells.
[0117] In one example, the input transcriptomic profiling data from said biological sample comprises a vector comprising values representative of an expression level of a set of genes of the corresponding biological sample before said input perturbation of a biological sample, said set of genes comprising said at least one target gene, and wherein said post-perturbation transcriptomic profiling data prediction comprises a vector comprising values representative of a synthetic expression level of said set of genes of said patient after said perturbation of a biological sample.
[0118] According to one example, the input molecular profiling data 32 of the patient further comprises proteomic profiling data, genomic profiling data, cell composition and / or metabolic profiling measured from the corresponding biological sample. The proteomic profiling data may be obtained from mass spectrometry (MS) or protein microarrays, genomic profiling data from Next-generation sequencing (NGS), wholegenome sequencing (WGS), whole-exome sequencing (WES), or microarrays, cell composition from flow cytometry, single-cell RNA sequencing (scRNA-seq), or cytometry by time of flight (CyTOF) while metabolic Profiling data may be obtained from nuclear magnetic resonance (NMR) spectroscopy, mass spectrometry (LC-MS, GC-MS).
[0119] The device 2 comprises at least one processor or a module 17 configured to feed the input molecular profiling data 32 and the at least one input perturbation 33 to the trained artificial intelligence system 31 so as to generate 19 a post-perturbation molecular profiling data prediction of the patient 41 and at least one clinical score of the patient 42.
[0120] The device 2 is further configured to predict and determine the effect of said drug on said patient during a predefined timeframe comprising at least two administrations of said drug by implementing a clock.
[0121] The implemented clock is meant to signal the completion of successive time steps and initiates successive iterations during the predefined timeframe.
[0122] By predetermined timeframe, it is meant a time interval beginning at a starting time instant and ending at an end time instant, during which the effect of the drug on thepatient is predicted at several time instants between the starting time instant and the end time instant. The time instants are defined by the implemented clock and the predefined interval.
[0123] By absorption phase, it is meant the measurement of a drug concentration in the body of a patient with respect to the time of administration. As the absorption data correspond to concentration data, preclinical data of drug concentrations are used to define the perturbation data, as will be explained further. Pharmacokinetics mathematical models or Plasma Concentration-Time Curve may be used to estimate this absorption phase.
[0124] By predefined interval, it is meant the time interval between two successive predictions by the device 2 of the effect of the drug on the patient.
[0125] By predefined administration interval, it is meant a time interval between two administrations of the drug to the patient.
[0126] In this example, the successive perturbation (at the i+lthiteration) is obtained by modifying the current perturbation (at the ithiteration) using the absorption phase. In this example, the given number of iterations is determined on the base of the predefined timeframe and the predefined interval. The successive perturbation (at the i+lthiteration) is obtained by feeding back the absorption phase to the total perturbation. For instance, if 50% of the treatment (i.e. drug concentration at initial time) was eliminated, the successive perturbation corresponds to the initial perturbation reduced by 50%.
[0127] In short, the implementation of the clock and of the cycles creates a comprehensive, time-resolved prediction of the drug’s effects from the molecular to the clinical level.
[0128] The device 2 may interact with a user interface 18, via which information can be entered and retrieved by a user. The user interface 18 includes any means appropriate for entering or retrieving data, information or instructions, notably visual, tactile and / or audio capacities that can encompass any or several of the following means as well known by a person skilled in the art: a screen, a keyboard, a trackball, a touchpad, a touchscreen, a loudspeaker, a voice recognition system.
[0129] In its automatic actions, the device 2 may for example execute the following process (Figure 5):- receiving the input molecular profiling data 32, comprising at least input transcriptomic profiling data from a biological sample of said patient, and current (i.e., input) perturbation of said at least one target gene 33 (step 61),- providing the input molecular profiling data 32 and the input perturbation 33 to the trained artificial intelligence model 31 so as to generate a post-perturbation molecular profiling data prediction of the patient 41 and at least one clinical score of the patient 42 (step 62),- transforming the input molecular profiling data 32 and the input perturbation 33 with the trained artificial intelligence model 31 so as to generate a postperturbation molecular profiling data prediction of the patient 41 and at least one clinical score of the patient 42 (step 63).
[0130] Figure 6 shows the result of implementation of the computer-implemented process executed by the device 2. This result can be represented as an array outputted by the first sub -architecture 15 of the trained artificial intelligence model 31, where each row represents gene expression levels of a patient after application of a perturbation of a rituximab treatment, each column corresponding to a gene. The rows 6 indicated in dotted patterns correspond to training data, while the rows 7 indicated with a grid pattern correspond to inference data corresponding to a benchmark dataset. Darker grey level in the array means inhibition of the gene and lighter grey level in the array means activation of the gene. It can be observed that for genes of interest, expression patterns are very similar between the training data and the inference data.
[0131] In its automatic actions, the device 2, when configured to determine an effect of a drug on a patient during a predefined timeframe, may further execute the following process (Figure 7):- receiving input data comprising the input molecular profiling data 32, an absorption phase of the patient, a predefined interval, a predefined administration interval representing an interval between two successive administrations, - implementing a clock, for a given number of iterations, on the base of the predefined interval and an autoregressive model, by iteratively providing currentinput molecular profiling data 32i to the trained artificial intelligence model 31 so as to generate a current post-perturbation molecular profiling data prediction of the patient 41i and feeding the current perturbation molecular profiling data prediction 41i as an updated input molecular profiling data 32i+ 1 at a successive iteration together with a successive perturbation.
[0132] Figure 8 shows the resulting clinical scores S(t) obtained with the trained artificial intelligence system 31 by application of a time spanning perturbation representative of rituximab administrations at several time instants over a predefined timeframe of 48 weeks on a cohort of patients. On Figure 8, the represented clinical score S(t) is representative of the activity of the disease at a clinical level and is the DAS28 score used in the assessment of rheumatoid arthritis. The clinical scores were predicted (violin plots 8 in lighter grey level on the right) with the trained artificial intelligence system 31 at four different time instants, namely, at an initial time instant to, 16 weeks after (ti), 24 weeks after (t2) and 48 weeks after (L). The ground truth violin plots 9 are displayed on the left of each prediction (darker grey level). It can be observed that the trained artificial intelligence system of prediction 31 is able to retrieve the trend of evolution of the disease activity score.
[0133] A particular apparatus may embody the device 1 as well as the device 2 described above. It corresponds for example to a workstation, a laptop, a tablet, a smartphone, or a head-mounted display (HMD).
[0134] That apparatus is suited to predictions of patient post-perturbation molecular profiling data prediction and of at least one clinical score, and to related ML training. It comprises the following elements, connected to each other by a bus 95 of addresses and data that also transports a clock signal:- a microprocessor (or CPU);- a graphics card comprising several Graphical Processing Units (or GPUs) and a Graphical Random Access Memory (GRAM); the GPUs are quite suited to image processing and also to deep learning computations, due to their highly parallel structure;- a non-volatile memory of ROM type;- a RAM;- one or several I / O (Input / Output) devices such as for example a keyboard, a mouse, a trackball, a webcam; other modes for introduction of commands such as for example vocal recognition are also possible;- a power source; and- a radiofrequency unit.
[0135] According to a variant, the power supply is external to the apparatus.
[0136] The apparatus also comprises a display device of display screen type directly connected to the graphics card to display synthesized images calculated and composed in the graphics card. According to a variant, a display device is external to the apparatus and is connected thereto by a cable or wirelessly for transmitting the display signals. The apparatus, for example through the graphics card, comprises an interface for transmission or connection adapted to transmit a display signal to an external display means such as for example an LCD or plasma screen or a video-projector. In this respect, the RF unit can be used for wireless transmissions.
[0137] When switched-on, the microprocessor loads and executes the instructions of the program contained in the RAM.
[0138] As will be understood by a skilled person, the presence of the graphics card is not mandatory, and can be replaced with entire CPU processing and / or simpler visualization implementations.
[0139] In variant modes, the apparatus may include only the functionalities of the device 1, and not the learning capacities of the device 2. In addition, the device 1 and / or the device 2 may be implemented differently than a standalone software, and an apparatus or set of apparatus comprising only parts of the apparatus may be exploited through an API call or via a cloud interface.
Claims
CLAIMS1. A device (1) for generating an artificial intelligence system (31) for prediction of an effect of a drug on a patient, said device (1) comprising:at least one input configured to receive:• a first training dataset (21) comprising at least one first training subdataset,said first training sub -dataset comprising a plurality of first training samples, each first training sample being associated to a subject from a first plurality of subjects and a perturbation from a plurality of perturbations; wherein each first training sample comprises first molecular profiling data, comprising at least first transcriptomic profiling data, measured from a biological sample of the subject before application of the perturbation, and second molecular profiling data measured after application of said perturbation on said biological sample;• a second training dataset (22) comprising a plurality of second training samples associated to a second plurality of subjects, each second training sample comprising subject biological data and clinical scores associated to the subject; wherein said subject biological data comprises at least second molecular profiling data, comprising at least second transcriptomic profiling data acquired from a biological sample of said subject;• an initial architecture (20) of said artificial intelligence system (31), said initial architecture (20) comprising:a) a first sub-architecture (16) configured to receive as input molecular profiling data from a biological sample of a patient comprising at least input transcriptomic profiling data and at least one input perturbation associated to at least one target gene, and to output a post-perturbation molecular profiling data prediction ;b) a second sub -architecture (18) configured to receive said postperturbation molecular profiling data prediction and to output at least one clinical score;at least one processor configured to generate said artificial intelligence system (31) via a first training process of said first sub-architecture, comprising a first phase using said first training sub-dataset and a second training process of said second sub -architecture (18) using said second training dataset (22); at least one output configured to provide as output said artificial intelligence system (31).
2. The device (1) according to claim 1, wherein said first training dataset (21) further comprises a second training sub-dataset,said second training sub-dataset, being associated to a predefined perturbation, comprising a plurality of third training samples from a third plurality of subjects, each third training sample comprising first specific molecular profiling data, comprising at least first specific transcriptomic profiling data, and second specific molecular profiling data, comprising at least second specific transcriptomic profiling data, measured from a biological sample of a subject from said third plurality of subjects, before and after application of said predefined perturbation on said biological sample;wherein said first training process comprises a second phase of fine-tuning learning based on said second training sub-dataset.
3. The device (1) according to either one of claim 1 or 2, wherein said input transcriptomic profiling data from said biological sample comprises a vector comprising values representative of an expression level of a set of genes of the corresponding biological sample before said input perturbation of a biological sample, said set of genes comprising said at least one target gene, and wherein said post-perturbation transcriptomic profiling data prediction comprises a vector comprising values representative of a synthetic expression level of said set of genes of said patient after said perturbation of a biological sample.
4. The device (1) according to claim 3, wherein said expression level comprises levels of RNA transcripts.
5. The device (1) according to any one of claims 1 to 4, wherein said effect of said drug is an activation of said at least one target gene or an inhibition of said at least one target gene.
6. The device (1) according to claim 5, wherein each perturbation of the plurality of perturbation, input perturbation or predefined perturbation, is associated to a one or more target genes, each perturbation is represented as a perturbation vector having a predefined number of components each associated to a gene, each component comprising a perturbation value comprised between a minimum value and a maximum value, said minimum value corresponding to a completely inhibited state of said corresponding gene, and said maximum value corresponding to an activated state of said corresponding gene.
7. The device (1) according to any one of claims 1 to 6, wherein the at least one clinical score comprises at least one among: a synthetic activity level of a disease, a synthetic inflammation biomarker, a synthetic blood cell count.
8. The device (1) according to any one of claims 1 to 7, wherein each of said first subarchitecture (16) and said second sub -architecture (18) is a foundation model based on a generative model.
9. The device (1) according to claim 8, wherein said first sub-architecture (16) and / or said second sub-architecture (18) is based on a Transformer architecture.
10. The device (1) according to any one of claims 1 to 9, wherein each first training sample, each second training sample and each third training sample further comprises proteomic profiling data, genomic profiling, cell composition and / or metabolic profiling measured from the corresponding biological sample.
11. A device (2) for predicting an effect of a drug on a patient, said drug modulating an expression of at least one target gene, by using the artificial intelligence system (31) obtained with a device (1) according to any one of claims 1 to 10 and comprising:at least one input configured to receive:• input molecular profiling data of said patient (32) comprising at least input transcriptomic profiling data from a biological sample of said patient; • a current perturbation of said at least one target gene (33), said current perturbation being representative of an action of said drug on said at least one target gene at a given time instant;at least one processor configured to:• provide, as input to said artificial intelligence system (31), said input molecular profiling data and said current perturbation,• obtain, as output, a patient post-perturbation molecular profiling data prediction and at least one clinical score, said at least one clinical score being representative of said effect of said drug on said patient at said given time instant,at least one output configured to provide said patient post-perturbation molecular profiling data and at least one clinical score.
12. The device (2) according to claim 11, wherein the effect of said drug on said patient is determined during a predefined timeframe comprising at least two administrations of said drug, wherein:said at least one input is further configured to receive an absorption phase for said patient, said predefined timeframe, a predefined interval, a predefined administration interval representing the interval between two successive administrations;said at least one processor is configured to implement a clock on the base of said predefined interval and an autoregressive model, by iteratively feeding the post-perturbation molecular profiling data prediction obtained at a current (ith) iteration as input molecular profiling data for the artificial intelligence model at a successive (i+lth) iteration together with a successive perturbation, wherein the successive perturbation is obtained by modifying the current perturbation using said absorption phase,wherein a number of iterations is determined on the base of the predefined timeframe and said predefined interval.
13. A computer-implemented method for generating an artificial intelligence system of prediction (31) of an effect of a drug on a patient, said method comprising:- receiving:• a first training dataset (21) comprising a first training sub -dataset, said first training sub-dataset comprising a plurality of first training samples, each first training sample being associated to a subject from a first plurality of subjects and a perturbation from a plurality of perturbations; wherein each first training sample comprises first molecular profiling data, comprising at least first transcriptomic profiling data, measured from a biological sample of said subject before application of said perturbation, and second molecular profiling data, measured after application of said perturbation on said biological sample;• a second training dataset (22) comprising a plurality of second training samples associated to a second plurality of subjects, each second training sample comprising subject biological data and clinical scores associated to said subject; wherein said subject biological data comprises at least second molecular profiling data, comprising at least second transcriptomic profiling data, acquired from a biological sample of said subject;• an initial architecture (20) of said artificial intelligence system (31), said initial architecture (20) comprising:a first sub -architecture configured to receive as input molecular profiling data from a biological sample of a patient comprising at least input transcriptomic profiling data and at least one input perturbation associated to at least one target gene, and to output a post-perturbation molecular profiling data prediction;a second sub-architecture configured to receive said post-perturbation molecular profiling data prediction and to output at least one clinical score;- generating said artificial intelligence system (31) via a first training process of said first sub-architecture (16) comprising a first phase using said first trainingsub-dataset and a second training process of said second sub -architecture (18) using second training dataset (22);- providing as output said artificial intelligence system of prediction (31).
14. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to claim 13.
15. A computer-readable storage medium, comprising instructions which, when executed by a computer, cause the computer to carry out the method according to claim 13.