System and method of interpretable prediction of a subject's condition

The FUSE model addresses the lack of interpretability in IVF prediction models by using embeddings and context-aware attention to generate significance matrices, improving accuracy and transparency in clinical decision-making.

WO2026133333A1PCT designated stage Publication Date: 2026-06-25RAMBAM MED TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RAMBAM MED TECH
Filing Date
2025-12-18
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing machine learning models for predicting mature oocyte retrieval in IVF cycles lack interpretability, leading to variability in outcomes due to subjective decision-making and a lack of transparency in clinical settings.

Method used

The FUSE model employs an interpretable deep learning architecture with embeddings and context-aware attention-based feature selection to provide transparent and accurate predictions by generating ad-hoc and global significance matrices, enhancing the understanding of parameter contributions.

Benefits of technology

FUSE achieves higher predictive accuracy and transparency, enabling data-driven, objective decision-making in IVF by providing clear explanations of how each parameter influences the outcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025051134_25062026_PF_FP_ABST
    Figure IL2025051134_25062026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method of providing an interpretable prediction of a condition of a subject may include receiving data including values of patient parameters, representing a subject's cunent physiology, and calculating an embedding matrix, representing said data in an embedding space. The embedding matrix may be processed through a cascade of stages. Each stage may include a respective Machine-Learning (ML) based, context-aware attention model, configured to generate an ad-hoc significance matrix pertaining to that stage. Embodiments may aggregate the ad-hoc significance matrices of these stages, to obtain a global significance matrix, representing contribution of each of said parameters in predicting the subject's condition. Embodiments may subsequently apply an ML-based regression model on the global significance matrix, to predict the condition of the subject, and provide the global significance matrix as an interpretation of that prediction.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD OF INTERPRETABLE PREDICTION OF A SUBJECT’S CONDITIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 736,091, filed 19 December 2024, entitled “SYSTEM AND METHOD OF INTERPRETABLE PREDICTION OF A SUBJECT’S CONDITION”. The contents of the above application is all incorporated by reference as if fully set forth herein in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates generally to the technological field of assistive diagnostics. More specifically, the present invention relates to a system and method of interpretable prediction of a subject’s condition.BACKGROUND OF THE INVENTION

[0003] Choosing the trigger date for in vitro fertilization (IVF) cycles represents a decision for physicians, aiming to maximize the number of mature oocytes retrieved. This decision relies heavily on subjective clinical judgment, influenced by various parameters such as patient age, treatment protocol, hormone levels, and individual response to medications. The subjective nature of this decision-making process can lead to variability in outcomes, highlighting the need for more objective and data-driven approaches.

[0004] Existing solutions have explored the use of machine learning models to transform decisionmaking in IVF into a more objective process. Models such as LightGBM, XGBoost, and Feed Forward Neural Networks have been employed to predict the number of mature oocytes retrieved. While these models have demonstrated potential in improving prediction accuracy, they often function as "black boxes," lacking transparency in their decision-making processes. This lack of interpretability poses significant challenges, particularly in clinical settings where understanding the rationale behind predictions is necessary for trust and validation.SUMMARY OF THE INVENTION

[0005] Embodiments of the invention introduce a novel, interpretable deep learning architecture, referred to herein as Feature Utilization via Selective Exchanges (FUSE).

[0006] As explained herein, FUSE was designed to predict a number of mature oocytes retrieved during IVF cycles.

[0007] However, it may be appreciated that embodiments of the invention may allow the FUSE model to provide interpretable prediction of other conditions of human subjects, given an appropriate dataset. Therefore, the example used herein, of predicting the number of mature oocytes during IVF treatment, should not be regarded as limiting in any way.

[0008] The inventors have experimentally demonstrated that FUSE achieves higher predictive accuracy compared to traditional models, while providing enhanced interpretability and transparency. The FUSE architecture incorporates advanced techniques such as embeddings for both categorical and continuous data, and context-aware attention-based feature selection. As elaborated herein, the components of the FUSE model enable learning complex relationships while maintaining clarity in the decision-making process, addressing the limitations of previous models and fostering trust in clinical applications.

[0009] As elaborated herein, embodiments of the invention enable interpretable predictions by utilizing an embedding matrix to represent patient data in a continuous space, allowing for a more nuanced understanding of the relationships between parameters.

[0010] This representation facilitates the processing of data through a cascade of context-aware attention models or stages, which generate respective ad-hoc significance matrices. Each ad-hoc significance matrix may define the contribution of each patient parameter in predicting the subject's condition, per the respective stage.

[0011] Embodiments of the invention may calculate a global significance matrix from these ad- hoc matrices. Embodiments of the invention may thereby provide a comprehensive view of parameter importance, enhancing the interpretability of the prediction. This approach addresses the limitations of traditional "black box" models by offering transparency in the decision-making process, which is crucial in clinical settings where understanding the rationale behind predictions is necessary for trust and validation.

[0012] The application of a pretrained ML-based classifier model on the global significance matrix ensures that the prediction is both accurate and interpretable, providing a clear explanation of howeach parameter influences the outcome. Embodiments of the invention may support clinical decision-making by offering a data-driven, objective approach to predicting a subject's condition, such as the number of mature oocytes available for retrieval in IVF treatment.

[0013] Embodiments of the invention may include a system for providing an interpretable prediction of a condition of a subject.

[0014] Embodiments of the system may include: a non-transitory memory device, wherein modules of instruction code may be stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code.

[0015] Upon execution of said modules of instruction code, the at least one processor may be configured to receive data that may include values of patient parameters, representing a subject’s current physiology, and calculate an embedding matrix, representing that data in an embedding space.

[0016] The at least one processor may process the embedding matrix through a cascade of stages, where each stage may include a respective ML based, context-aware attention model. Each context-aware attention model may be configured to generate an ad-hoc significance matrix pertaining to that stage.

[0017] The at least one processor may aggregate the ad-hoc significance matrices of said stages to obtain a global significance matrix, representing contribution of each of said parameters in predicting the subject’s condition, and apply an ML-based regression model on the global significance matrix, to predict the condition of the subject. The at least one processor may further provide the global significance matrix as an interpretation of said prediction, as elaborated herein.

[0018] Additionally, or alternatively, the at least one processor may apply the ML-based regression model further on the embedding matrix, to predict the condition of the subject.

[0019] The predicted condition of the subject may include, for example a number of mature oocytes that may be available for retrieval from that subject, during an IVF treatment, in relation to a predetermined trigger date.

[0020] Additionally, or alternatively, the patient parameters may include a first set of categorical parameters pertaining to ovarian stimulation regimen, which the patient is undergoing, or undergone. The first set of categorical parameters may include, for example, an antagonist regimen, a long agonist regimen, a microflare regimen, an OI, as a preparatory step for IVF, and the like.

[0021] Additionally, or alternatively, the patient parameters may include a second set of categorical parameters pertaining to patient diagnoses. The second set of categorical parameters may include, for example diminished ovarian reserve, endometriosis, a male factor, an ovulatory dysfunction, a tubal disease, and a uterine factor.

[0022] Additionally, or alternatively, the patient parameters may include one or more continuous parameters selected from a list consisting of: the patient’s age, their Body Mass Index (BMI), their basal Follicle-Stimulating Hormone (FSH) level, their trigger estradiol (E2) level, their endometrial thickness on trigger day, a number of stimulation days during a current IVF cycle, and the like.

[0023] According to some embodiments, the at least one processor may be further configured to, for each patient parameter, calculate a parameter embedding vector having a plurality of parameter entries. The at least one processor may subsequently aggregate the parameter embedding vectors of all patient parameters, to obtain the embedding matrix.

[0024] Additionally, or alternatively, the at least one processor may be further configured to calculate a patient parameter’s embedding vector by multiplying that parameter’s value with a respective, uniquely indicative vector, representing a type of that patient parameter.

[0025] According to some embodiments, each context-aware attention model may include a feature selection model, configured to calculate a respective mask matrix. The mask matrix may represent a selection of parameter entries of the embedding matrix as features of interest, at the respective stage.

[0026] The at least one processor may employ a feature selection model of a current stage of the cascade to receive one or more previous mask matrices, pertaining to respective, previous stages of the cascade. The at least one processor may calculate the mask matrix of the current stage based on the one or more mask matrices of previous stages. This calculation may include suppression of re-selection of parameter entries that were already selected by the mask matrices of previous stages.

[0027] According to some embodiments, each context-aware attention model may further include a feature interaction model. The at least one processor may be further configured to employ a feature interaction model of a current stage to obtain the selection of parameter entries, based on the mask matrix of the current stage; apply a self-attention algorithm on the selected parameter entries, so as to assign respective weights to one or more parameter embedding vectors of theembedding matrix; and generate the ad-hoc significance matrix of the current stage based on said weights of parameter embedding vectors.

[0028] Embodiments of the invention may include a method of providing an interpretable prediction of a condition of a subject. Embodiments of the method may include receiving data comprising values of patient parameters, representing a subject’s current physiology; calculating an embedding matrix, representing said data in an embedding space; processing said embedding matrix through a cascade of stages, each comprising a respective Machine-Learning (ML) based, context-aware attention model, wherein each context-aware attention model is configured to generate an ad-hoc significance matrix pertaining to that stage; aggregating the ad-hoc significance matrices of said stages to obtain a global significance matrix, representing contribution of each of said parameters in predicting the subject’s condition; applying an ML-based regression model on the global significance matrix, to predict the condition of the subject; and providing the global significance matrix as an interpretation of said prediction.

[0029] Embodiments of the invention may include a method of providing an interpretable prediction of a number of mature oocytes available for retrieval during In Vitro Fertilization (IVF) treatment, by at least one processor. Embodiments of the method may include receiving data comprising values of patient parameters representing a subject undergoing ovarian stimulation, and calculating an embedding matrix, representing said data in an embedding space. The at least one processor may process the embedding matrix through a cascade of stages, each including a respective Machine-Learning (ML) based, context-aware attention model. Each context-aware attention model may be configured to generate an ad-hoc significance matrix pertaining to that stage. The at least one processor may aggregate the ad-hoc significance matrices of the stages to obtain a global significance matrix that may represent contribution of each of said parameters in predicting the number of mature oocytes.

[0030] Additionally, or alternatively, the at least one processor may apply an ML-based regression model on the global significance matrix to predict the number of mature oocytes available for retrieval in relation to a predetermined trigger date. The at least one processor may subsequently providing the global significance matrix, e.g., via a user interface (UI), as an interpretation of said prediction to support clinical decision-making regarding trigger timing for oocyte retrieval.

[0031] According to some embodiments, the patient parameters may include one or more categorical parameters such as an ovarian stimulation regimen, a diagnosis of diminished ovarianreserve, endometriosis, a male factor, ovulatory dysfunction, a tubal disease, and a uterine factor. Additionally, or alternatively, the patient parameters may include one or more continuous parameters such as age, Body Mass Index (BMI), basal Follicle-Stimulating Hormone (FSH) level, trigger estradiol (E2) level, endometrial thickness on trigger day, a number of stimulation days, follicle counts by size range, average follicle size, and total follicle count.

[0032] According to some embodiments, each context-aware attention model may include a feature selection model configured to calculate a respective mask matrix representing a selection of parameter entries of the embedding matrix as features of interest at the respective stage, and a feature interaction model configured to apply a self-attention algorithm on selected parameter entries to generate the ad-hoc significance matrix of that stage.BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features, and advantages thereof, may best be understood by reference to the following detailed description when read with the accompanying drawings in which:

[0034] Fig. 1 is a block diagram, depicting a computing device which may be included in a system for interpretably predicting a condition of a subject, according to some embodiments;

[0035] Fig. 2 is a block diagram, depicting a system for interpretably predicting a condition of a subject, according to some embodiments;

[0036] Fig. 3 is a schematic chart, showing results of predictions of conditions of subjects in two cases, according to some embodiments;

[0037] Fig. 4 is a flow diagram, depicting a method of providing interpretable prediction of a condition of a subject, according to some embodiments; and

[0038] Fig. 5 is a flow diagram, depicting a method of providing an interpretable prediction of a number of mature oocytes available for retrieval during In Vitro Fertilization (IVF) treatment, according to some embodiments.

[0039] It will be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where consideredappropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION OF THE PRESENT INVENTION

[0040] One skilled in the art will realize the invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting of the invention described herein. Scope of the invention is thus indicated by the appended claims, rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.

[0041] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the present invention. Some features or elements described with respect to one embodiment may be combined with features or elements described with respect to other embodiments. For the sake of clarity, discussion of same or similar features or elements may not be repeated.

[0042] Although embodiments of the invention are not limited in this regard, discussions utilizing terms such as, for example, “processing,” “computing,” “calculating,” “determining,” “establishing”, “analyzing”, “checking”, or the like, may refer to operation(s) and / or process(es) of a computer, a computing platform, a computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer’s registers and / or memories into other data similarly represented as physical quantities within the computer’s registers and / or memories or other information non-transitory storage medium that may store instructions to perform operations and / or processes.

[0043] Although embodiments of the invention are not limited in this regard, the terms “plurality” and “a plurality” as used herein may include, for example, “multiple” or “two or more”. The terms “plurality” or “a plurality” may be used throughout the specification to describe two or more components, devices, elements, units, parameters, or the like. The term “set” when used herein may include one or more items.

[0044] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Additionally, some of the described method embodiments or elements thereof can occur or be performed simultaneously, at the same point in time, or concurrently.

[0045] Reference is now made to Fig. 1, which is a block diagram depicting a computing device, which may be included within an embodiment of a system for providing interpretable prediction of a condition of a subject, according to some embodiments.

[0046] Computing device 1 may include a processor or controller 2 that may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing or computational device, an operating system 3, a memory 4, executable code 5, a storage system 6, input devices 7 and output devices 8. Processor 2 (or one or more controllers or processors, possibly across multiple units or devices) may be configured to carry out methods described herein, and / or to execute or act as the various modules, units, etc. More than one computing device 1 may be included in, and one or more computing devices 1 may act as the components of, a system according to embodiments of the invention.

[0047] Operating system 3 may be or may include any code segment (e.g., one similar to executable code 5 described herein) designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 1, for example, scheduling execution of software programs or tasks or enabling software programs or other modules or units to communicate. Operating system 3 may be a commercial operating system. It will be noted that an operating system 3 may be an optional component, e.g., in some embodiments, a system may include a computing device that does not require or include an operating system 3.

[0048] Memory 4 may be or may include, for example, a Random- Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Memory 4 may be or may include a plurality of possibly different memory units. Memory 4 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM. In one embodiment, a non-transitory storage medium such as memory 4, a hard disk drive, another storage device, etc. may store instructions or codewhich when executed by a processor may cause the processor to carry out methods as described herein.

[0049] Executable code 5 may be any executable code, e.g., an application, a program, a process, task, or script. Executable code 5 may be executed by processor or controller 2 possibly under control of operating system 3. For example, executable code 5 may be an application that may provide interpretable prediction of a condition of a subject as further described herein. Although, for the sake of clarity, a single item of executable code 5 is shown in Fig. 1 , a system according to some embodiments of the invention may include a plurality of executable code segments similar to executable code 5 that may be loaded into memory 4 and cause processor 2 to carry out methods described herein.

[0050] Storage system 6 may be or may include, for example, a flash memory as known in the art, a memory that is internal to, or embedded in, a micro controller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disk (BD), a universal serial bus (USB) device or other suitable removable and / or fixed storage unit. Data pertaining to a subject may be stored in storage system 6 and may be loaded from storage system 6 into memory 4 where it may be processed by processor or controller 2. In some embodiments, some of the components shown in Fig. 1 may be omitted. For example, memory 4 may be a non-volatile memory having the storage capacity of storage system 6. Accordingly, although shown as a separate component, storage system 6 may be embedded or included in memory 4.

[0051] Input devices 7 may be or may include any suitable input devices, components, or systems, e.g., a detachable keyboard or keypad, a mouse and the like. Output devices 8 may include one or more (possibly detachable) displays or monitors, speakers and / or any other suitable output devices. Any applicable input / output (I / O) devices may be connected to Computing device 1 as shown by blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device or external hard drive may be included in input devices 7 and / or output devices 8. It will be recognized that any suitable number of input devices 7 and output device 8 may be operatively connected to Computing device 1 as shown by blocks 7 and 8.

[0052] A system according to some embodiments of the invention may include components such as, but not limited to, a plurality of central processing units (CPU) or any other suitable multipurpose or specific processors or controllers (e.g., similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.

[0053] The term neural network (NN) or artificial neural network (ANN), e.g., a neural network implementing a machine learning (ML) or artificial intelligence (Al) function, may be used herein to refer to an information processing paradigm that may include nodes, referred to as neurons, organized into layers, with links between the neurons. The links may transfer signals between neurons and may be associated with weights. An NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training an NN for the specific task may involve adjusting these weights based on examples. Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN. Typically, the neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights. At least one processor (e.g., processor 2 of Fig. 1) such as one or more CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.

[0054] Reference is now made to Fig. 2, which depicts a system 10 for providing interpretable prediction of a condition of a subject, according to some embodiments.

[0055] According to some embodiments of the invention, system 10 may be implemented as a software module, a hardware module, or any combination thereof. For example, system may be or may include a computing device such as element 1 of Fig. 1 , and may be adapted to execute one or more modules of executable code (e.g., element 5 of Fig. 1) to provide interpretable prediction of a condition of a subject, as further described herein.

[0056] As shown in Fig. 2, arrows may represent flow of one or more data elements to and from system 10 and / or among modules or elements of system 10. Some arrows have been omitted on Fig. 2 for the purpose of clarity.

[0057] As shown in Fig. 2, system 10 may receive data comprising values of patient parameters 20, representing a subject’s current physiology. These patient parameters 20 may be used for subsequent processing stages within system 10, to determine the subjects condition.

[0058] For example, the predicted condition of the subject may include a prediction a of a number of mature oocytes that may be available for retrieval from a patient, during an In Vitro Fertilization (IVF) treatment cycle, in relation to predetermined trigger dates. In this example, patientparameters 20 data may include value(s) of a plurality of relevant numerical and categorical patient parameters.

[0059] For example, patient parameters 20 may include a set of categorical parameters pertaining to an ovarian stimulation regimen. Such categorical parameters may include, for example an antagonist regimen, a long agonist regimen, a microflare regimen, an Ovulation Induction (01), as a preparatory step for IVF, and the like.

[0060] Additionally, or alternatively, patient parameters 20 may include a set of categorical patient parameters pertaining to patient diagnoses such as diagnosis of a diminished ovarian reserve, diagnosis of endometriosis, indication of a male factor, diagnosis of ovulatory dysfunction, diagnosis of a tubal disease, indication of a uterine factor, and the like.

[0061] Additionally, or alternatively, patient parameters 20 may include one or more continuous, numerical parameters describing the patient, such as an age of the patient, the patient’s Body Mass Index (BMI), the patient’s basal Follicle-Stimulating Hormone (FSH) level, the patient’s trigger estradiol (E2) level, indication of endometrial thickness on trigger day, a number of stimulation days during a current IVF cycle, and the like.

[0062] Additionally, or alternatively, system 10 may adjust a target variable to account for potential inaccuracies in follicle count reporting. For example, during ovarian stimulation, ultrasound scanning may be used to monitor ovarian follicular development. In some cases, particularly in high responders, follicle counts may be underreported as smaller follicles may not be accurately counted. To mitigate bias from such reporting errors, the target variable may be adjusted to represent a number of follicles that yield a mature oocyte, calculated as a minimum value between a total follicle count and a number of mature oocytes retrieved. This adjustment may ensure that a ratio between oocytes and follicles remains consistent, helping to reduce impact of underreporting errors and improving prediction accuracy across different patient response profiles.

[0063] Embedding model 200 may be responsible for calculating an embedding matrix 200EM, representing the patient data in an embedding space. As elaborated herein, embedding model 200 may process both categorical and continuous, numerical parameters, transforming them into numerical representations known as embedding vectors 200EV. These embedding vectors 200EV may capture meanings and relationships of patient parameters’ data 20, and may allow subsequent analysis by machine learning algorithms, as elaborated herein.

[0064] According to some embodiments, for one or more (e.g., each) patient parameter, embedding model 200 may calculate a parameter embedding vector 200EV having a plurality of parameter entries 200ENT.

[0065] Embedding model 200 may calculate, for each of the one or more parameters 20, a respective, uniquely indicative vector 200IND. Vector 200IND, may uniquely representing, or indicate the specific type of that patient parameter.

[0066] For example, a first indicative vector 200IND may indicate a specific, categorical patient parameter, having a binary value (e.g., whether or not the patient is receiving a specific treatment). In another example, a second indicative vector 200IND may indicate a specific, continuous patient parameter, having a continuous value, such as the patient’s age.

[0067] When the patient parameter is a categorical one, embedding model 200 may calculate a parameter embedding vector 200EV based on the binary manifestation of that patient parameter, e.g., as equal to that patient parameter’s indicative vector 200IND.

[0068] When the patient parameter is a continuous one, embedding model 200 may calculate a parameter embedding vector 200EV based on the numeric manifestation of that patient parameter, e.g., as a multiplication product of (i) the patient parameter’s indicative vector 200IND (ii) the numeric value 20 of that parameter as manifested in the patient data 20.

[0069] Embedding model 200 may aggregate the parameter embedding vectors of two or more (e.g., all) patient parameters 20, to obtain the embedding matrix 200EM.

[0070] After embedding the patient parameters 20, each parameter 20 may be represented by a respective embedding vector 200EV, which may correspond to a row in the embedding matrix 200EM.

[0071] In other words, a row (or column) of embedding matrix 200EM may correspond to an embedding vector 200EV, which represents both a patient-parameter type 20 (e.g., BMI), and a value of the parameter of that type (e.g., BMI=27). As used herein, the term “entry” may refer to a single numerical value of embedding matrix 200EM, e.g., a single component of an embedding vector 200EV within embedding matrix 200EM.

[0072] Feature selection 120 may not be a binary process; instead, a parameter may be partially selected by utilizing only certain entries 200ENT in its respective row within embedding matrix 200EM. Typically, in a single step, most parameters may exhibit low utilization and thus may not be selected, while a few features may show high utilization, indicating they were chosen.Calculations may be performed either at the row level of embedding matrix 200EM, or at an entry 200ENT level of that matrix.

[0073] System 10 may process embedding matrix 200EM through a cascade of stages (denoted Si, S2, Sn), each comprising a respective machine-learning (ML) based, context-aware attention model 100. As elaborated herein, each context-aware attention model 100 may be configured to generate an ad-hoc significance matrix 110M, representing contribution of each parameter 20 in predicting the subject’s condition.

[0074] As shown in Fig. 2, one or more (e.g., each) context-aware attention model 100 may include (a) a feature selection module 120, and (b) a feature interaction module 110.

[0075] The feature selection module 120 may be configured to calculate a respective mask matrix 120M. Mask matrix 120M may represent a selection of parameter entries 200ENT of the embedding matrix 200EM as features of interest at the respective stage.

[0076] According to some embodiments, feature selection module 120 of a current stage of the cascade may receive one or more previous mask matrices 120M, pertaining to respective, previous stages of the cascade. As elaborated herein, feature selection module 120 may calculate the mask matrix 120M of the current stage based on the one or more mask matrices of previous stages. This calculation may include suppression of re-selection of parameter entries that were already selected by the mask matrices 120M of previous stages.

[0077] Feature interaction module 110 may obtain the selection of parameter entries 200ENT based on the mask matrix 120M of the current stage, and may apply a self-attention algorithm on the selected parameter entries 200ENT, to assign respective weights to one or more parameter embedding vectors 200EV of the embedding matrix 200EM. Feature interaction module 110 may thereby generate the ad-hoc significance matrix 110M of the current stage based on the weights of parameter embedding vectors. The term “ad-hoc” may be used herein in relation to significance matrices 110M to indicate their pertinence to a specific context of a current stage in the cascade. In other words, ad-hoc significance matrices 110M may highlight contribution of each parameter 20 in predicting the subject’s condition, within the context of their respective stages of the cascade.

[0078] According to some embodiments, system 10 may proceed to calculate a global significance matrix 110GM based on the ad-hoc significance matrices 110M of context-aware attention models 100 of one or more (e.g., all) respective stages of the cascade.

[0079] Additionally, or alternatively, system 10 may aggregate the ad-hoc significance matrices 110M of context-aware attention models 100 of one or more stages of the cascade to calculate a global significance matrix 110GM. The aggregation may include summing the ad-hoc significance matrices 110M across the stages. By summing the outputs from each stage, system 10 may quantify how much of each parameter 20 is utilized at each stage of the cascade, which may enable system 10 to determine an overall utilization of each parameter 20 across all stages.

[0080] Global significance matrix 110GM may thereby represent contribution of each parameter 20 in predicting the subject’s condition, providing an overall view of parameter importance. This approach may enhance interpretability of the predicted condition, addressing the limitations of traditional "black box" models by offering transparency in the decision-making process.

[0081] According to some embodiments, regression model 300 may apply a machine-learning (ML) based regression, or classification algorithm on global significance matrix 110GM, to produce a prediction 300P of a condition of the subject. As explained herein, regression model 300 may ensure that prediction 300P is both accurate and interpretable, providing clear explanation of how each parameter may influence the outcome.

[0082] For example, during a training stage, system 10 may receive a training dataset 300DS that may include a plurality of annotated patient parameters 20. The patient parameters 20 may be annotated in a sense that they may be associated with respective annotations or labels, which may indicate ground-truth values of the subject’s condition. Model 300 may be trained using a supervised regression objective. For each input sample, regression model 300 may output a single continuous prediction. The training may minimize a Mean Squared Error (MSE) between the predicted values and the ground-truth values of the subject's condition.

[0083] Pertaining to the example of predicting a number of oocytes during an IVF treatment cycle, annotations of parameters 20 in the training dataset 300DS may include ground truth numbers of eventually extracted oocytes.

[0084] As known in the art, system 10 may subsequently utilize a training scheme (e.g., a backward propagation scheme), to train model 300 so as to produce interim predictions 300P, while using training dataset 300DS as supervisory information.

[0085] Additionally, or alternatively, one or more (e.g., all) components of the neural network, including embeddings 200, context-aware attention models 100, and regression model 300, may be trained end-to-end using gradient descent via backpropagation. The training data may be splitinto mini-batches, and regression model 300 may produce predictions 300P for each batch. The MSE loss for each batch may be computed, and gradients of the loss with respect to all model parameters may be computed using backpropagation. The parameters may be updated using an optimizer, such as the Adam optimizer. Additionally, or alternatively, a cyclical learning-rate scheduler may be used, which may oscillate the learning rate between a minimum and maximum value to improve training stability and reduce overfitting. Additionally, or alternatively, gradient clipping may be applied to ensure training stability, wherein gradient norms may be clipped to a predetermined threshold value. Additionally, or alternatively, early stopping may be applied based on validation MSE to prevent overfitting.

[0086] In a subsequent, inference stage, regression model 300 may be configured to receive parameters 20 of target patients. Regression model 300 may then produce prediction 300P of the condition of the subject, based on its training.

[0087] According to some embodiments, regression model 300 may apply ML-based regression model 300 further on the embedding matrix 200EM, to predict the condition of the subject.

[0088] Additionally, or alternatively, system 10 may provide global significance matrix 110GM as an interpretation of prediction 300P. It may be appreciated that system 10 may thereby support clinical decision-making by offering a data-driven, objective approach to predicting the subject’s condition.

[0089] Following is a formal mathematical explanation, elaborating on the processes of feature selection 120 and feature interaction 110, that may be employed by embodiments of the invention. Definition of symbols used for this explanation are found in Table 1, below:Table 1

[0090] According to some embodiments, feature selection module 120 of a specific stage i may initially calculate prior P[i] according to Equation 1, below:where the function ones_like(e) outputs a matrix of all ones, having the same shape as that of e.

[0091] Prior matrix P[i] may be intuitively understood as a record of entries 200ENT of embedding matrix e (200EM) that were selected by selection matrices M[i] of previous stages in the cascade. The y parameter may be understood as a diminishing factor, which controls the likelihood of entries 200ENT of embedding matrix e (200EM), that were selected in previous stages, to be re-selected in a current stage of the cascade.

[0092] Feature selection module 120 may calculate mask M[i] of each stage i based on Equation Eq. 2, below:Eq. 2M[i = sparsemax(P[i — 1] • / ij(a[i — 1]))

[0093] In the first stage (e.g., Si), feature selection module 120 may define a[i- 1] as equal to e. In subsequent stages (e.g., S2...Sn), feature selection module 120 of each stage i may use prior matrix P[i- 1 ] of the previous stage, and the parameters selected in previous stages a[i- 1 ] , to calculate mask M[i] of the current stage i.

[0094] According to some embodiments, h[i] may be a single FC layer with batch normalization (normalization of the output values) and a non-linear activation function. The input for h[i] may be the selected parameters 20 of the previous stage (a[i]).

[0095] Eq. 2 may thereby allow feature selection module 120 to create a new mask selection matrix M[i] (120M), for a current stage i, while utilizing the knowledge of (a) which entries 200ENT ofembedding matrix 200EM were already selected in all the previous stages (through P[i-1]), and (b) which features were selected in the previous step (through a[i- 1 ]).

[0096] As known in the art, Sparsemax is a mathematical function used in machine learning, particularly in attention mechanisms, to transform a vector of values into a probability distribution. Unlike the traditional Softmax function, which produces dense probability distributions, Sparsemax generates sparse probability distributions, meaning that some of the output values are exactly zero. This sparsity can enhance both performance and interpretability by focusing on the most relevant features and ignoring the less important ones.

[0097] In the context of the present invention, Sparsemax may be used to project the mask’s values onto the probabilistic simplex, while encouraging sparsity. This may help in selecting the most significant features while zeroing out the less important ones, thereby improving the model's performance and making the results more interpretable.

[0098] As known in the art, the term “attention” may be used in reference to a mechanism that dynamically determines the importance of different parts of input data by comparing a query (representing a "first" feature or focus point) to a set of keys (representing "second" features or candidate context points). The query may be compared to each key using a similarity measure, and the resulting scores are used to compute attention weights. These weights indicate how much each key contributes to the final representation. For example, in text processing, the query could be a word or token, and the keys could be other words in the sequence, allowing the model to focus on contextually relevant parts of the input when processing or generating language. This selective focus enables improved performance in tasks like translation, summarization, and understanding complex relationships. A “classical” representation of an attention algorithm may be described by equation Eq. 3, below:Eq. 3A[i] = Softmax Q KT) • V where Q, K and V represent Query, Key and Value projections of the attention algorithm A[i], respectively.

[0099] According to some embodiments, feature interaction module 110 may implement attention-based algorithm A[i] based on a reduced version of Eq. 3, e.g., when Q=K=V, and all are equal to the selected embeddings (i.e., M[i] • e).

[0100] Feature interaction 110 may thus apply attention-based algorithm A[i] on the masked embeddings, to obtain a[i] (ad-hoc significance matrix 110M), according to Equation Eq. 4 below: Eq. 4 a[i] = A[i](M[i] • e)

[0101] As seen in Eq. 4, mask matrix M[i] (120M) may be applied to the entries 200ENT of embedding matrix e (200EM), effectively zeroing out some of them. This application may influence the subsequent attention calculation A[i], which operates on the rows of the embedding matrix 200EM (i.e., on embedding vectors 200EV). The attention calculation algorithm A[i] of feature interaction module 110 may perform a dot-product between pairs of parameters 20 (e.g., represented as embedding vectors 200EV in embedding matrix 200EM).

[0102] As explained herein, ad-hoc significance matrix a[i] (110M) may represent weights of a subset of parameters 20, as selected by mask matrix M[i] (120M) of the stage i in the cascade.

[0103] A higher dot product value may indicate greater utilization of that parameter 20 in the ad- hoc significance matrix a[i] (110M) current step. The dot product value may be determined by the entries 200ENT that are not masked for that parameter 20. For instance, if a first embedding vector 200EV has its first half of entries 200ENT zeroed out, and a second embedding vector 200EV has its second half zeroed out, the resulting dot product of multiplications between these two embedding vectors 200EV may be zero. Thus, mask matrix M[i] (120M) may dictate the dotproduct calculation.

[0104] According to some embodiments, system 10 may present the results of prediction 300P, and interpretation 110GM via a User Interface 40 (UI, such as input 7 and output 8 of Fig. 1), to be studied by a physician or care giver, as elaborated herein (e.g., in relation to Fig. 3).

[0105] Reference is now made to Fig. 3, which is a schematic chart, showing results of predictions of conditions of subjects in two cases, denoted “Patient A” and “Patient B”, according to some embodiments.

[0106] For each of the patients, the type, and value of each parameter is provided in a parameter table. For example, the parameters 20 of patient A include her age (28.9 years), her stimulation regimen (antagonist), etc. The parameters 20 of patient B include different parameter 20 values such her age (29.5 years), her stimulation regimen (long agonist), etc.

[0107] The prediction 300P of a number of extracted oocytes is also presented in the respective tables: For patient A, the predicted 300P number was 16, whereas the ground-truth number (e.g.,the number of eventually extracted oocytes) was 20. In the case of patient B, the predicted 300P number was 4, whereas the ground-truth number was eventually 3.

[0108] The bar-charts on the right-hand side of Fig. 3 present the global significance matrices 110GM of each of the patients, and demonstrates the aspect of interpretability provided by embodiments of the invention: In the case of patient A, the larger follicles (16-17 mm and 18+mm) found in an ultrasound test were key contributors in the regression model’s 300 prediction 300P of 16 extracted oocytes. In the case of patient B the predominant contributor to the prediction 300P was the stimulation regimen.

[0109] It may be appreciated that a user (e.g., a physician) may therefore use UI 40 to view information as presented in Fig. 3. The user may then utilize prediction 300P of system 10, to determine an optimal timing to trigger the IVF treatment, e.g., to result in a maximal amount of treatable oocytes. Additionally, the user may also use the presentation of global significance matrix 110GM as in Fig. 3, to analyze, or study the causes for the predicted value 300P, and possibly apply changes to the scheduled treatment accordingly.

[0110] Reference is now made to Fig. 4, which is a flow diagram, depicting a method of providing interpretable prediction of a condition of a subject, by at least one processor (e.g., processor 2 of Fig. 1), according to some embodiments of the invention.

[0111] As shown on step S1005, the at least one processor 2 may receive data (e.g., patient data 20 of Fig. 2). Data 20 may include values of patient parameters, representing a patient’s or subject’s current physiology.

[0112] As shown on step S1010, the at least one processor 2 may employ an embedding model (e.g., embedding model 200 of Fig. 2), to calculate an embedding matrix (e.g., 200EM of Fig. 2). Embedding matrix 200EM may representing the patient parameters’ data 20 in an embedding space.

[0113] As shown on step S 1015, the at least one processor 2 may process embedding matrix 200EM through a cascade of stages (e.g., cascade of stages SI, S2, ...,Sn, as in Fig. 2). Each stage (SI, S2, ...,Sn) may include a respective ML based, context-aware attention model (e.g., Context- aware attention model 100 of Fig. 2). Each context-aware attention model may be configured to generate an ad-hoc significance matrix (e.g., 110M of Fig. 2) pertaining to that stage.

[0114] As shown on step S 1020, the at least one processor 2 may aggregate the ad-hoc significance matrices 110M of stages (e.g., all stages SI, S2, ...,Sn) of the cascade of stages, to obtain a globalsignificance matrix (e.g., 110GM of Fig. 2). Global significance matrix 110GM may represent contribution of each of said parameters in predicting the subject’s condition.

[0115] As shown on step S 1025, the at least one processor 2 may apply an ML-based classification or regression model (e.g., regression model 300 of Fig. 2) on the global significance matrix 110GM, to predict (e.g., produce prediction 300P of Fig. 2) the condition of the subject. The at least one processor 2 may further provide the global significance matrix 110GM as an interpretation of prediction 300P (step S1030).

[0116] As explained herein, every stage (Si, ...Sn) of the cascade may realize a synergistic relation between feature selection 120 and feature interaction 110 to provide a mechanism of context-aware attention, resulting in an interpretable selection of parameters 20 for performing a classification, or regression task 300.

[0117] On one hand, feature selection module 120 may produce a selection mask matrix 120M (M[i], e.g., according to Eq. 2), which represents a choice of entries 200ENT in embedding matrix 200EM as features of interest pertaining to the current stage (e.g., Si). Feature interaction 110 may use the selected entries 200ENT, as expressed by mask matrix 120M, to calculate an ad-hoc significance matrix a[i] (110M) (e.g., according to Eq. 4), which represents the most significant parameters 20 in terms of self-attention as known in the art, in the context of that stage (e.g., Si).

[0118] On the other hand, feature selection module 120 of the following stage (e.g., S2) may use significance matrix a[i] (110M) of the previous stage (e.g., Si), e.g., according to Eq. 1 and Eq. 2), to select new entries 200ENTof embedding matrix 200EM, while suppressing re-selection of previously selected entries 200ENTof embedding matrix 200EM.

[0119] Currently available systems for multi-stage feature selection also exist in the art. Some publications regarding these systems even use the term “attention” to describe components of these systems. However, as known to persons skilled in the art, this use of the term “attention” in these publications is erroneous, as it only reflects a mechanism for selecting entries in each decision step, and does not reflect the mathematical process of attention, as commonly referred to in the art, e.g., by the seminal article “Attention Is All You Need” (http8: / / anfiy.org / pdf / l 71)6.(8762), and as described herein.

[0120] For example, currently available systems for multi-stage feature selection do not teach embedding of numerical, continuous parameters 20, as elaborated herein. Currently available systems thereby disallow production of vector representations for continuous numerical values,and therefore cannot provide a context-based attention mechanism for calculating ad-hoc significance matrix a[i] (HOM), comprising weights of parameter significance, as explained herein.

[0121] Moreover, currently available systems for multi-stage feature selection do not include integration of an attention-based mechanism into each the feature selection stages, and may therefore do not provide the benefit of synergy between selection of entries (e.g., by feature selection 120), and the attention-based weighing of parameters (e.g., by feature interaction 110), as elaborated herein.

[0122] Reference is now made to Fig. 5, which is a flow diagram, depicting a method of providing an interpretable prediction of a number of mature oocytes available for retrieval during In Vitro Fertilization (IVF) treatment, by at least one processor (e.g., processor 2 of Fig. 1), according to some embodiments of the invention.

[0123] As shown on step S2005, the at least one processor 2 may receive data (e.g., patient data 20 of Fig. 2). Data 20 may include values of patient parameters, representing a subject's current physiology during ovarian stimulation. The patient parameters may include, for example, one or more categorical parameters selected from a list consisting of: an ovarian stimulation regimen, a diagnosis of diminished ovarian reserve, endometriosis, a male factor, ovulatory dysfunction, a tubal disease, and a uterine factor. Additionally, or alternatively, the patient parameters may include one or more continuous parameters selected from a list consisting of: age, Body Mass Index (BMI), basal Follicle-Stimulating Hormone (FSH) level, trigger estradiol (E2) level, endometrial thickness on trigger day, a number of stimulation days, follicle counts by size range, average follicle size, and total follicle count.

[0124] As shown on step S2010, the at least one processor 2 may employ an embedding model (e.g., embedding model 200 of Fig. 2), to calculate an embedding matrix (e.g., 200EM of Fig. 2). Embedding matrix 200EM may represent the patient parameters' data 20 in an embedding space.

[0125] As shown on step S2015, the at least one processor 2 may process embedding matrix 200EM through a cascade of stages (e.g., cascade of stages SI, S2, ...,Sn, as in Fig. 2). Each stage (SI, S2, ...,Sn) may include a respective ML based, context-aware attention model (e.g., Context- aware attention model 100 of Fig. 2). Each context-aware attention model may include a feature selection model configured to calculate a respective mask matrix representing a selection of parameter entries of the embedding matrix as features of interest at the respective stage, and afeature interaction model configured to apply a self-attention algorithm on selected parameter entries to generate an ad-hoc significance matrix (e.g., 110M of Fig. 2) pertaining to that stage.

[0126] As shown on step S2020, the at least one processor 2 may aggregate the ad-hoc significance matrices 110M of stages (e.g., all stages SI, S2, ...,Sn) of the cascade of stages, to obtain a global significance matrix (e.g., 110GM of Fig. 2). Global significance matrix 110GM may represent contribution of each of said parameters in predicting the number of mature oocytes.

[0127] As shown on step S2025, the at least one processor 2 may apply an ML-based regression model (e.g., regression model 300 of Fig. 2) on the global significance matrix 110GM, to predict (e.g., produce prediction 300P of Fig. 2) the number of mature oocytes available for retrieval in relation to a predetermined trigger date. The at least one processor 2 may further provide the global significance matrix 110GM as an interpretation of prediction 300P to support clinical decisionmaking regarding trigger timing for oocyte retrieval (step S2030).

[0128] According to some embodiments, providing the global significance matrix 110GM as an interpretation may include presenting the contribution of each patient parameter 20 in a visually interpretable format via a user interface (UI, e.g., output 8 of computing device 1 of Fig. 1).

[0129] For example, the at least one processor 2 may generate a bar chart displaying the percentage contribution of each parameter 20 to the prediction 300P, enabling a clinician to quickly identify which parameters most strongly influenced the predicted number of mature oocytes for that particular patient. Additionally, or alternatively, the at least one processor 2 may generate a heatmap visualization of the global significance matrix 110GM, wherein colors or shading intensity may represent the magnitude of each parameter's contribution, with darker colors indicating negative contributions and lighter colors indicating positive contributions to the predicted oocyte yield. Additionally, or alternatively, the at least one processor 2 may highlight or emphasize parameters whose contribution exceeds a predetermined threshold or differs significantly from average contribution values across a patient population, thereby alerting the clinician to patient-specific factors that may warrant particular attention in determining optimal trigger timing. For example, if estradiol levels at trigger contribute 20% to the prediction for a particular patient compared to a 9% average contribution, this parameter may be visually highlighted or flagged on the user interface to indicate its elevated importance in that specific case.

[0130] As elaborated herein, embodiments of the present invention provide a technological improvement to the field of assistive diagnostics by enabling accurate, interpretable predictions ofa subject's condition through a novel deep learning architecture that addresses fundamental limitations of existing diagnostic support systems. Embodiments of the invention may transform subjective clinical decision-making into an objective, data-driven process while maintaining transparency in the reasoning behind predictions, thereby enabling clinical validation that was not previously possible with conventional "black box" machine learning models.

[0131] The technological advancement provided by embodiments of the invention may be understood in the context of IVF treatment optimization. Traditional clinical practice relies on subjective assessment of multiple parameters including patient age, hormone levels, follicle counts, and treatment protocols to determine optimal trigger timing for oocyte retrieval. This subjective approach may lead to variability in outcomes and suboptimal oocyte yields. Prior machine learning approaches, while offering improved prediction accuracy, functioned as opaque "black boxes" that provided predictions without explanation, making them unsuitable for clinical adoption where understanding the rationale behind medical decisions is necessary for obtaining successful outcomes.

[0132] Embodiments of the present invention overcome these limitations by providing a system that not only predicts outcomes with high accuracy but also generates a global significance matrix 110GM that quantifies the contribution of each patient parameter 20 to the prediction. This dual capability represents a significant technological advancement over prior art. For example, in predicting the number of mature oocytes available for retrieval, the system may identify that for a particular patient, estradiol levels at trigger contribute 20% to the prediction compared to a 9% average contribution, while for another patient with elevated basal FSH, that parameter may contribute 11% compared to a 5% average. This parameter-specific, patient-specific interpretability enables clinicians to validate predictions against established medical knowledge and identify cases where the model may not have adequately considered all relevant factors.

[0133] The invention may not be performed as a mental process or with pen and paper due to the computational complexity and scale of the operations required. The embedding model 200 transforms multiple patient parameters 20 into high-dimensional embedding vectors 200EV, where each parameter may be represented by dozens or hundreds of numerical values in embedding space. The cascade of context-aware attention models 100 performs iterative, nonlinear transformations on these embeddings, with each stage calculating mask matrices 120M through Sparsemax projections and generating ad-hoc significance matrices 110M through self-attention algorithms involving matrix multiplications and dot products across all parameter combinations. These operations may involve millions of numerical calculations per prediction, operating on matrices with dimensions that scale with the number of parameters and embedding dimensions.

[0134] Furthermore, the training process requires processing thousands of patient records through the neural network architecture, computing gradients via backpropagation through multiple layers, and iteratively updating potentially thousands of model parameters using optimization algorithms such as Adam optimizer further render mental computation entirely impractical for clinical use where timely decisions are necessary.

[0135] The practical application of the present invention as a tool that improves current technology of assistive diagnostics may be demonstrated through several concrete improvements. First, the invention enables more accurate prediction of treatment outcomes compared to existing machine learning approaches. For example, in predicting mature oocyte yield during IVF cycles, embodiments of the invention have been shown to achieve a mean absolute error that was superior to other, currently applied approaches.

[0136] Additionally, embodiments of the invention have provided interpretability that enabled the inventors to identify clinically relevant patterns that were previously hidden in "black box" models. For example, through analysis of global significance matrices 110GM across multiple patients, embodiments of the invention have revealed that follicles in the 12- 15mm size range consistently exhibit negative correlation with mature oocyte yield, suggesting these follicles have lower likelihood of producing mature oocytes. This insight, derived from the model's transparent decision-making process, may inform clinical strategies for optimizing trigger timing and may guide future research into the relationship between follicle size distribution and oocyte maturation.

[0137] In another example, embodiments of the invention have revealed that basal FSH may exert greater influence on outcomes when values are elevated rather than when they fall below average, providing clinically actionable information about which patient populations may benefit from particular interventions.

[0138] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Furthermore, all formulas described herein are intended as examples only and other or different formulas may be used. Additionally, some of the described method embodiments or elements thereof may occur or be performed at the same point in time.

[0139] While certain features of the invention have been illustrated and described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention.

[0140] Various embodiments have been presented. Each of these embodiments may of course include features from other embodiments presented, and embodiments not specifically described may include various features described herein.

Claims

CLAIMS1. A method of providing an interpretable prediction of a condition of a subject, the method comprising: receiving data comprising values of patient parameters, representing a subject’s current physiology; calculating an embedding matrix, representing said data in an embedding space; processing said embedding matrix through a cascade of stages, each comprising a respective Machine-Learning (ML) based, context-aware attention model, wherein each context- aware attention model is configured to generate an ad-hoc significance matrix pertaining to that stage; aggregating the ad-hoc significance matrices of said stages to obtain a global significance matrix, representing contribution of each of said parameters in predicting the subject’s condition; applying an ML-based regression model on the global significance matrix, to predict the condition of the subject; and providing the global significance matrix as an interpretation of said prediction.

2. The method of claim 1 , further comprising applying the ML-based regression model further on the embedding matrix, to predict the condition of the subject.

3. The method according to any one of claims 1-2, wherein the condition of the subject is a number of mature oocytes that are available for retrieval from the subject, during an In Vitro Fertilization (IVF) treatment, in relation to a predetermined trigger date.

4. The method of claim 3, wherein said patient parameters comprise a first set of categorical parameters pertaining to ovarian stimulation regimen, said parameters selected from a list consisting of: an antagonist regimen, a long agonist regimen, a microflare regimen, and an Ovulation Induction (01) as a preparatory step for IVF.

5. The method according to any one of claims 3-4, wherein said patient parameters comprise a second set of categorical parameters pertaining to patient diagnoses, said parameters selectedfrom a list consisting of: diminished ovarian reserve, endometriosis, a male factor, ovulatory dysfunction, a tubal disease, and a uterine factor.

6. The method according to any one of claims 3-5, wherein said patient parameters comprise one or more continuous parameters selected from a list consisting of: age, Body Mass Index (BMI), basal Follicle-Stimulating Hormone (FSH) level, trigger estradiol (E2) level, endometrial thickness on trigger day, and a number of stimulation days during a current IVF cycle.

7. The method according to any one of claims 1-6, further comprising: for each patient parameter, calculating a parameter embedding vector having a plurality of parameter entries; and aggregating the parameter embedding vectors of all patient parameters, to obtain the embedding matrix.

8. The method of claim 7, wherein calculating a patient parameter’s embedding vector comprises multiplying that parameter’s value with a respective, uniquely indicative vector, representing a type of that patient parameter.

9. The method according to any one of claims 1-8 wherein each context-aware attention model comprises a feature selection model, configured to calculate a respective mask matrix, representing a selection of parameter entries of the embedding matrix as features of interest, at the respective stage.

10. The method of claim 9, wherein a feature selection model of a current stage of the cascade is adapted to: receive one or more previous mask matrices, pertaining to respective, previous stages of the cascade; and calculate the mask matrix of the current stage based on the one or more mask matrices of previous stages, wherein said calculation comprises suppression of re-selection of parameter entries that were already selected by the mask matrices of previous stages.

11. The method of claim 10, wherein each context-aware attention model further comprises a feature interaction model, and wherein the feature interaction model of the current stage is configured to: obtain the selection of parameter entries, based on the mask matrix of the current stage; apply a self-attention algorithm on the selected parameter entries, so as to assign respective weights to one or more parameter embedding vectors of the embedding matrix; and generate the ad-hoc significance matrix of the current stage based on said weights of parameter embedding vectors.

12. A system for providing an interpretable prediction of a condition of a subject, the system comprising: a non-transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to: receive data comprising values of patient parameters, representing a subject’s current physiology; calculate an embedding matrix, representing said data in an embedding space; process said embedding matrix through a cascade of stages, each comprising a respective ML based, context-aware attention model, wherein each context-aware attention model is configured to generate an ad-hoc significance matrix pertaining to that stage; aggregate the ad-hoc significance matrices of said stages to obtain a global significance matrix, representing contribution of each of said parameters in predicting the subject’s condition; apply an ML-based regression model on the global significance matrix, to predict the condition of the subject; and provide the global significance matrix as an interpretation of said prediction.

13. The system of claim 12, wherein the at least one processor is further configured to apply the ML-based regression model further on the embedding matrix, to predict the condition of the subject.

14. The system according to any one of claims 12-13, wherein the condition of the subject is a number of mature oocytes that are available for retrieval from the subject, during an IVF treatment, in relation to a predetermined trigger date.

15. The system of claim 14, wherein said patient parameters comprise a first set of categorical parameters pertaining to ovarian stimulation regimen, said parameters selected from a list consisting of: an antagonist regimen, a long agonist regimen, a microflare regimen, and an 01, as a preparatory step for IVF.

16. The system of claim 14-15, wherein said patient parameters comprise a second set of categorical parameters pertaining to patient diagnoses, wherein said second set of categorical parameters is selected from a list consisting of: diminished ovarian reserve, endometriosis, a male factor, ovulatory dysfunction, a tubal disease, and a uterine factor.

17. The system of claim 14-16, wherein said patient parameters comprise one or more continuous parameters selected from a list consisting of: age, BMI, basal FSH level, trigger estradiol (E2) level, endometrial thickness on trigger day, and a number of stimulation days during a current IVF cycle.

18. The system of claim 12-17, wherein the at least one processor is further configured to: for each patient parameter, calculate a parameter embedding vector having a plurality of parameter entries; and aggregate the parameter embedding vectors of all patient parameters, to obtain the embedding matrix.

19. The system of claim 18, wherein the at least one processor is further configured to calculate a patient parameter’s embedding vector by multiplying that parameter’s value with a respective, uniquely indicative vector, representing a type of that patient parameter.

20. The system of claim 12-19 wherein each context-aware attention model comprises a feature selection model, configured to calculate a respective mask matrix, representing a selection of parameter entries of the embedding matrix as features of interest, at the respective stage.

21. The system of claim 20, wherein the at least one processor is further configured to employ a feature selection model of a current stage of the cascade to: receive one or more previous mask matrices, pertaining to respective, previous stages of the cascade; and calculate the mask matrix of the current stage based on the one or more mask matrices of previous stages, wherein said calculation comprises suppression of re-selection of parameter entries that were already selected by the mask matrices of previous stages.

22. The system of claim 21, wherein each context-aware attention model further comprises a feature interaction model, and wherein the at least one processor is further configured to employ a feature interaction model of the current stage to: obtain the selection of parameter entries, based on the mask matrix of the current stage; apply a self- attention algorithm on the selected parameter entries, so as to assign respective weights to one or more parameter embedding vectors of the embedding matrix; and generate the ad-hoc significance matrix of the current stage based on said weights of parameter embedding vectors.

23. A method of providing an interpretable prediction of a number of mature oocytes available for retrieval during In Vitro Fertilization (IVF) treatment, by at least one processor, the method comprising: receiving data comprising values of patient parameters representing a subject undergoing ovarian stimulation; calculating an embedding matrix, representing said data in an embedding space; processing said embedding matrix through a cascade of stages, each comprising a respective Machine-Learning (ML) based, context-aware attention model, wherein each context-aware attention model is configured to generate an ad-hoc significance matrix pertaining to that stage;aggregating the ad-hoc significance matrices of said stages to obtain a global significance matrix, representing contribution of each of said parameters in predicting the number of mature oocytes; applying an ML-based regression model on the global significance matrix to predict the number of mature oocytes available for retrieval in relation to a predetermined trigger date; and providing the global significance matrix as an interpretation of said prediction to support clinical decision-making regarding trigger timing for oocyte retrieval.

24. The method of claim 23, wherein said patient parameters comprise: (i) one or more categorical parameters selected from a list consisting of: an ovarian stimulation regimen, a diagnosis of diminished ovarian reserve, endometriosis, a male factor, ovulatory dysfunction, a tubal disease, and a uterine factor; and (ii) one or more continuous parameters selected from a list consisting of: age, Body Mass Index (BMI), basal Follicle-Stimulating Hormone (FSH) level, trigger estradiol (E2) level, endometrial thickness on trigger day, a number of stimulation days, follicle counts by size range, average follicle size, and total follicle count.

25. The method of claim 23, wherein each context-aware attention model comprises: a feature selection model configured to calculate a respective mask matrix representing a selection of parameter entries of the embedding matrix as features of interest at the respective stage; and a feature interaction model configured to apply a self-attention algorithm on selected parameter entries to generate the ad-hoc significance matrix of that stage.